Method and System for Constructing a Log Storage System Based on Multi-Dimensional Logical Relationship Interaction

By performing multi-dimensional analysis and classified storage of log files, the problem of waste of storage space and low retrieval efficiency in the log storage system is solved, and efficient log storage and retrieval is achieved.

CN119621672BActive Publication Date: 2025-07-08CHINA NAT BUILDING MATERIAL GRP FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411700974.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-07-08
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing log storage systems fail to distinguish the importance and priorities of logs, resulting in wasted storage space and inefficient retrieval.

Method used

By receiving log storage instructions, analyzing the original log set into a cube, calculating user weights, time weights, and position weights, classifying log files into common and backlog files according to log priority thresholds, and building a log graph database and a compressed database.

Benefits of technology

Save storage space, improve log retrieval efficiency, and optimize log storage and retrieval process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621672B_ABST
    Figure CN119621672B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information technology, and a method and system for constructing a log storage system based on multi-dimensional logical relationship interaction, including: receiving a log storage instruction, collecting an original log set based on the log storage instruction, performing a parsing operation on the original log set to obtain a multi-dimensional data set, calculating user weights according to the multi-dimensional data and the priority user sequence, calculating time weights according to the transaction time and the priority time sequence in the multi-dimensional data, obtaining the headquarters location, calculating location weights according to the initiating location and the headquarters location in the multi-dimensional data, calculating the log priority according to the user weights, time weights, location weights and transaction amount values, performing a classification and compression operation on the backlog log set to obtain a log compression package, and constructing a log graph database by using the common log set to complete the construction of the log storage system. The present invention can save the storage space of logs and improve the retrieval efficiency of logs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a method, a system, an electronic device, and a computer-readable storage medium for constructing a log storage system based on multi-dimensional logical relationship interaction. Background Art

[0002] In modern information society, with the expansion of business scale and the popularization of network services, a large amount of log data has been generated in various complex business scenarios. Logs are important texts that record information such as system operation and business transactions, and are crucial for system monitoring and analysis of transaction information.

[0003] Currently, log storage systems usually simply record information on system operation or business transactions based on time or event sequence, and directly store them in the memory in the form of original log files.

[0004] Although existing log storage systems can complete the storage of logs, they treat different logs equally when storing, without distinguishing the importance and priority of logs. For example, logs with higher transaction amounts or involving key operations need to be stored preferentially, while ordinary logs should be reasonably archived or compressed. At the same time, key information in the logs is not extracted during storage, and a large amount of redundant text data is retained, consuming more storage space and affecting the retrieval speed. Summary of the Invention

[0005] The present invention provides a method and a computer-readable storage medium for constructing a log storage system based on multi-dimensional logical relationship interaction, and its main purpose is to save the storage space of logs and improve the retrieval efficiency of logs.

[0006] To achieve the above object, a method for constructing a log storage system based on multi-dimensional logical relationship interaction provided by the present invention includes:

[0007] Receiving a log storage instruction, collecting an original log set based on the log storage instruction, and performing a parsing operation on the original log set to obtain a multi-dimensional data set, where the original log set includes a plurality of original log files, the multi-dimensional data set includes a plurality of multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. The multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value;

[0008] Obtaining a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively obtaining a preferred user sequence and a preferred time sequence by using the a user retrieval records and the b time retrieval records;

[0009] Performing the following operations on each multi-dimensional data in the multi-dimensional data set:

[0010] Calculate the user weight according to the multi-dimensional data and the priority user sequence, and calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence;

[0011] Obtain the headquarters location, and calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location;

[0012] Calculate the log priority according to the user weight, time weight, location weight and transaction amount value;

[0013] Compare the log priority with the preset priority threshold. If the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a common log file. If the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlog log file;

[0014] Aggregate the common log files and the backlog log files respectively to obtain a common log set and a backlog log set;

[0015] Perform a classification and compression operation on the backlog log set to obtain a log compression package, and construct a log graph database using the common log set;

[0016] Import the log compression package into a pre-constructed basic server to obtain a compressed database, and import the log graph database into a pre-constructed high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

[0017] Optionally, the collection of the original log set based on the log storage instruction includes:

[0018] Send the log storage instruction to a pre-constructed log management system, where the log management system includes: a log database;

[0019] When the log management system receives the log storage instruction, extract the general log set from the log database, where the general log set includes multiple general log files;

[0020] Perform the following operations on each general log file in the multiple general log files:

[0021] Identify the log storage capacity of the general log file, and compare the log storage capacity with the preset storage capacity threshold;

[0022] If the log storage capacity is greater than the storage capacity threshold, mark the general log file as a basic log file;

[0023] Aggregate the basic log files to obtain multiple basic log files;

[0024] Perform the following operations on each basic log file in the multiple basic log files:

[0025] Read the log text of the basic log file, and retrieve whether a preset first keyword or a preset second keyword exists in the log text;

[0026] If the first keyword or the second keyword exists in the log text, mark the basic log file corresponding to the log text as the original log file;

[0027] Summarize the original log files to obtain the original log set.

[0028] Optionally, the obtaining of the priority user sequence and the priority time sequence by using a user retrieval records and b time retrieval records respectively includes:

[0029] Identify a initial user IDs in the a user retrieval records, and perform a uniquification operation on the a initial user IDs to obtain A unique user IDs, where the user retrieval records and the initial user IDs are in one-to-one correspondence, and a≥A;

[0030] Perform the following operations on each of the A unique user IDs:

[0031] Based on the unique user ID and the a initial user IDs, confirm the user retrieval quantity, where the user retrieval quantity is the number of initial user IDs that are the same as the unique user ID among the a initial user IDs;

[0032] Sort the A unique user IDs in descending order according to the user retrieval quantity corresponding to the unique user ID to obtain the priority user sequence;

[0033] Identify b retrieval times in the b time retrieval records, and based on the b retrieval times, confirm the earliest time and the latest time, where the time retrieval records and the retrieval times are in one-to-one correspondence, the earliest time is the earliest retrieval time among the b retrieval times, and the latest time is the latest retrieval time among the b retrieval times;

[0034] Based on the earliest time and the latest time, confirm the retrieval time range, where the minimum value of the retrieval time range is the earliest time, and the maximum value of the retrieval time range is the latest time;

[0035] Perform a uniform sampling operation on the retrieval time range by using a preset time sampling interval to obtain B sampled time ranges;

[0036] Perform the following operations on each of the B sampled time ranges:

[0037] Based on the sampled time range and the b retrieval times, confirm the time retrieval quantity, where the time retrieval quantity is the number of retrieval times that are within the sampled time range among the b retrieval times;

[0038] Perform a sorting operation on B sampling time ranges in the order of the number of retrieved times corresponding to the sampling time ranges from more to less to obtain a preferred time series.

[0039] Optionally, the calculating the user weight according to the multi-dimensional data and the preferred user sequence includes:

[0040] Judge whether there is a unique user ID in the preferred user sequence that is the same as the initiating user ID in the multi-dimensional data;

[0041] If there is a unique user ID in the preferred user sequence that is the same as the initiating user ID, use the unique user ID in the preferred user sequence that is the same as the initiating user ID as the target user ID;

[0042] If there is no unique user ID in the preferred user sequence that is the same as the initiating user ID, judge whether there is a unique user ID in the preferred user sequence that is the same as the receiving user ID in the multi-dimensional data;

[0043] If there is a unique user ID in the preferred user sequence that is the same as the receiving user ID, use the unique user ID in the preferred user sequence that is the same as the receiving user ID as the target user ID;

[0044] Retrieve the user level number from the pre-constructed user database based on the target user ID;

[0045] Confirm the user ranking value in the preferred user sequence based on the target user ID, where the user ranking value is the position of the target user ID in the preferred user sequence;

[0046] Calculate the user weight according to the user ranking value and the user level number, and the calculation formula is as follows:

[0047]

[0048] Wherein, is the user weight, is the user ranking value, is the user level number, is the preset maximum user level number, is the hyperbolic tangent function;

[0049] If there is no unique user ID in the preferred user sequence that is the same as the receiving user ID, record the user weight as 0.

[0050] Optionally, the calculating the time weight according to the transaction time in the multi-dimensional data and the preferred time series includes:

[0051] Judge whether the transaction time is within the retrieved time range;

[0052] If the transaction time is within the retrieval time range, the following operations are sequentially performed for each of the B sampling time ranges:

[0053] Determine whether the transaction time is within the sampling time range. If the transaction time is within the sampling time range, use the sampling time range as the target time range;

[0054] Based on the target time range, confirm the time ranking value in the priority time sequence, where the time ranking value is the position of the target time range in the priority time sequence;

[0055] Obtain the current time, and calculate the time difference according to the current time and the transaction time, where the time difference is the absolute difference between the current time and the transaction time;

[0056] Calculate the time weight according to the time difference and the time ranking value. The calculation formula is as follows:

[0057]

[0058] Where, is the time weight, is the time ranking value, is the time difference, is the natural constant;

[0059] If the transaction time is not within the retrieval time range, record the time weight as 0.

[0060] Optionally, the calculating the location weight according to the initiating location and the headquarters location in the multi-dimensional data includes:

[0061] Obtain the transaction distance using the initiating location and the headquarters location, and calculate the location weight using the transaction distance. The calculation formula is as follows:

[0062]

[0063] Where, is the location weight, is the transaction distance, is the preset standard distance.

[0064] Optionally, the calculation formula of the log priority is as follows:

[0065]

[0066] Where, is the log priority, is the transaction amount value, is the natural logarithm.

[0067] Optionally, performing a classification and compression operation on the backlog log set to obtain a log compression package, including:

[0068] Performing a merging operation on multiple backlog log files in the backlog log set to obtain multiple log folders, where each log folder in the multiple log folders includes: one or more backlog log files;

[0069] Obtaining multiple file storage amounts by using the multiple log folders, where the file storage amounts and the log folders correspond one by one;

[0070] Identifying the maximum storage amount and the minimum storage amount based on the multiple file storage amounts, where the maximum storage amount is the file storage amount with the largest storage capacity among the multiple file storage amounts, and the minimum storage amount is the file storage amount with the smallest storage capacity among the multiple file storage amounts;

[0071] Performing the following operations on each of the multiple log folders:

[0072] Calculating the compression level by using the maximum storage amount, the minimum storage amount, and the file storage amount corresponding to the log folder, and the calculation formula is as follows:

[0073]

[0074] Wherein, is the compression level, is the maximum storage amount, is the minimum storage amount, is the file storage amount corresponding to the log folder, is the preset maximum compression level, represents the ceiling operation;

[0075] Performing a compression operation on the log folder based on the compression level to obtain a compressed file;

[0076] Summarizing the compressed files to obtain a log compression package.

[0077] Optionally, constructing a log graph database by using the common log set, including:

[0078] Performing the following operations on each common log file in the common log set:

[0079] Combining the initiating user ID and the initiating location corresponding to the common log file into an initiating information group, and combining the receiving user ID and the receiving location corresponding to the common log file into a receiving information group;

[0080] Identifying both the initiating information group and the receiving information group as user information groups;

[0081] Summarizing the user information groups to obtain multiple user information groups;

[0082] Perform a uniquification operation on multiple user information groups to obtain multiple unique user information, and use the multiple unique user information to create multiple unique user nodes in a pre-constructed graph database construction software, where the unique user nodes and the unique user information are in one-to-one correspondence;

[0083] Perform the following operations on each common log file in the common log set:

[0084] Use the initiating user ID corresponding to the common log file as the indexing initiating ID, and use the indexing initiating ID to retrieve the initiating node among the multiple unique user nodes, where the initiating user ID or the receiving user ID in the unique user information corresponding to the initiating node is the same as the indexing initiating ID;

[0085] Obtain the receiving node based on the receiving user ID corresponding to the common log file;

[0086] Create a relationship edge in the graph database construction software based on the transaction time, transaction amount value, initiating node, and receiving node corresponding to the common log file;

[0087] Summarize the relationship edges to obtain multiple relationship edges, and confirm the log graph database based on the multiple unique user nodes and the multiple relationship edges.

[0088] To achieve the above object, the present invention also provides a log storage system construction system based on multi-dimensional logical relationship interaction, including:

[0089] A multi-dimensional data parsing module, configured to receive a log storage instruction, collect an original log set based on the log storage instruction, and perform a parsing operation on the original log set to obtain a multi-dimensional data set, where the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence, where the multi-dimensional data includes: initiating user ID, initiating location, receiving user ID, receiving location, transaction time, and transaction amount value;

[0090] A retrieval record analysis module, configured to obtain a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively use the a user retrieval records and the b time retrieval records to obtain a priority user sequence and a priority time sequence;

[0091] A priority log classification module is used to perform the following operations on each multi-dimensional data in a multi-dimensional dataset: calculate user weights based on the multi-dimensional data and the priority user sequence, calculate time weights based on the transaction time in the multi-dimensional data and the priority time sequence, obtain the headquarters location, calculate location weights based on the initiating location in the multi-dimensional data and the headquarters location, calculate the log priority based on the user weights, time weights, location weights, and transaction amount values, compare the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a frequently used log file. If the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlogged log file;

[0092] A log database construction module is used to separately summarize the frequently used log files and the backlogged log files to obtain a frequently used log set and a backlogged log set, perform a classification and compression operation on the backlogged log set to obtain a log compression package, construct a log graph database using the frequently used log set, import the log compression package into a pre-constructed basic server to obtain a compressed database, and import the log graph database into a pre-constructed high-speed server to obtain a high-speed database, thereby completing the construction of the log storage system.

[0093] To solve the above problems, the present invention also provides an electronic device, which includes:

[0094] A memory that stores at least one instruction; and a processor that executes the instruction stored in the memory to implement the above-mentioned method for constructing a log storage system based on multi-dimensional logical relationship interaction.

[0095] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned method for constructing a log storage system based on multi-dimensional logical relationship interaction.

[0096] To solve the problems described in the background art, the present invention receives a log storage instruction, collects an original log set based on the log storage instruction, and performs a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value. It can be seen that in the embodiment of the present invention, by parsing the original log files into multiple multi-dimensional data, redundant text data is eliminated, the storage space of the logs is saved, and at the same time, it is convenient to classify the original log files according to the multi-dimensional data subsequently, and then obtain a retrieval record set. Among them, the retrieval record set includes: a user retrieval records and b time retrieval records. The priority user sequence and the priority time sequence are obtained by using the a user retrieval records and the b time retrieval records respectively. It can be seen that in the embodiment of the present invention, by analyzing the historical retrieval records, reference data is provided for calculating the log priority subsequently, and the accuracy of classifying the original log files subsequently is improved. The following operations are performed on each multi-dimensional data in the multi-dimensional data set: calculate the user weight according to the multi-dimensional data and the priority user sequence, calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence, obtain the headquarters location, and calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location. It can be seen that in the embodiment of the present invention, by comprehensively analyzing the data in the three dimensions of the user, the transaction time, and the initiating location, the user weight, the time weight, and the location weight are confirmed, which is convenient for subsequently calculating the log priority comprehensively and accurately according to the user weight, the time weight, and the location weight. Calculate the log priority according to the user weight, the time weight, the location weight, and the transaction amount value, compare the log priority with the preset priority threshold. If the log priority is greater than or equal to the priority threshold, the original log file corresponding to the log priority is marked as a frequently used log file. If the log priority is less than the priority threshold, the original log file corresponding to the log priority is marked as a backlogged log file. It can be seen that in the embodiment of the present invention, the original log files are divided by the log priority and the priority threshold, the files that are frequently retrieved and important are marked as frequently used log files, and the files that are occasionally retrieved and unimportant are marked as backlogged log files, which is convenient for subsequently adopting targeted storage methods for the two types of files to save the storage space of the logs and improve the retrieval efficiency of the logs. The frequently used log files and the backlogged log files are summarized respectively to obtain a frequently used log set and a backlogged log set. A classification compression operation is performed on the backlogged log set to obtain a log compression package. A log graph database is constructed using the frequently used log set. It can be seen that in the embodiment of the present invention, by classifying and compressing the backlogged log files, the storage space is saved, and the log graph database is constructed using the multi-dimensional data corresponding to the frequently used log files, so that only the key multi-dimensional data is stored in the log graph database, redundant text data is removed, the storage space is saved, and at the same time, the graph database is a database that is convenient for retrieval.Therefore, the retrieval efficiency of the logs is also improved. The log compressed package is imported into the pre-built basic server to obtain a compressed database, and the log graph database is imported into the pre-built high-speed server to obtain a high-speed database, completing the construction of the log storage system. It can be seen that in the embodiment of the present invention, by importing the log graph database into a high-speed server with a faster reading speed, the retrieval efficiency of the logs is improved, and the log compressed package is stored in the basic server, saving the computing power and storage space of other high-performance servers, and at the same time properly storing the infrequently used log data. Therefore, the present invention can save the storage space of the logs and improve the retrieval efficiency of the logs. Description of the Drawings

[0097] Figure 1 It is a schematic flowchart of a method for constructing a log storage system based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention;

[0098] Figure 2 It is a functional module diagram of a log storage system construction system based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention;

[0099] Figure 3 It is a schematic structural diagram of an electronic device for implementing the method for constructing a log storage system based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention.

[0100] Description of the reference numerals:

[0101] 1. Electronic device; 10. Processor; 11. Storage; 12. Bus.

[0102] The realization, functional characteristics and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments

[0103] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0104] An embodiment of the present application provides a method for constructing a log storage system based on multi-dimensional logical relationship interaction. The execution subject of the method for constructing a log storage system based on multi-dimensional logical relationship interaction includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for constructing a log storage system based on multi-dimensional logical relationship interaction can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0105] Refer to Figure 1As shown in the figure, it is a schematic flowchart of a method for constructing a log storage system based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention. In this embodiment, the method for constructing a log storage system based on multi-dimensional logical relationship interaction includes:

[0106] S1. Receive a log storage instruction, collect an original log set based on the log storage instruction, and perform a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value.

[0107] In the embodiment of the present invention, the constructed log storage system is used to store transaction logs in the financial field. A transaction log refers to a series of logs containing transaction information (such as the user's ID, the amount of the transaction) automatically recorded by the server when a transaction (such as payment, transfer) occurs, and one transaction corresponds to one transaction log.

[0108] It should be explained that the log storage instruction is generally initiated by the system administrator. Exemplarily, Xiao Zhang is a system administrator of a certain financial company. Every once in a while, a large number of transaction logs will accumulate in the company's server. If these transaction logs are not archived to other storage devices in time, it will cause insufficient storage space on the server hard disk, affecting the system operation, and when the transaction logs accumulate too much, the retrieval efficiency of the transaction logs will also decrease. Therefore, in order to store the transaction logs and facilitate future retrieval and access, Xiao Zhang initiates the log storage instruction.

[0109] Specifically, collecting the original log set based on the log storage instruction includes:

[0110] Send the log storage instruction to a pre-constructed log management system, where the log management system includes: a log database;

[0111] After the log management system receives the log storage instruction, extract a general log set from the log database, where the general log set includes multiple general log files;

[0112] Perform the following operations on each general log file in the multiple general log files:

[0113] Identify the log storage capacity of the general log file, and compare the log storage capacity with a preset storage capacity threshold;

[0114] If the log storage capacity is greater than the storage capacity threshold, mark the general log file as a basic log file;

[0115] Summarize the basic log files to obtain multiple basic log files;

[0116] For each of multiple basic log files, perform the following operations:

[0117] Read the log text of the basic log file, and retrieve whether a preset first keyword or a preset second keyword exists in the log text;

[0118] If the first keyword or the second keyword exists in the log text, mark the basic log file corresponding to the log text as the original log file;

[0119] Summarize the original log files to obtain the original log set.

[0120] It should be explained that the log management system is a server used by the finance company to record various logs (such as running logs, error logs, and transaction logs), and the log management system includes a log database. The log database is a database used in the log management system to store logs. The general log set consists of multiple general log files, and the general log file refers to the log file stored in the log database during a certain period in the past.

[0121] It should be understood that the embodiments of the present invention screen out transaction logs from various logs (such as running logs or error logs) by comparing the storage capacity and the retrieval keywords.

[0122] It should be explained that the log storage capacity refers to the size of the general log file. For example, if the size of the general log file is 1 kb, that is, the log storage capacity is 1 kb. The storage capacity threshold is related to the file size of the transaction log and is set by the system administrator.

[0123] It should be understood that the technology of reading the log text of the basic log file is a prior art and will not be elaborated here.

[0124] It can be understood that the log text refers to the string recorded in the basic log file, which contains key event information so that developers, operation and maintenance personnel, or system administrators can understand the behavior of the system. Optionally, select user_id as the first keyword and select amount as the second keyword. Since the difference between transaction logs and other logs is that data such as user IDs and bill information will appear in the text, the transaction log files can be screened out from multiple basic log files by setting the first keyword and the second keyword. The original log file is the screened-out transaction log file.

[0125] It should be understood that performing a parsing operation on the original log set to obtain a multi-dimensional dataset means sequentially performing a parsing operation on the original log files in the original log set to obtain multi-dimensional data and aggregating the multi-dimensional data to obtain a multi-dimensional dataset. Performing a parsing operation on the original log files in the original log set means parsing out the information contained in the original log files from the text of the original log files. For example: The user ID of Xiao Zhao is 12345, and the latitude and longitude coordinates of the location where Xiao Zhao is located are (39.9062, 116.3912), the user ID of Xiao Wang is 67890, and the latitude and longitude coordinates of the location where Xiao Wang is located are (39.5518, 116.2324). At 15:00:01 on November 21, 2024, Xiao Zhao transferred 500 yuan to Xiao Wang, then the text of the original log file will contain the following information: "timestamp": "2024-11-21T15:00:01Z"

[0126] "amount": 500.00

[0127] "sender": {"user_id": "12345", "coordinates": {"latitude": 39.9062, "longitude": 116.3912}}

[0128] "receiver": {"user_id": "67890", "coordinates": {"latitude": 39.5518, "longitude": 116.2324}}

[0129] Among them, timestamp represents the transaction time, amount represents the transaction amount value, sender represents the initiator of the transfer, receiver represents the recipient of the transfer, user_id represents the user ID, coordinates represents the longitude and latitude coordinates, latitude represents the latitude, and longitude represents the longitude. The information in the parentheses after sender represents the information of the initiator. For example, "user_id" in the parentheses after "sender" refers to the initiating ID, and "coordinates" in the parentheses after "sender" refers to the initiating location. Similarly, the receiving ID and receiving coordinates can be obtained. By parsing the text of the original log file, it can be obtained that the transaction time is 15:00:01 on November 21, 2024, the transaction amount value is 500, the initiating user ID is 12345, the initiating location is the location corresponding to the longitude and latitude coordinates (39.9062, 116.3912) in reality, the receiving user ID is 67890, and the receiving location is the location corresponding to the longitude and latitude coordinates (39.5518, 116.2324) in reality. Among them, the latitude comes before the longitude in the longitude and latitude coordinates. And the technology of performing the parsing operation on the original log file in the original log set is an existing technology, which will not be elaborated here.

[0130] S2. Obtain a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and use the a user retrieval records and b time retrieval records to obtain a preferred user sequence and a preferred time sequence respectively.

[0131] Exemplarily, when the staff of a finance company needs to query the transaction records of a certain user through retrieval of the log, they will retrieve the user ID of the user in the log database, and the log management system will record the user ID during the retrieval. Count the a user IDs recorded by the log management system within a period of time as a user retrieval records. When the staff of the finance company needs to query the transaction records near a certain time point, they will retrieve that time point in the log database. Similarly, count the b time points recorded by the log management system during the retrieval within a period of time as b time retrieval records. The retrieval record set consists of a user retrieval records and b time retrieval records and can be directly read from the log management system.

[0132] Generally, in many retrievals within a period of time, it is possible to retrieve the same user ID multiple times, so that the user IDs in the recorded a user IDs may be repeated. Therefore, subsequent uniqueness operations are required to remove the duplicate user IDs.

[0133] Specifically, the obtaining of the preferred user sequence and the preferred time sequence by using the a user retrieval records and b time retrieval records respectively includes:

[0134] Identify a initial user IDs from a user retrieval records, perform a uniquification operation on the a initial user IDs to obtain A unique user IDs, where the user retrieval records correspond to the initial user IDs one by one, and a ≥ A;

[0135] Perform the following operations on each of the A unique user IDs:

[0136] Based on the unique user ID and the a initial user IDs, confirm the user retrieval quantity, where the user retrieval quantity is the number of initial user IDs among the a initial user IDs that are the same as the unique user ID;

[0137] Perform a sorting operation on the A unique user IDs in descending order of the user retrieval quantity corresponding to the unique user ID to obtain a preferred user sequence;

[0138] Identify b retrieval times from b time retrieval records, and based on the b retrieval times, confirm the earliest time and the latest time, where the time retrieval records correspond to the retrieval times one by one, the earliest time is the earliest retrieval time among the b retrieval times, and the latest time is the latest retrieval time among the b retrieval times;

[0139] Based on the earliest time and the latest time, confirm the retrieval time range, where the minimum value of the retrieval time range is the earliest time and the maximum value of the retrieval time range is the latest time;

[0140] Perform a uniform sampling operation on the retrieval time range using a preset time sampling interval to obtain B sampling time ranges;

[0141] Perform the following operations on each of the B sampling time ranges:

[0142] Based on the sampling time range and the b retrieval times, confirm the time retrieval quantity, where the time retrieval quantity is the number of retrieval times among the b retrieval times that are within the sampling time range;

[0143] Perform a sorting operation on the B sampling time ranges in descending order of the time retrieval quantity corresponding to the sampling time range to obtain a preferred time sequence.

[0144] It should be explained that the initial user ID refers to the user ID in the user retrieval record, and the retrieval time refers to the time point in the time retrieval record. The operation of performing uniquification on a initial user IDs to obtain A unique user IDs means: screening out duplicate initial user IDs among the a initial user IDs, so that only one unique initial user ID is retained for multiple identical initial user IDs. For example, if the 6 initial user IDs are: 123, 124, 124, 125, 125, 125, one duplicate user ID, which is 124, and two duplicate user IDs, which are 125, are screened out, obtaining 3 unique user IDs: 123, 124, 125, and the uniquification operation can be implemented using the numpy.unique() function in python.

[0145] It can be understood that the sampling interval is related to the earliest time and the latest time. Optionally, the sampling interval is set to one-thirtieth of the absolute difference between the earliest time and the latest time.

[0146] Exemplarily, the retrieval time range is: (00:00:00 on October 1, 2024, 00:00:00 on October 31, 2024). If the sampling interval is set to 24 hours, then 30 sampling time ranges are obtained: (00:00:00 on October 1, 2024, 00:00:00 on October 2, 2024), …, (00:00:00 on October 30, 2024, 00:00:00 on October 31, 2024).

[0147] S3. Perform the following operations on each multi-dimensional data in the multi-dimensional dataset: calculate the user weight according to the multi-dimensional data and the priority user sequence, and calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence.

[0148] Specifically, the calculation of the user weight according to the multi-dimensional data and the priority user sequence includes:

[0149] Judge whether there is a unique user ID in the priority user sequence that is the same as the initiating user ID in the multi-dimensional data;

[0150] If there is a unique user ID in the priority user sequence that is the same as the initiating user ID, then use the unique user ID in the priority user sequence that is the same as the initiating user ID as the target user ID;

[0151] If there is no unique user ID in the priority user sequence that is the same as the initiating user ID, then judge whether there is a unique user ID in the priority user sequence that is the same as the receiving user ID in the multi-dimensional data;

[0152] If there is a unique user ID in the priority user sequence that is the same as the received user ID, then use the unique user ID in the priority user sequence that is the same as the received user ID as the target user ID;

[0153] Retrieve the user level number from the pre-built user database based on the target user ID;

[0154] Confirm the user ranking value in the priority user sequence based on the target user ID, where the user ranking value is the position of the target user ID in the priority user sequence;

[0155] Calculate the user weight according to the user ranking value and the user level number, and the calculation formula is as follows:

[0156]

[0157] Where, is the user weight, is the user ranking value, is the user level number, is the preset maximum user level number, is the hyperbolic tangent function;

[0158] If there is no unique user ID in the priority user sequence that is the same as the received user ID, then record the user weight as 0.

[0159] It should be explained that the user database is composed of user information stored in the server in advance, and the user information includes the user level. The user level is the level evaluated by the financial company for each user based on the user's account transactions and credit, etc. The specific rating criteria are related to the financial company.

[0160] Exemplarily, Xiao Zhao is a user of the financial company, with a user level of 8, that is, the user level number is 8. His user ID is 1234567, and it is found that the position of the user ID 1234567 in the priority user sequence is the 10th, that is, the user ranking value is 10. If there are 100 unique user IDs in the priority user sequence, that is, A is 100, and if the highest user level is 10 when the financial company conducts ratings, that is, the maximum user level number is 10, then the calculated user weight is 0.5976.

[0161] It should be understood that the more frequently the user ID corresponding to an original log file is retrieved, and the higher the level of the user corresponding to the original log file, the more important the original log file represents. Therefore, the user weight reflects the importance of the original log file. The higher the user weight, the higher the importance of the original log file, and the higher the priority in subsequent storage.

[0162] Specifically, calculating the time weight according to the transaction time and the priority time series in the multi-dimensional data includes:

[0163] Determine whether the transaction time is within the retrieval time range;

[0164] If the transaction time is within the retrieval time range, the following operations are sequentially performed for each of the B sampling time ranges:

[0165] Determine whether the transaction time is within the sampling time range. If the transaction time is within the sampling time range, then use the sampling time range as the target time range;

[0166] Based on the target time range, confirm the time ranking value in the priority time series, where the time ranking value is the position of the target time range in the priority time series;

[0167] Obtain the current time, and calculate the time difference according to the current time and the transaction time, where the time difference is the absolute difference between the current time and the transaction time;

[0168] Calculate the time weight according to the time difference and the time ranking value. The calculation formula is as follows:

[0169]

[0170] Wherein, is the time weight, is the time ranking value, is the time difference, is the natural constant;

[0171] If the transaction time is not within the retrieval time range, then record the time weight as 0.

[0172] It should be understood that the method of confirming the time ranking value based on the target time range in the priority time series is the same as the method of confirming the user ranking value based on the target user ID in the priority user series, and will not be elaborated here.

[0173] It should be explained that the current time refers to the time when the time ranking value is confirmed. Optionally, when the time ranking value is confirmed, query the Beijing time at this time as the current time.

[0174] It should be understood that the smaller the time difference, the later the generation time of the original log file, the newer the original log file, and the greater the probability of being used for retrieval, indicating that the importance of the original log file is higher. Therefore, the time weight reflects the importance of the original log file. The higher the time weight, the higher the importance of the original log file and the higher the priority in subsequent storage.

[0175] S4. Obtain the headquarters location, and calculate the location weight according to the initiation location and the headquarters location in the multi-dimensional data.

[0176] It should be explained that the headquarters location refers to the longitude and latitude coordinates of the finance company in the real world. Optionally, the headquarters location is obtained through the Global Positioning System.

[0177] Specifically, calculating the location weight according to the initiation location and the headquarters location in the multi-dimensional data includes:

[0178] Obtain the transaction distance using the initiation location and the headquarters location, and calculate the location weight using the transaction distance. The calculation formula is as follows:

[0179]

[0180] Where, is the location weight, is the transaction distance, is the preset standard distance.

[0181] It should be explained that the transaction distance refers to the distance between the initiation location and the headquarters location in the real world. The transaction distance can be calculated through the longitude and latitude coordinates of the initiation location and the longitude and latitude coordinates of the headquarters location. And the technology of calculating the transaction distance through the longitude and latitude of the initiation location and the longitude and latitude of the headquarters location is an existing technology, which will not be elaborated here.

[0182] Exemplarily, for the convenience of communication between the finance company and users, the finance company only accepts users within a distance of 50 km from the finance company. Therefore, the standard distance is set to 50 km.

[0183] It should be understood that the smaller the transaction distance, the closer the user is to the finance company during the transaction, and the greater the probability that the user will come to the finance company for consultation. That is, the greater the probability that the original log file corresponding to the transaction distance is used for retrieval, indicating that the importance of the original log file is higher. Therefore, the location weight reflects the importance of the original log file. The higher the location weight, the higher the importance of the original log file, and the higher the priority during subsequent storage.

[0184] S5. Calculate the log priority according to the user weight, time weight, location weight and transaction amount value.

[0185] Specifically, the calculation formula of the log priority is as follows:

[0186]

[0187] Where, is the log priority, is the transaction amount value, is the natural logarithm.

[0188] It should be understood that the log priority reflects the importance of the original log file. The higher the log priority, the higher the importance of the original log file.

[0189] S6. Compare the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a frequently used log file. If the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlogged log file.

[0190] Optionally, the system administrator selects multiple original log files that have been retrieved, calculates the log priority of each original log file in the multiple original log files in turn, and finally uses the minimum log priority among the calculated multiple log priorities as the priority threshold.

[0191] It should be understood that in the embodiments of the present invention, the original log files are divided into frequently used log files and backlogged log files through the priority threshold, so as to store the frequently used log files and backlogged log files separately in the subsequent process. Among them, the frequently used log files are original log files with a very high degree of importance and are estimated to be frequently retrieved. Therefore, in the subsequent process, a graph database storage method that is convenient for retrieval is adopted, while the backlogged log files are original log files with a relatively low degree of importance and are estimated to be occasionally retrieved. Therefore, in the subsequent process, they are stored in a classified compression manner.

[0192] S7. Aggregate the frequently used log files and backlogged log files respectively to obtain a frequently used log set and a backlogged log set. Perform a classified compression operation on the backlogged log set to obtain a log compression package, and construct a log graph database using the frequently used log set.

[0193] Specifically, the performing a classified compression operation on the backlogged log set to obtain a log compression package includes:

[0194] Perform a merge operation on multiple backlogged log files in the backlogged log set to obtain multiple log folders. Each log folder in the multiple log folders includes: one or more backlogged log files;

[0195] Obtain multiple file storage amounts using the multiple log folders, where the file storage amounts correspond to the log folders one by one;

[0196] Based on the multiple file storage amounts, confirm the maximum storage amount and the minimum storage amount. The maximum storage amount is the file storage amount with the largest storage capacity among the multiple file storage amounts, and the minimum storage amount is the file storage amount with the smallest storage capacity among the multiple file storage amounts;

[0197] Perform the following operations on each log folder in the multiple log folders:

[0198] Calculate the compression level using the maximum storage, minimum storage, and the file storage corresponding to the log folder. The calculation formula is as follows:

[0199]

[0200] Among them, is the compression level, is the maximum storage, is the minimum storage, is the file storage corresponding to the log folder, is the preset maximum compression level, represents the ceiling operation;

[0201] Perform a compression operation on the log folder based on the compression level to obtain a compressed file;

[0202] Summarize the compressed files to obtain a log compression package.

[0203] It should be explained that performing a merge operation on multiple backlog log files in the backlog log set means: integrating the backlog log files with the same initiating user ID among multiple backlog log files into one folder, and the merged folder is the log folder. For example: there are 5 backlog log files. The initiating user ID corresponding to the first backlog log file is 123, the initiating user ID corresponding to the second backlog log file is 124, the initiating user ID corresponding to the third backlog log file is 125, the initiating user ID corresponding to the fourth backlog log file is 123, and the initiating user ID corresponding to the fifth backlog log file is 124. Then, integrate the first backlog log file and the fourth backlog log file into one folder, integrate the second backlog log file and the fifth backlog log file into one folder, and integrate the third backlog log file into one folder. Finally, 3 log folders are obtained.

[0204] It can be understood that the file storage refers to the size of the log folder. For example, if the size of the log folder is 100kb, then the file storage is 100kb. The compression level reflects the compression ratio when compressing the log folder. The higher the compression level, the greater the compression ratio, and different compression software can select different compression levels. For example: 7-Zip can select compression levels from 0 to 9, Zstandard can select compression levels from 1 to 22, and the maximum compression level is related to the compression software used. If the compression software used is 7-Zip, then the maximum compression level is 9.

[0205] It should be understood that the larger the file storage capacity of the log folder corresponding to a compressed file, the more backlogged log files there are in the log folder. When the financial company retrieves the log, the probability of retrieving the compressed file is greater. When the compressed file is retrieved, it is necessary to decompress the compressed file to extract the backlogged log files required by the financial company. Also, because the higher the compression level, the longer the time consumed for decompressing the compressed file. Therefore, when compressing the log folder with a greater retrieval probability, that is, the log folder with a larger file storage capacity, a lower compression level will be adopted to reduce the time consumed for decompression.

[0206] Exemplarily, if the compression level is 8, then when performing the compression operation on the log folder, set the compression level of 7-Zip to level 8.

[0207] Optionally, use the following command to set the compression level of 7-Zip to level 8 when performing the compression operation on the log folder: 7z a -t7z -mx=8 example.7z example.txt. The technology for performing the compression operation on the log folder is an existing technology and will not be elaborated here.

[0208] Specifically, constructing the log graph database using the common log set includes:

[0209] Perform the following operations on each common log file in the common log set:

[0210] Combine the initiating user ID and initiating location corresponding to the common log file into an initiating information group, and combine the receiving user ID and receiving location corresponding to the common log file into a receiving information group;

[0211] Mark both the initiating information group and the receiving information group as user information groups;

[0212] Summarize the user information groups to obtain multiple user information groups;

[0213] Perform a uniquification operation on the multiple user information groups to obtain multiple unique user information, and use the multiple unique user information to create multiple unique user nodes in the pre-constructed graph database construction software, where the unique user nodes and the unique user information correspond one by one;

[0214] Perform the following operations on each common log file in the common log set:

[0215] Use the initiating user ID corresponding to the common log file as the indexing initiating ID, and use the indexing initiating ID to retrieve the initiating node among the multiple unique user nodes, where the initiating user ID or receiving user ID in the unique user information corresponding to the initiating node is the same as the indexing initiating ID;

[0216] Obtain the receiving node based on the receiving user ID corresponding to the common log file;

[0217] Create a relationship edge in the graph database construction software based on the transaction time, transaction amount value, initiating node, and receiving node corresponding to the common log file;

[0218] Summarize the relationship edges to obtain multiple relationship edges, and confirm the log graph database based on the multiple unique user nodes and multiple relationship edges.

[0219] It should be understood that the method of performing the uniqueness operation on multiple user information groups to obtain multiple unique user information is the same as the method of performing the uniqueness operation on a initial user IDs to obtain A unique user IDs, which will not be elaborated here.

[0220] In the embodiment of the present invention, the graph database construction software used is Neo4j graph database. The unique user node is a basic unit in the Neo4j graph database for storing unique user information. The relationship edge is a structure located between two unique user nodes in the Neo4j graph database, reflecting the relationship between the two unique user nodes.

[0221] Exemplarily, if in a unique user information, the initiating user ID or the receiving user ID is 12345, and the initiating location or the receiving location is the longitude and latitude coordinates (39.9062, 116.3912), then write the following code in the command line of Neo4j Browser and run it to create a unique user node:

[0222] CREATE (U1:User {

[0223] id: "12345",

[0224] latitude: 39.9062,

[0225] longitude: 116.3912

[0226] }) where CREATE means create, User is the label of the unique user node, used to indicate that the unique user node is of the "user" type. The label can help the system administrator classify the unique user nodes or optimize the query process. id represents the initiating user ID or the receiving user ID, latitude represents the latitude in the longitude and latitude coordinates, and longitude represents the longitude in the longitude and latitude coordinates. U1 is the reference name of the node. The reference name is assigned by the system administrator when creating the unique user node, and each unique user node corresponds to a reference name. When creating the relationship edge later, the code will be written using the reference name corresponding to the initiating node and the reference name corresponding to the receiving node to create the relationship edge.

[0227] It should be understood that the method for obtaining the receiving node based on the received user ID corresponding to the common log file is the same as the method for obtaining the originating node using the originating user ID corresponding to the common file, which will not be elaborated here.

[0228] Exemplarily, if the transaction time is 10:00:00 on November 21, 2024, the transaction amount value is 200, the reference name corresponding to the originating node is U1, and the reference name corresponding to the receiving node is U2, then the following code is written and run in the command line of Neo4j Browser to create a relationship edge:

[0229] CREATE (U1)-[:TRANSACTION {

[0230] amount: 200,

[0231] timestamp: datetime("2024-11-21T10:00:00")

[0232] }]->(U2), where TRANSACTION is used to indicate that the type of the relationship edge is a "transaction" relationship, amount represents the transaction amount value, timestamp represents the transaction time, and datetime() is a function used to define the data within the parentheses as the date-time type. It can be understood that the confirmation of the log graph database based on multiple unique user nodes and multiple relationship edges means that when multiple unique user nodes and multiple relationship edges are successfully created in the graph database construction software, the log graph database can be confirmed to be successfully constructed.

[0233] S8. Import the log compression package into the pre-built basic server to obtain a compressed database, and import the log graph database into the pre-built high-speed server to obtain a high-speed database, thereby completing the construction of the log storage system.

[0234] It should be understood that the basic server uses a mechanical hard disk and is the server for storing log compression packages. Since the log compression packages store infrequently used and unimportant original log files and do not require a very high retrieval or reading speed, a mechanical hard disk is used in the basic server to prioritize storage capacity and economically store more log compression packages. The high-speed server uses a solid-state drive and is the server for storing the log graph database. Since the log graph database corresponds to frequently used and important original log files and needs to be frequently retrieved and read, a solid-state drive is used in the high-speed server to prioritize the reading speed, thereby efficiently realizing the retrieval of logs. The compression database is the database in the basic server that stores log compression packages. The high-speed database is the log graph database stored in the high-speed server. And the technology of importing the log graph database into the pre-built high-speed server is an existing technology and will not be elaborated here.

[0235] To solve the problems described in the background art, the present invention receives a log storage instruction, collects an original log set based on the log storage instruction, and performs a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value. It can be seen that in the embodiment of the present invention, by parsing the original log files into multiple multi-dimensional data, redundant text data is eliminated, the storage space of the logs is saved, and at the same time, it is convenient to classify the original log files according to the multi-dimensional data subsequently, and then obtain a retrieval record set. Among them, the retrieval record set includes: a user retrieval records and b time retrieval records. The priority user sequence and the priority time sequence are obtained by using the a user retrieval records and the b time retrieval records respectively. It can be seen that in the embodiment of the present invention, by analyzing the historical retrieval records, reference data is provided for calculating the log priority subsequently, and the accuracy of classifying the original log files subsequently is improved. The following operations are performed on each multi-dimensional data in the multi-dimensional data set: calculate the user weight according to the multi-dimensional data and the priority user sequence, calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence, obtain the headquarters location, and calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location. It can be seen that in the embodiment of the present invention, by comprehensively analyzing the data in the three dimensions of the user, the transaction time, and the initiating location, the user weight, the time weight, and the location weight are confirmed, which is convenient for subsequently calculating the log priority comprehensively and accurately according to the user weight, the time weight, and the location weight. Calculate the log priority according to the user weight, the time weight, the location weight, and the transaction amount value, compare the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, the original log file corresponding to the log priority is marked as a frequently used log file. If the log priority is less than the priority threshold, the original log file corresponding to the log priority is marked as a backlogged log file. It can be seen that in the embodiment of the present invention, the original log files are divided by the log priority and the priority threshold, the files that are frequently retrieved and important are marked as frequently used log files, and the files that are occasionally retrieved and unimportant are marked as backlogged log files, which is convenient for subsequently adopting targeted storage methods for the two types of files to save the storage space of the logs and improve the retrieval efficiency of the logs. The frequently used log files and the backlogged log files are summarized respectively to obtain a frequently used log set and a backlogged log set. A classification compression operation is performed on the backlogged log set to obtain a log compression package, and a log graph database is constructed using the frequently used log set. It can be seen that in the embodiment of the present invention, by classifying and compressing the backlogged log files, the storage space is saved, and the log graph database is constructed using the multi-dimensional data corresponding to the frequently used log files, so that only the key multi-dimensional data is stored in the log graph database, redundant text data is removed, the storage space is saved, and at the same time, the graph database is a database that is convenient for retrieval.Therefore, the retrieval efficiency of the logs is also improved. The log compression package is imported into the pre-built basic server to obtain a compressed database, and the log graph database is imported into the pre-built high-speed server to obtain a high-speed database, completing the construction of the log storage system. It can be seen that in the embodiment of the present invention, by importing the log graph database into a high-speed server with a faster reading speed, the retrieval efficiency of the logs is improved, and the log compression package is stored in the basic server, saving the computing power and storage space of other high-performance servers, and at the same time properly storing the infrequently used log data. Therefore, the present invention can save the storage space of the logs and improve the retrieval efficiency of the logs.

[0236] As Figure 2 shown, it is a functional module diagram of a log storage system construction system based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention.

[0237] The log storage system construction system 100 based on multi-dimensional logical relationship interaction described in the present invention can be installed in an electronic device. According to the implemented functions, the log storage system construction system 100 based on multi-dimensional logical relationship interaction can include a multi-dimensional data parsing module 101, a retrieval record analysis module 102, a priority log classification module 103, and a log database construction module 104. The modules described in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0238] The multi-dimensional data parsing module 101 is configured to receive a log storage instruction, collect an original log set based on the log storage instruction, and perform a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: initiating user ID, initiating location, receiving user ID, receiving location, transaction time, and transaction amount value;

[0239] The retrieval record analysis module 102 is configured to obtain a retrieval record set. Among them, the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively use the a user retrieval records and the b time retrieval records to obtain a priority user sequence and a priority time sequence;

[0240] The priority log classification module 103 is used to perform the following operations on each multi-dimensional data in the multi-dimensional dataset: calculate the user weight according to the multi-dimensional data and the priority user sequence, calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence, obtain the headquarters location, calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location, calculate the log priority according to the user weight, time weight, location weight and transaction amount value, compare the log priority with a preset priority threshold, if the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a common log file, and if the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlog log file;

[0241] The log database construction module 104 is used to respectively summarize the common log files and the backlog log files to obtain a common log set and a backlog log set, perform a classification and compression operation on the backlog log set to obtain a log compression package, construct a log graph database using the common log set, import the log compression package into a pre-constructed basic server to obtain a compressed database, and import the log graph database into a pre-constructed high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

[0242] Specifically, each module in the log storage system construction system 100 based on multi-dimensional logical relationship interaction in the embodiments of the present invention adopts the same technical means as those in the Figure 1 log storage system construction method based on multi-dimensional logical relationship interaction described above, and can produce the same technical effects, which will not be elaborated here.

[0243] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the log storage system construction method based on multi-dimensional logical relationship interaction provided by an embodiment of the present invention.

[0244] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may further include a computer program stored in the memory 11 and operable on the processor 10, such as a log storage system construction method program based on multi-dimensional logical relationship interaction.

[0245] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 11 can also be an external storage device of the electronic device 1 in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 also includes the internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code of the method program for constructing a log storage system based on multi-dimensional logical relationship interaction, etc., but also to temporarily store data that has been output or will be output.

[0246] The processor 10 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting all components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the method program for constructing a log storage system based on multi-dimensional logical relationship interaction, etc.), and calling data stored in the memory 11, to execute various functions of the electronic device 1 and process data.

[0247] The bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0248] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3The structures shown do not constitute a limitation on the electronic device 1, and it may include fewer or more components than those shown, or combine certain components, or have different component arrangements.

[0249] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0250] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0251] Optionally, the electronic device 1 may also include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0252] The method program for constructing a log storage system based on multi-dimensional logical relationship interaction stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0253] Receiving a log storage instruction, collecting an original log set based on the log storage instruction, and performing a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: initiating user ID, initiating location, receiving user ID, receiving location, transaction time, and transaction amount value;

[0254] Obtain a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively obtain a priority user sequence and a priority time sequence by using the a user retrieval records and the b time retrieval records;

[0255] Perform the following operations on each multi-dimensional data in the multi-dimensional data set:

[0256] Calculate the user weight according to the multi-dimensional data and the priority user sequence, and calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence;

[0257] Obtain the headquarters location, and calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location;

[0258] Calculate the log priority according to the user weight, time weight, location weight and transaction amount value;

[0259] Compare the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a common log file. If the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlog log file;

[0260] Aggregate the common log files and the backlog log files respectively to obtain a common log set and a backlog log set;

[0261] Perform a classification compression operation on the backlog log set to obtain a log compression package, and construct a log graph database by using the common log set;

[0262] Import the log compression package into a pre-constructed basic server to obtain a compressed database, and import the log graph database into a pre-constructed high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

[0263] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figures 1 to 3 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0264] Further, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0265] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor of an electronic device, can implement:

[0266] Receiving a log storage instruction, collecting an original log set based on the log storage instruction, and performing a parsing operation on the original log set to obtain a multi-dimensional data set, where the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. The multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value;

[0267] Obtaining a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively using the a user retrieval records and the b time retrieval records to obtain a priority user sequence and a priority time sequence;

[0268] Performing the following operations on each multi-dimensional data in the multi-dimensional data set:

[0269] Calculating a user weight according to the multi-dimensional data and the priority user sequence, and calculating a time weight according to the transaction time in the multi-dimensional data and the priority time sequence;

[0270] Obtaining the headquarters location, and calculating a location weight according to the initiating location in the multi-dimensional data and the headquarters location;

[0271] Calculating a log priority according to the user weight, the time weight, the location weight, and the transaction amount value;

[0272] Comparing the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, marking the original log file corresponding to the log priority as a common log file. If the log priority is less than the priority threshold, marking the original log file corresponding to the log priority as a backlogged log file;

[0273] Respectively summarizing the common log files and the backlogged log files to obtain a common log set and a backlogged log set;

[0274] Performing a classification and compression operation on the backlogged log set to obtain a log compression package, and constructing a log graph database using the common log set;

[0275] Importing the log compression package into a pre-constructed basic server to obtain a compressed database, and importing the log graph database into a pre-constructed high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

[0276] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and there may be other partitioning methods in actual implementation.

[0277] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0278] In addition, the functional modules in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0279] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0280] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a log storage system based on multi-dimensional logical relationship interaction, characterized in that, The method includes: Receiving a log storage instruction, collecting an original log set based on the log storage instruction, and performing a parsing operation on the original log set to obtain a multi-dimensional data set. The original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. The multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value; Obtaining a retrieval record set, where the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively using the a user retrieval records and the b time retrieval records to obtain a priority user sequence and a priority time sequence; Performing the following operations on each multi-dimensional data in the multi-dimensional data set: Calculating a user weight according to the multi-dimensional data and the priority user sequence, and calculating a time weight according to the transaction time in the multi-dimensional data and the priority time sequence; Obtaining the headquarters location, and calculating a location weight according to the initiating location in the multi-dimensional data and the headquarters location; Calculating a log priority according to the user weight, the time weight, the location weight, and the transaction amount value; Comparing the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, identifying the original log file corresponding to the log priority as a commonly used log file. If the log priority is less than the priority threshold, identifying the original log file corresponding to the log priority as a backlogged log file; Respectively summarizing the commonly used log files and the backlogged log files to obtain a commonly used log set and a backlogged log set; Performing a classification compression operation on the backlogged log set to obtain a log compression package, and constructing a log graph database using the commonly used log set. The constructing the log graph database using the commonly used log set includes: Performing the following operations on each commonly used log file in the commonly used log set: Combining the initiating user ID and the initiating location corresponding to the commonly used log file into an initiating information group, and combining the receiving user ID and the receiving location corresponding to the commonly used log file into a receiving information group; Identifying both the initiating information group and the receiving information group as user information groups; Summarizing the user information groups to obtain multiple user information groups; Performing a uniqueness operation on the multiple user information groups to obtain multiple unique user information, and creating multiple unique user nodes in a pre-constructed graph database construction software using the multiple unique user information. The unique user nodes and the unique user information are in one-to-one correspondence; Performing the following operations on each commonly used log file in the commonly used log set: Using the initiating user ID corresponding to the commonly used log file as a search initiating ID, and retrieving an initiating node from the multiple unique user nodes using the search initiating ID. The initiating user ID or the receiving user ID in the unique user information corresponding to the initiating node is the same as the search initiating ID; Obtaining a receiving node based on the receiving user ID corresponding to the commonly used log file; Creating a relationship edge in the graph database construction software based on the transaction time, the transaction amount value, the initiating node, and the receiving node corresponding to the commonly used log file; Summarizing the relationship edges to obtain multiple relationship edges, and confirming the log graph database based on the multiple unique user nodes and the multiple relationship edges; Import the log compression package into a pre-built basic server to obtain a compressed database, and import the log graph database into a pre-built high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

2. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 1, wherein The collection of the original log set based on the log storage instruction includes: Send the log storage instruction to a pre-built log management system, where the log management system includes: a log database; After the log management system receives the log storage instruction, extract the general log set from the log database, where the general log set includes multiple general log files; Perform the following operations on each of the multiple general log files: Identify the log storage capacity of the general log file and compare the log storage capacity with a preset storage capacity threshold; If the log storage capacity is greater than the storage capacity threshold, mark the general log file as a basic log file; Summarize the basic log files to obtain multiple basic log files; Perform the following operations on each of the multiple basic log files: Read the log text of the basic log file and retrieve whether there is a preset first keyword or a preset second keyword in the log text; If there is a first keyword or a second keyword in the log text, mark the basic log file corresponding to the log text as an original log file; Summarize the original log files to obtain the original log set.

3. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 2, wherein The obtaining of the priority user sequence and the priority time sequence by using a user retrieval records and b time retrieval records respectively includes: Identify a initial user IDs in a user retrieval records, perform a uniqueness operation on the a initial user IDs to obtain A unique user IDs, where the user retrieval records and the initial user IDs are in one-to-one correspondence, and a ≥ A; Perform the following operations on each of the A unique user IDs: Based on the unique user ID and the a initial user IDs, confirm the user retrieval quantity, where the user retrieval quantity is the number of initial user IDs that are the same as the unique user ID among the a initial user IDs; Perform a sorting operation on the A unique user IDs in the order of the user retrieval quantity corresponding to the unique user ID from more to less to obtain the priority user sequence; Identify b retrieval times in the b time retrieval records, and based on the b retrieval times, confirm the earliest time and the latest time, where the time retrieval records and the retrieval times are in one-to-one correspondence, the earliest time is the earliest retrieval time among the b retrieval times, and the latest time is the latest retrieval time among the b retrieval times; Based on the earliest time and the latest time, confirm the retrieval time range, where the minimum value of the retrieval time range is the earliest time and the maximum value of the retrieval time range is the latest time; Perform a uniform sampling operation on the retrieval time range by using a preset time sampling interval to obtain B sampling time ranges; Perform the following operations on each of the B sampling time ranges: Based on the sampling time range and the b retrieval times, confirm the time retrieval quantity, where the time retrieval quantity is the number of retrieval times among the b retrieval times that are within the sampling time range; Perform a sorting operation on B sampling time ranges in the order of the number of retrieved data corresponding to the sampling time range from more to less to obtain a preferred time sequence.

4. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 3, wherein The calculating of the user weight according to the multi-dimensional data and the preferred user sequence includes: Determine whether there is a unique user ID in the preferred user sequence that is the same as the initiating user ID in the multi-dimensional data; If there is a unique user ID in the preferred user sequence that is the same as the initiating user ID, use the unique user ID in the preferred user sequence that is the same as the initiating user ID as the target user ID; If there is no unique user ID in the preferred user sequence that is the same as the initiating user ID, determine whether there is a unique user ID in the preferred user sequence that is the same as the receiving user ID in the multi-dimensional data; If there is a unique user ID in the preferred user sequence that is the same as the receiving user ID, use the unique user ID in the preferred user sequence that is the same as the receiving user ID as the target user ID; Retrieve the user level number from the pre-constructed user database based on the target user ID; Confirm the user ranking value in the preferred user sequence based on the target user ID, where the user ranking value is the position of the target user ID in the preferred user sequence; Calculate the user weight according to the user ranking value and the user level number, and the calculation formula is as follows: Among them, is the user weight, is the user ranking value, is the user level number, is the preset maximum user level number, is the hyperbolic tangent function; If there is no unique user ID in the preferred user sequence that is the same as the receiving user ID, record the user weight as 0.

5. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 4, wherein The calculating of the time weight according to the transaction time in the multi-dimensional data and the preferred time sequence includes: Determine whether the transaction time is within the retrieved time range; If the transaction time is within the retrieved time range, perform the following operations on each of the B sampling time ranges in turn: Determine whether the transaction time is within the sampling time range. If the transaction time is within the sampling time range, use the sampling time range as the target time range; Confirm the time ranking value in the preferred time sequence based on the target time range, where the time ranking value is the position of the target time range in the preferred time sequence; Obtain the current time, and calculate the time difference according to the current time and the transaction time, where the time difference is the absolute difference between the current time and the transaction time; Calculate the time weight according to the time difference and the time ranking value, and the calculation formula is as follows: Among them, is the time weight, is the time ranking value, is the time difference, is the natural constant; If the transaction time is not within the retrieved time range, record the time weight as 0.

6. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 5, wherein The calculating of the location weight according to the initiating location and the headquarters location in the multi-dimensional data includes: Obtain the transaction distance using the initiating location and the headquarters location, and calculate the location weight using the transaction distance. The calculation formula is as follows: Among them, is the position weight, is the transaction distance, is the preset standard distance.

7. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 6, characterized in that, The calculation formula of the log priority is as follows: Among them, is the log priority, is the transaction amount value, is the natural logarithm.

8. The method for constructing a log storage system based on multi-dimensional logical relationship interaction according to claim 7, wherein The performing of the classification and compression operation on the backlog log set to obtain a log compression package includes: Perform a merging operation on multiple backlog log files in the backlog log set to obtain multiple log folders, where each of the multiple log folders includes: one or more backlog log files; Obtain multiple file storage amounts using the multiple log folders, where the file storage amounts and the log folders correspond one by one; The maximum storage capacity and the minimum storage capacity are confirmed based on multiple file storage capacities. Among them, the maximum storage capacity is the file storage capacity with the largest storage capacity among the multiple file storage capacities, and the minimum storage capacity is the file storage capacity with the smallest storage capacity among the multiple file storage capacities; Perform the following operations on each log folder among multiple log folders: Calculate the compression level by using the maximum storage capacity, the minimum storage capacity, and the file storage capacity corresponding to the log folder. The calculation formula is as follows: Among them, is the compression level, is the maximum storage capacity, is the minimum storage capacity, is the file storage capacity corresponding to the log folder, is the preset maximum compression level, represents the ceiling operation; Perform a compression operation on the log folder based on the compression level to obtain a compressed file; Aggregate the compressed files to obtain a log compression package.

9. A system for constructing a log storage system based on multi-dimensional logical relationship interaction, characterized in that The system includes: A multi-dimensional data parsing module, which is used to receive a log storage instruction, collect an original log set based on the log storage instruction, and perform a parsing operation on the original log set to obtain a multi-dimensional data set. Among them, the original log set includes multiple original log files, the multi-dimensional data set includes multiple multi-dimensional data, and the original log files and the multi-dimensional data are in one-to-one correspondence. Among them, the multi-dimensional data includes: the initiating user ID, the initiating location, the receiving user ID, the receiving location, the transaction time, and the transaction amount value; A retrieval record analysis module, which is used to obtain a retrieval record set. Among them, the retrieval record set includes: a user retrieval records and b time retrieval records, and respectively use the a user retrieval records and the b time retrieval records to obtain a priority user sequence and a priority time sequence; A priority log classification module, which is used to perform the following operations on each multi-dimensional data in the multi-dimensional data set: calculate the user weight according to the multi-dimensional data and the priority user sequence, calculate the time weight according to the transaction time in the multi-dimensional data and the priority time sequence, obtain the headquarters location, calculate the location weight according to the initiating location in the multi-dimensional data and the headquarters location, calculate the log priority according to the user weight, the time weight, the location weight, and the transaction amount value, compare the log priority with a preset priority threshold. If the log priority is greater than or equal to the priority threshold, mark the original log file corresponding to the log priority as a common log file. If the log priority is less than the priority threshold, mark the original log file corresponding to the log priority as a backlogged log file; A log database construction module, which is used to aggregate the common log files and the backlogged log files respectively to obtain a common log set and a backlogged log set, perform a classification compression operation on the backlogged log set to obtain a log compression package, and construct a log graph database by using the common log set. Among them, the construction of the log graph database by using the common log set includes: Perform the following operations on each common log file in the common log set: Combine the initiating user ID and the initiating location corresponding to the common log file into an initiating information group, and combine the receiving user ID and the receiving location corresponding to the common log file into a receiving information group; Mark both the initiating information group and the receiving information group as user information groups; Aggregate the user information groups to obtain multiple user information groups; Perform a uniquification operation on the multiple user information groups to obtain multiple unique user information, and create multiple unique user nodes in a pre-constructed graph database construction software by using the multiple unique user information. Among them, the unique user nodes and the unique user information are in one-to-one correspondence; Perform the following operations on each common log file in the common log set: Use the initiating user ID corresponding to the common log file as the indexing initiating ID, and use the indexing initiating ID to retrieve the initiating node among multiple unique user nodes. Among them, the initiating user ID or the receiving user ID in the unique user information corresponding to the initiating node is the same as the indexing initiating ID; Obtain the receiving node based on the receiving user ID corresponding to the common log file; Create relationship edges in the graph database construction software based on the transaction time, transaction amount value, initiating node, and receiving node corresponding to the common log file; Summarize the relationship edges to obtain multiple relationship edges, and confirm the log graph database based on the multiple unique user nodes and the multiple relationship edges; Import the log compression package into the pre-built basic server to obtain a compressed database, and import the log graph database into the pre-built high-speed server to obtain a high-speed database, thus completing the construction of the log storage system.

Citation Information

Patent Citations

  • Log storage method, device and system, electronic equipment and computer readable medium

    CN111046010A

  • Real-time storage method for log data

    CN118964916A