Train data classified storage method
By constructing a set of correlation patterns and optimizing the preloading strategy in the train monitoring system, the problem of low retrieval efficiency after massive data storage is solved, realizing efficient and fast data access and query, and adapting to the data storage needs of multiple sources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2026-04-21
AI Technical Summary
The problem of low efficiency in retrieving massive amounts of data after data storage in train monitoring systems, especially when data is stored and retrieved from multiple sources, is that existing technologies cannot handle it efficiently.
By identifying data sources, setting category labels, constructing a set of association patterns, analyzing call pattern characteristics, classifying association interference categories, prioritizing the loading of high-frequency associated data, and dynamically adjusting the preloading order, the efficiency of retrieving stored data can be improved.
In the train monitoring system, data access efficiency has been improved, query latency and resource waste have been reduced, high-frequency and low-frequency correlated data have been processed adaptively, repeated query storms have been avoided, and real-time requirements have been met.
Smart Images

Figure CN120849681B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage, and more particularly to a method for classifying and storing train data. Background Technology
[0002] Currently, due to the combined effects of various technological factors and operational needs, the number of sensors and sampling frequency of train systems have increased explosively. The operation of a single train generates a large amount of data every day, such as sensor data, video surveillance, and operation logs. At the same time, the amount of data continues to accumulate as the train's operating time and mileage increase, forming a "data tsunami." This increases the latency when querying and retrieving targeted data. Therefore, it is necessary to classify and store train data to meet the real-time query and retrieval needs of massive amounts of train data.
[0003] Chinese Patent Application Publication No. CN119027236A discloses a method for classifying and storing information technology data, specifically relating to the field of online classification and storage technology. The method includes the following steps: representing historical financial data in the form of a knowledge graph to obtain the correlation between real-time financial data and historical financial data, and obtaining the timeliness of historical financial data based on the timestamps of the historical financial data; comprehensively analyzing the correlation between real-time financial data and historical financial data and the timeliness of historical financial data to obtain reference data that can be used as current financial institutions; merging and analyzing the obtained reference data with real-time financial data to obtain information on the fluctuation of financial amounts, transaction frequency, and speed of single transactions, and determining the financial data that needs to be stored separately in the merged financial data. This helps financial institutions to filter historical financial data from other financial institutions, retain meaningful financial information of customers, and more accurately locate financial information with potential risks.
[0004] However, the following problems still exist in the existing technology.
[0005] The train monitoring system needs to acquire data from several sensors deployed on the train. Due to the large number of trains and the continuous monitoring by some sensors, the acquired data is massive. This massive amount of data needs to be stored for subsequent retrieval and analysis. However, when the amount of data is large, the centralized storage method results in low retrieval efficiency during subsequent retrieval. Summary of the Invention
[0006] To address this issue, the present invention provides a method for classifying and storing train data, thereby overcoming the problem of low retrieval efficiency in subsequent calls to massive amounts of data from multiple sources in the prior art.
[0007] To achieve the above objectives, the present invention provides a method for classifying and storing train data, comprising:
[0008] Determine the data source corresponding to the data to be stored, set category labels for the data to be stored based on the data source, and store the data to be stored in the database;
[0009] Active analysis is performed based on the call records of the stored data corresponding to each category of labels. This includes determining the call pattern characteristics of the stored data corresponding to different categories of labels within the time domain segment, verifying the effectiveness of the call pattern, and constructing a set of association patterns for each category of labels.
[0010] Determine the number of elements in the association pattern set corresponding to each category label, and calculate the association activity characterization value in combination with the call frequency of the stored data corresponding to the category label, so as to classify the association interference category of each category label;
[0011] In response to the retrieval of stored data, the data is retrieved based on the associated interference category corresponding to the category label of the retrieved stored data, including,
[0012] Based on the set of association rules corresponding to the category labels, several pre-call category labels are determined. Based on the call records of the stored data corresponding to each pre-call category label, the call characteristics are determined to analyze the loading activity characterization parameters, determine the loading priority sequence of each pre-call category label, and preload the stored data corresponding to each pre-call category label according to the loading priority sequence, and prioritize calling the preloaded stored data.
[0013] Alternatively, the corresponding stored data can be retrieved from the database;
[0014] The calling characteristics include the amount of data and the calling frequency.
[0015] Furthermore, the process of determining the retrieval patterns of stored data corresponding to different categories of labels within a time domain segment includes,
[0016] Determine the stored data of each category label invoked in a single invocation event, and construct the association relationship between each category label and other category labels;
[0017] Calculate the probability of occurrence of the association relationship corresponding to each category label, and determine the occurrence probability as the calling pattern feature;
[0018] The single call event refers to the storage data of several category tags being called by a single caller within a predetermined time. If the storage data of a single category tag and the storage data of any category tag are both called, then it is determined that there is a correlation.
[0019] Furthermore, verify the validity of the call pattern, including,
[0020] If the call pattern feature of the association corresponding to the category label is greater than or equal to the call pattern feature threshold, then the call pattern of the association corresponding to the category label is verified as valid.
[0021] Furthermore, the process of constructing a set of association patterns for each category of labels includes,
[0022] Identify other category labels that have a valid association with the category labels based on calling patterns;
[0023] Each category label is stored in the same data set to form the association pattern set.
[0024] Furthermore, the process of calculating the associated activity representation value includes,
[0025] The ratio of the number of elements in the association pattern set corresponding to the category label to the element number threshold is used as the first association activity feature;
[0026] The ratio of the call frequency of the stored data corresponding to the category label to the call frequency threshold is used as the second associated activity feature;
[0027] The sum of the first associated activity feature and the second associated activity feature is determined as the associated activity representation value.
[0028] Furthermore, the association interference categories of each category label are divided, including,
[0029] If the association activity value is greater than or equal to the association activity threshold, the category label will be classified as a strong association interference category.
[0030] If the association activity value is less than the association activity threshold, the category label will be classified as a weak association interference category.
[0031] Furthermore, the data is retrieved based on the association interference category corresponding to the category label of the retrieved stored data, including:
[0032] If the correlation interference category is a strong correlation interference category, then based on the correlation rule set corresponding to the category label, several pre-call category labels are determined, and the call characteristics are determined based on the call records of the stored data corresponding to each pre-call category label, so as to analyze the loading activity characterization parameters, determine the loading priority sequence of each pre-call category label, and preload the stored data corresponding to each pre-call category label according to the loading priority sequence, and preferentially call several preloaded stored data.
[0033] If the correlation interference category is a weak correlation interference category, the corresponding stored data will be retrieved from the database.
[0034] Furthermore, the process of determining several pre-invoked category labels includes,
[0035] Obtain the set of association patterns corresponding to category labels;
[0036] The category labels within the set of association rules are determined as the pre-call category labels.
[0037] Furthermore, the process of loading active characterization parameters is analyzed, including,
[0038] The ratio of the amount of stored data corresponding to the pre-call category label to the data amount threshold is used as the first loading activity feature;
[0039] The ratio of the frequency of access to the stored data to a baseline threshold for the frequency of access is used as the second loading activity feature;
[0040] The first loading activity feature and the second loading activity feature are weighted and summed to determine the loading activity representation parameter of the pre-call category label.
[0041] Furthermore, the process of determining the loading priority sequence of each pre-invoked category label includes,
[0042] Sort the loading activity representation parameters of each of the pre-invoked category tags in descending order;
[0043] The descending order is used as the loading priority sequence for the corresponding pre-call category tags.
[0044] Compared with existing technologies, this invention determines the data source corresponding to the required stored data, sets category labels for the required stored data based on the data source, and stores the required stored data in a database; it performs activity analysis based on the call records of the stored data corresponding to each category label; it determines the number of elements in the association pattern set corresponding to each category label, and calculates the association activity characterization value in combination with the call frequency of the stored data corresponding to the corresponding category label, so as to classify the association interference categories of each category label; in response to the storage data being called, it adaptively calls the data based on the association interference category corresponding to the category label of the called storage data. Thus, under the premise of a large number of trains and many sensors involved, when accessing and querying targeted data from a database storing a large amount of data, it improves access efficiency while avoiding repeated query storms on the database.
[0045] In particular, this invention considers the retrieval status of stored data under already stored category tags, and then performs activity analysis on the category tags to identify the regular characteristics of the stored data, verify its validity, and construct a set of association patterns for high-frequency associated stored data combinations. This adapts to changes in stored data retrieval patterns, avoids constant database queries, reduces I / O overhead, and improves the access speed of stored data. Furthermore, the validity verification can eliminate data with low correlation to category tags, avoiding invalid preloading in the future, reducing the waste of storage and computing resources, and improving the efficiency of stored data caching. The construction of the set of association patterns facilitates efficient and rapid preloading of related stored data, especially in scenarios with high real-time requirements such as train monitoring, reducing query access latency and improving database response speed. Therefore, this invention optimizes the stored data preloading strategy by analyzing the correlation between the retrieval of stored data under each category tag, thereby improving the efficiency of stored data retrieval.
[0046] In particular, based on the construction of a set of association patterns, this invention considers the number of elements in the set and the frequency of calls to the stored data corresponding to each element to evaluate the activity level between frequently associated category tags. Specifically, for cases with a large number of elements and high call frequency, it indicates that the association pattern set has a wider range of associations and covers more comprehensive scenarios, or that several category tags within the set are called synchronously more frequently and the logical dependencies between stored data are strong. This places stricter requirements on the real-time performance of stored data queries, necessitating pre-loading of stored data to reduce real-time query pressure and improve query access efficiency. Conversely, for cases with a small number of elements and low call frequency, the associations between category tags may be occasional, requiring on-demand queries to avoid resource waste. Therefore, this application uses an association activity characterization value to represent the activity level of association calls between category tags within the set of association patterns, providing data support for subsequent classification of association interference categories for each category tag. Furthermore, in response to the call of stored data, the invention adaptively calls stored data based on the activity level of the category tags of the called stored data. This invention improves access efficiency while avoiding repeated query storms when selectively accessing and querying from databases storing large amounts of data.
[0047] In particular, when category labels of strongly correlated interference categories are invoked, this invention identifies pre-invocation category labels that have a high degree of regular correlation with them. By determining the invocation characteristics of the stored data corresponding to the pre-invocation category labels, namely, the data volume and invocation frequency, it reflects the degree of matching between the pre-loaded stored data and the actual query invocation requirements. In train data application scenarios with high real-time requirements, this ensures fast access to high-frequency pre-loaded stored data and reduces invocation latency. Therefore, this invention loads active characterization parameters to characterize the loading frequency of pre-invocation category labels and the corresponding data volume, providing data support for subsequently determining the corresponding loading priority sequence, dynamically adjusting the pre-loading order of stored data, and improving the query invocation efficiency of stored data. This invention improves access efficiency while avoiding repeated query storms on the database when specifically accessing and querying from databases storing large amounts of data. Attached Figure Description
[0048] Figure 1 A schematic diagram illustrating the steps of a train data classification and storage method according to an embodiment of the invention;
[0049] Figure 2 A logic diagram for verifying the validity of the calling pattern in an embodiment of the invention;
[0050] Figure 3 A logical decision diagram for classifying the correlation interference categories of labels in various categories according to embodiments of the invention;
[0051] Figure 4 This is a logic decision diagram for retrieving data according to an embodiment of the invention. Detailed Implementation
[0052] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0053] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0054] It should be noted that in the description of this invention, the terms "inner" and "other" indicate directions or positional relationships based on the directions or positional relationships shown in the drawings. This is merely for ease of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0055] Please see Figure 1The diagram illustrates the steps of a train data classification and storage method according to an embodiment of the present invention. The train data classification and storage method according to an embodiment of the present invention includes:
[0056] Step S1: Determine the data source corresponding to the data to be stored, set category labels for the data to be stored based on the data source, and store the data to be stored in the database;
[0057] Step S2 involves conducting an activity analysis based on the call records of the stored data corresponding to each category of labels. This includes determining the call pattern characteristics of the stored data corresponding to different categories of labels within the time domain segment, verifying the validity of the call patterns, and constructing a set of association patterns for each category of labels.
[0058] Step S3: Determine the number of elements in the association pattern set corresponding to each category label, and calculate the association activity characterization value in combination with the call frequency of the stored data corresponding to the category label, so as to classify the association interference category of each category label;
[0059] Step S4: In response to the retrieval of stored data, the data is retrieved based on the association interference category corresponding to the category label of the retrieved stored data, including:
[0060] Based on the set of association rules corresponding to the category labels, several pre-call category labels are determined. Based on the call records of the stored data corresponding to each pre-call category label, the call characteristics are determined to analyze the loading activity characterization parameters, determine the loading priority sequence of each pre-call category label, and preload the stored data corresponding to each pre-call category label according to the loading priority sequence, and prioritize calling the preloaded stored data.
[0061] Alternatively, the corresponding stored data can be retrieved from the database;
[0062] The calling characteristics include the amount of data and the calling frequency.
[0063] In this embodiment, the data sources include several sensors deployed on the train, such as a brake pressure sensor that collects and receives the pressure of the train's brake cylinder / pipe, and a track gauge detection sensor that monitors the track condition. Here, several stored data obtained based on a single data source are set as a category label, which will not be elaborated further.
[0064] It is understandable that during the process of calling stored data, there is a sequence of calls. Also, based on the purpose of calling stored data, due to the relationship between the stored data to be called, it may be that stored data corresponding to several data sources is called. Therefore, in order to reflect the calling pattern of stored data as much as possible, in this embodiment, the time domain segment must be greater than 48 hours.
[0065] In some possible implementations, those skilled in the art need to analyze data from multiple data sources to obtain analysis results. For example, a train fault analysis model containing several data points may be trained, which may require calling data from multiple data sources as input to the train fault analysis model. Of course, there may be other methods, which will not be elaborated here.
[0066] Specifically, the process of determining the retrieval patterns of stored data corresponding to different categories of labels within a time domain segment includes,
[0067] Determine the stored data of each category label invoked in a single invocation event, and construct the association relationship between each category label and other category labels;
[0068] Calculate the probability of occurrence of the association relationship corresponding to each category label, and determine the occurrence probability as the calling pattern feature;
[0069] The single call event refers to the storage data of several category tags being called by a single caller within a predetermined time. If the storage data of a single category tag and the storage data of any category tag are both called, then it is determined that there is a correlation.
[0070] Specifically, the purpose of setting a scheduled time is to observe the user's demand for multiple data sources within a short period of time, and the scheduled time is selected within the range of [3min, 5min].
[0071] Specifically, please refer to Figure 2 As shown, this is a logic decision diagram for verifying the validity of the calling pattern in an embodiment of the present invention. Verifying the validity of the calling pattern includes:
[0072] If the calling pattern feature of the association corresponding to the category label is greater than or equal to the calling pattern feature threshold, then the calling pattern of the association corresponding to the category label is verified as valid.
[0073] If the call pattern feature of the association corresponding to the category label is less than the call pattern feature threshold, then verifying the call pattern of the association corresponding to the category label is invalid.
[0074] Specifically, the process of constructing a set of association patterns for each category of tags includes,
[0075] Identify other category labels that have a valid association with the category labels based on calling patterns;
[0076] Each category label is stored in the same data set to form the association pattern set.
[0077] Specifically, this invention considers the retrieval status of stored data under already stored category tags, and then performs activity analysis on the category tags to identify the regular characteristics of the stored data, verify its validity, and construct a set of association patterns for high-frequency associated stored data combinations. This adapts to changes in stored data retrieval patterns, avoids constant database queries, reduces I / O overhead, and improves the access speed of stored data. Furthermore, the validity verification can eliminate data with low correlation to category tags, avoiding invalid preloading in the future, reducing the waste of storage and computing resources, and improving the efficiency of stored data caching. The construction of the set of association patterns facilitates efficient and rapid preloading of related stored data, especially in scenarios with high real-time requirements such as train monitoring, reducing query access latency and improving database response speed. Therefore, this invention optimizes the stored data preloading strategy by analyzing the correlation between the retrieval of stored data under each category tag, thereby improving the efficiency of stored data retrieval.
[0078] Specifically, the process of calculating the association activity representation value includes,
[0079] The ratio of the number of elements in the association pattern set corresponding to the category label to the element number threshold is used as the first association activity feature;
[0080] The ratio of the call frequency of the stored data corresponding to the category label to the call frequency threshold is used as the second associated activity feature;
[0081] The sum of the first associated activity feature and the second associated activity feature is determined as the associated activity representation value.
[0082] In this embodiment, the purpose of setting the element quantity threshold and the call frequency threshold is to characterize the situation where there are many other category tags associated with the category tag, the association is close, and the call is frequent. By obtaining the call records of the stored data corresponding to each category tag, the element quantity data of several association rule sets corresponding to the same category tag and the call frequency data of several corresponding category tag stored data are extracted. The mean element quantity and the mean call frequency are calculated. Based on the purpose of setting the above two thresholds, the element quantity threshold is determined as the product of the mean element quantity and the quantity deviation coefficient, and the call frequency threshold is determined as the product of the mean call frequency and the frequency deviation coefficient. The quantity deviation coefficient is selected in the interval [1.2, 1.4], and the frequency deviation coefficient is selected in the interval [1.2, 1.25].
[0083] Specifically, please refer to Figure 3 The above describes a logical decision diagram for classifying the correlation interference categories of each category label according to an embodiment of the present invention. The classification of the correlation interference categories of each category label includes:
[0084] If the association activity value is greater than or equal to the association activity threshold, the category label will be classified as a strong association interference category.
[0085] If the association activity value is less than the association activity threshold, the category label will be classified as a weak association interference category.
[0086] The threshold for identifying active associations is selected within the range [2.17, 2.23].
[0087] Specifically, please refer to Figure 4 As shown, this is a logic decision diagram for data retrieval in an embodiment of the present invention. Data retrieval is based on the association interference category corresponding to the category label of the stored data being retrieved, including:
[0088] If the correlation interference category is a strong correlation interference category, then based on the correlation rule set corresponding to the category label, several pre-call category labels are determined, and the call characteristics are determined based on the call records of the stored data corresponding to each pre-call category label, so as to analyze the loading activity characterization parameters, determine the loading priority sequence of each pre-call category label, and preload the stored data corresponding to each pre-call category label according to the loading priority sequence, and preferentially call several preloaded stored data.
[0089] If the correlation interference category is a weak correlation interference category, the corresponding stored data will be retrieved from the database.
[0090] Specifically, based on the construction of a set of association rules, this invention considers the number of elements in the set and the frequency of calling the stored data corresponding to each element to evaluate the activity level between frequently associated category tags. In the case of a large number of elements and a high call frequency, it can be said that the association range of the set of association rules is wider and the coverage of scenarios is more comprehensive, or that several category tags in the set of association rules are called synchronously more frequently and the logical dependencies between the stored data are stronger. The real-time requirements for querying and calling the stored data are more stringent, and it is necessary to intervene in advance to preload the stored data to reduce the real-time query pressure and improve the efficiency of query access. In the case of a small number of elements and a low call frequency, the association between category tags may be occasional, and queries and calls should be made on demand to avoid wasting resources.
[0091] Therefore, this application uses the association activity representation value to characterize the activity level of association calls between category labels within the association pattern set, providing data support for the subsequent division of association interference categories of each category label. Furthermore, in response to the retrieval of stored data, the storage data is adaptively retrieved based on the activity level of the category labels of the retrieved stored data. This invention improves access efficiency while avoiding repeated query storms on the database when accessing and querying data from a database containing large amounts of data.
[0092] Specifically, the process of determining several pre-invoked category labels includes,
[0093] Obtain the set of association patterns corresponding to category labels;
[0094] The category labels within the set of association rules are determined as the pre-call category labels.
[0095] Specifically, the analysis of the process of loading active characterization parameters includes,
[0096] The ratio of the amount of stored data corresponding to the pre-call category label to the data amount threshold is used as the first loading activity feature;
[0097] The ratio of the frequency of access to the stored data to a baseline threshold for the frequency of access is used as the second loading activity feature;
[0098] The first loading activity feature and the second loading activity feature are weighted and summed to determine the loading activity representation parameter of the pre-call category label.
[0099] Specifically, during the preloading of stored data, given a strong correlation with the data being retrieved, prioritizing the preloading of large amounts of stored data can ensure the efficiency of data query and retrieval, and guarantee fast access to stored data. Therefore, in implementation, the amount of data is taken into consideration first, and a slightly higher weight is assigned to the first loading activity feature calculated based on the quantity. Thus, when performing weighted summation, the weight of the first loading activity feature is set to 0.6, and the weight of the second loading activity feature is set to 0.4.
[0100] In this embodiment, the purpose of setting the data volume threshold and the call frequency benchmark threshold is to characterize the situation where the preloaded data matches the actual query call demand to a high degree and the call is given priority. By obtaining the call records of the stored data corresponding to each category label, the data volume data and the corresponding call frequency data of the stored data corresponding to the same preloaded category label are extracted several times. The average value of the data volume and the average value of the call frequency are calculated. Based on the purpose of setting the above two thresholds, the data volume threshold is determined to be the product of the average value of the data volume and the data offset coefficient, and the call frequency benchmark threshold is determined to be the product of the average value of the call frequency and the call deviation coefficient. The data offset coefficient is selected in the interval [1.25, 1.3], and the call deviation coefficient is selected in the interval [1.3, 1.35].
[0101] Specifically, when a category label of a strongly correlated interference category is invoked, this invention identifies a pre-invocation category label that has a high degree of regular correlation with it. By determining the invocation characteristics of the stored data corresponding to the pre-invocation category label, namely the data volume and invocation frequency, it reflects the degree of matching between the pre-loaded stored data and the actual query invocation requirements. In train data application scenarios with high real-time requirements, this ensures fast access to high-frequency pre-loaded stored data and reduces invocation latency. Therefore, this invention loads an active characterization parameter to characterize the loading frequency of the pre-invocation category label and the corresponding data volume, providing data support for subsequently determining the corresponding loading priority sequence, dynamically adjusting the pre-loading order of stored data, and improving the query invocation efficiency of stored data. This invention improves access efficiency while avoiding repeated query storms on the database when specifically accessing and querying from a database storing large amounts of data.
[0102] Specifically, the process of determining the loading priority sequence of each pre-invoked category tag includes,
[0103] Sort the loading activity representation parameters of each of the pre-invoked category tags in descending order;
[0104] The descending order is used as the loading priority sequence for the corresponding pre-call category tags.
[0105] If the train data classification and storage method of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for classifying and storing train data, characterized in that, include: Determine the data source corresponding to the data to be stored, set category labels for the data to be stored based on the data source, and store the data to be stored in the database; Active analysis is performed based on the call records of the stored data corresponding to each category of labels. This includes determining the call pattern characteristics of the stored data corresponding to different categories of labels within the time domain segment, verifying the effectiveness of the call pattern, and constructing a set of association patterns for each category of labels. Determine the number of elements in the association pattern set corresponding to each category label, and calculate the association activity characterization value in combination with the call frequency of the stored data corresponding to the category label, so as to classify the association interference category of each category label; In response to the retrieval of stored data, the data is retrieved based on the associated interference category corresponding to the category label of the retrieved stored data, including, If the correlation interference category is a strong correlation interference category, then based on the correlation rule set corresponding to the category label, several pre-call category labels are determined, and the call characteristics are determined based on the call records of the stored data corresponding to each pre-call category label, so as to analyze the loading activity characterization parameters, determine the loading priority sequence of each pre-call category label, and preload the stored data corresponding to each pre-call category label according to the loading priority sequence, and preferentially call several preloaded stored data. If the correlation interference category is a weak correlation interference category, the corresponding stored data will be retrieved from the database; The calling characteristics include data volume and calling frequency; The process of determining the retrieval patterns of stored data corresponding to different categories of labels within a time domain segment includes, Determine the stored data of each category label invoked in a single invocation event, and construct the association relationship between each category label and other category labels; Calculate the probability of occurrence of the association relationship corresponding to each category label, and determine the occurrence probability as the calling pattern feature; The single call event refers to the storage data of several category tags being called by a single caller within a predetermined time. If the storage data of a single category tag and the storage data of any category tag are both called, it is determined that there is a correlation. The association interference categories for each of the aforementioned category labels include, If the association activity value is greater than or equal to the association activity threshold, the category label will be classified as a strong association interference category. If the association activity value is less than the association activity threshold, the category label will be classified as a weak association interference category. The process of constructing a set of association patterns for each category of tags includes, Identify other category labels that have a valid association with the category labels based on calling patterns; Each category label is stored in the same data set to form the association pattern set.
2. The train data classification and storage method according to claim 1, characterized in that, Verify the validity of the call pattern, including: If the call pattern feature of the association corresponding to the category label is greater than or equal to the call pattern feature threshold, then the call pattern of the association corresponding to the category label is verified as valid.
3. The train data classification and storage method according to claim 1, characterized in that, The process of calculating the associated activity representation value includes, The ratio of the number of elements in the association pattern set corresponding to the category label to the element number threshold is used as the first association activity feature; The ratio of the call frequency of the stored data corresponding to the category label to the call frequency threshold is used as the second associated activity feature; The sum of the first associated activity feature and the second associated activity feature is determined as the associated activity representation value.
4. The train data classification and storage method according to claim 1, characterized in that, The process of determining several pre-invocation category labels includes, Obtain the set of association patterns corresponding to category labels; The category labels within the set of association rules are determined as the pre-call category labels.
5. The train data classification and storage method according to claim 1, characterized in that, The process of analyzing and loading active characterization parameters includes, The ratio of the amount of stored data corresponding to the pre-call category label to the data amount threshold is used as the first loading activity feature; The ratio of the call frequency of the stored data corresponding to the pre-call category label to the call frequency baseline threshold is used as the second loading activity feature; The first loading activity feature and the second loading activity feature are weighted and summed to determine the loading activity representation parameter of the pre-call category label.
6. The train data classification and storage method according to claim 5, characterized in that, The process of determining the loading priority sequence of each pre-invoked category tag includes, Sort the loading activity representation parameters of each of the pre-invoked category tags in descending order; The descending order is used as the loading priority sequence for the corresponding pre-call category tags.
Citation Information
Patent Citations
Classified storage method for information technology data
CN119027236A
Medical data query method and device, equipment and storage medium
CN113569012A
Storage method and system based on e-commerce label data
CN115687353A