A database management method based on data analysis
By verifying the identity of the data source and conducting pre-transmission experiments, screening adjacent data packet groups, determining dependency characterization parameters, dividing data packet dependency tendencies, and performing protected transmission, the problem of database overload in data transmission is solved and the accuracy and efficiency of data transmission are improved.
Patent Information
- Application Number
- CN202510766288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies perform indiscriminate analysis, encryption, or backup of all data sources during data transmission, causing the database load to exceed its carrying capacity, reducing the speed and efficiency of data transmission, and affecting the accuracy of data transmission.
By acquiring data packets during the historical transmission process of the data source, constructing a high-frequency data category sequence, performing identity verification, conducting pre-transmission experiments, screening adjacent data packet groups, determining time domain and combination dependency characterization parameters, dividing data packet dependency tendencies, monitoring data category sensitive parameters, and performing protected transmission.
It improves the accuracy and efficiency of data transmission, reduces computing load, avoids database overload caused by indiscriminate analysis, and ensures the efficiency and security of data transmission.
Smart Images

Figure CN120296007B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a database management method based on data analysis. Background Art
[0002] With the rapid development of information technology, all industries are generating massive amounts of data. Traditional manual data management methods can no longer cope with such a huge amount of data. There is an urgent need to use database management systems to efficiently store and manage this data. In particular, as the value of data becomes increasingly prominent, the accuracy and security of data transmission are receiving more and more attention. Data transmission has become the cornerstone of operation in various fields. During the data transmission process, with the help of artificial intelligence and machine learning technology, the data transmission process can be comprehensively and finely optimized and managed in real time to ensure the efficiency, stability and security of data during the transmission process, thereby maximizing the efficiency and quality of data transmission, meeting the growing digital needs, and promoting the digitalization process of various industries to a new level.
[0003] Chinese patent publication number CN119149783A discloses a data management method and system for a multimodal database, comprising a multimodal data analysis module, a multimodal data storage module, and a multimodal data update module. This invention utilizes a multimodal data feature analysis model to extract deep-level features, improving the accuracy and depth of data recognition. It also optimizes the data storage structure and enhances query efficiency by allocating database storage using feature matching technology. It also introduces a feature-adaptive threshold factor Z and a knowledge graph-based CNN model to enhance the model's generalization and adaptability. Furthermore, through a multimodal database management and update model, it enables dynamic updates of database partitions, responds to user operation changes, maintains the database's timeliness and relevance, and enhances the user experience.
[0004] Chinese Patent Publication No.: CN114138933A, discloses an NLP-based big data analysis management system and method, including an analysis management system, a cloud server and a display unit, wherein the analysis management system is connected to the display unit, and the analysis management system includes a dedicated database, a data acquisition module, a data query module, a data processing module, a graphics processing module, an interactive processing module and a search engine simulation training module. This invention belongs to the field of big data analysis management technology, and specifically provides a data query based on keywords, key words, and instructions, which can simultaneously meet the requirements of real-time data display, real-time variable-dimensional data display, full-text search based on keywords, key words, and instructions, and information selection of data presentation indicators. Through massive data retrieval and large-scale data rendering, a good user experience effect is achieved. At the same time, the dedicated corpus can be updated, the accuracy of data collection is continuously improved, and the reliance on human errors is reduced.
[0005] However, the prior art still has the following problems:
[0006] In actual situations, when data is transmitted, most of the time, all data sources are analyzed indiscriminately, and data sources that may be at risk of loss or delay are encrypted or backed up indiscriminately. This will cause the database load to exceed its own carrying capacity, causing the database to be overloaded, thereby reducing the speed and efficiency of data transmission and affecting the accuracy of data transmission. Summary of the Invention
[0007] To this end, the present invention provides a database management method based on data analysis to solve the problem that in actual situations, when performing data transmission, most of the time, all data sources are analyzed indiscriminately, and data sources that may be at risk of loss or delay are encrypted or backed up indiscriminately. This will cause the database load to exceed its own carrying capacity, causing the database to be overloaded, thereby reducing the speed and efficiency of data transmission and affecting the accuracy of data transmission.
[0008] To achieve the above objectives, the present invention provides a database management method based on data analysis, which includes:
[0009] Obtaining several data packets corresponding to several historical transmission processes of the data source, extracting the data category of each data packet, constructing a high-frequency data category sequence, recording the number of occurrences of the same high-frequency data category sequence to analyze the historical transmission pattern of the data source, and verifying the identity of the data source;
[0010] In response to a data source requiring data transmission, a pre-transmission experiment is conducted based on an identity verification result to randomly extract a predetermined proportion of data packets of different data categories transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine an arrival time deviation, screen adjacent data packet groups, and determine a time-domain dependency characterization parameter based on decoding results of the adjacent data packet groups;
[0011] Decoding the data packets of the pre-transmission experiment to determine the number of undecodable data packets to determine the combination dependency characterization parameter;
[0012] Calculating a data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter, and classifying the data packet dependency tendency;
[0013] In response to the division result of the data packet dependency tendency, controlling the data transmission of the data source, including,
[0014] Monitor the pre-transmission experiment, analyze the data packet arrival value and decoding characterization value of the data packet corresponding to each data category based on the results of the pre-transmission experiment of the data source, calculate the data category sensitive parameters to locate the sensitive data category, select the protection processing data packet based on the positioning result, and perform protection transmission on the data packet.
[0015] Furthermore, the process of analyzing the historical transmission patterns of the data source and verifying the identity of the data source includes:
[0016] Record the number of similar transmission behaviors of the data source during several historical transmissions;
[0017] Calculate the ratio of the number of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio;
[0018] If the similar behavior transmission ratio is greater than a predetermined similar behavior transmission ratio threshold, it is determined that the data sources are identical;
[0019] The similar transmission behavior must satisfy the requirement that the high-frequency data category sequence corresponding to each data packet transmitted during the historical transmission process is the same as the high-frequency data category sequence corresponding to each data packet during any other historical transmission process.
[0020] Furthermore, based on the identity verification result, a pre-transmission experiment is performed, wherein:
[0021] If the data sources are identical, a preliminary transmission experiment is required to randomly select a predetermined proportion of different data category data packets transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time domain dependence characterization parameters based on the decoding results of the adjacent data packet groups;
[0022] If the data sources are not identical, no pre-transmission experiment is required.
[0023] Furthermore, the process of screening the adjacent data packet groups and determining the time-domain dependency characterization parameters based on the decoding results of the adjacent data packet groups includes:
[0024] Determine the arrival time deviation corresponding to each data packet;
[0025] If there is a data packet that meets the deviation condition, combining the data packet and the adjacent data packet into an adjacent data packet group;
[0026] determining whether each data packet in the adjacent data packet group can be completely decoded to determine the time-domain dependent data packet group;
[0027] determining a ratio of the number of time-domain dependent data packet groups to the total number of data packets as a time-domain dependence characterization parameter;
[0028] The deviation condition is that the deviation time corresponding to the data packet is greater than a predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.
[0029] Furthermore, the process of determining the combination dependency characterization parameter includes:
[0030] determining the difference between the number of undecodable packets and the number of corrupted packets;
[0031] The ratio of the difference to the number of undecodable data packets is determined as a combination dependency characterization parameter.
[0032] Furthermore, the process of calculating the data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter includes:
[0033] Determine the ratio of the time-domain dependency characterization parameter to the benchmark time-domain dependency characterization parameter as the time-domain dependency impact factor;
[0034] Determine the ratio of the combination dependency characterization parameter to the benchmark combination dependency characterization parameter as the combination dependency impact factor;
[0035] Determine the weighted sum of the time domain dependency impact factor and the combined dependency impact factor as a data packet dependency characteristic value.
[0036] Furthermore, the data transmission of the data source is controlled in response to the division result of the data packet dependency tendency, wherein,
[0037] If the data packet dependency characteristic value is greater than the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a strong dependency tendency, the pre-transmission experiment is monitored, and based on the result of the pre-transmission experiment of the data source, the data packet arrival value and the decoding representation value of the data packet corresponding to each data category are analyzed, and the data category sensitive parameter is calculated to locate the sensitive data category, and a protection processing data packet is selected based on the positioning result, and the data packet is protected and transmitted;
[0038] If the data packet dependency characteristic value is less than or equal to the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a weak dependency tendency.
[0039] Furthermore, the process of analyzing the packet arrival value and the decoding representation value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source includes:
[0040] Determine the number of decodable data packets corresponding to each data category as a data packet arrival value;
[0041] The ratio of the number of decodable data packets to the total number of data packets of the same data category is determined as a decoding characterization value.
[0042] Furthermore, the process of calculating the data category sensitive parameters includes:
[0043] Determine the ratio of the baseline data packet arrival value to the data packet arrival value corresponding to the data category as the arrival value impact factor;
[0044] Determine the ratio of the benchmark decoding representation value corresponding to the data category to the decoding representation value as the decoding representation value influencing factor;
[0045] A weighted sum of the arrival value impact factor and the decoding representation value impact factor is determined as the data category sensitive parameter.
[0046] Furthermore, the process of locating the sensitive data category, selecting a protection processing data packet based on the locating result, and performing protection transmission on the data packet includes:
[0047] If the data category sensitive parameter is greater than a preset sensitive parameter threshold, the data category is determined to be a sensitive data category;
[0048] Selecting data packets corresponding to sensitive data categories from untransmitted data packets as protection processing data packets;
[0049] Performing protection transmission on the protection-processed data packet;
[0050] The protection transmission is to construct the same data packet based on the protection processing data packet for synchronous transmission.
[0051] Compared with the prior art, the present invention obtains several data packets, extracts the data category of each of the data packets, constructs a high-frequency data category sequence, and verifies the identity of the data source to conduct a pre-transmission experiment; determines the arrival time deviation, screens adjacent data packet groups, and determines the time domain dependency characterization parameters; determines the number of undecodable data packets to determine the combination dependency characterization parameters; calculates the data packet dependency characteristic values and divides the data packet dependency tendencies; in response to the division results of the data packet dependency tendencies, controls the data transmission of the data source, monitors the pre-transmission experiment, calculates the data category sensitive parameters to locate the sensitive data category, selects the protection processing data packet based on the positioning result, and performs protection transmission on the data packet. The present invention improves the accuracy of data transmission by analyzing the data packet, locating the sensitive data category, and protecting the sensitive data.
[0052] In particular, by verifying the identity of the data source, a data basis is provided for determining whether to conduct a pre-transmission experiment. In actual situations, most of the data transmission status of the data source is judged by obtaining data from the data transmission operation process. However, for massive data, the data of the data transmission operation process is numerous. Directly analyzing it will waste a lot of computing power, increase the computing load, and even cause errors in the analysis. Based on this, the present invention first verifies the identity of the data source, and conducts a pre-transmission experiment on the data source with the same identity, providing a classification theoretical basis for the subsequent analysis of the degree of dependence of the data packet, thereby improving the accuracy and efficiency of data transmission.
[0053] In particular, by determining the time domain dependency characterization parameters, a data basis is provided for calculating the packet dependency characteristic values. During data transmission, if a slow transmission problem is encountered during the transmission process, the transmission process of the packet will change, resulting in the arrival time of the packet being inconsistent with the specified time. Adjacent packet groups with a high degree of dependence may not be successfully decoded due to the large difference in arrival time, which in turn causes the transmitted data to be unable to be identified or started. Based on this, the present invention considers determining the time domain dependency characterization parameters according to the results of the preliminary transmission experiment, providing a classification theoretical basis for the subsequent analysis of the dependency degree of the packet, thereby improving the accuracy and efficiency of data transmission.
[0054] In particular, by determining the combined dependency characterization parameters to provide a data basis for calculating the packet dependency characteristic values, during data transmission, if transmission loss or damage occurs during the transmission process, some data packets may not be successfully decoded due to failure to match with other data packets, thereby causing the transmitted data to be unable to be identified or started. Based on this, the difference between the number of undecodable data packets and the number of damaged data packets is calculated to obtain potential situations where decoding failure may occur due to the need to combine with other data packets for decoding, and then calculate the combined dependency characterization parameters, providing a classification theoretical basis for subsequent analysis of the degree of dependency of data packets, thereby improving the accuracy and efficiency of data transmission.
[0055] In particular, by calculating the sensitive parameters of the data category, the sensitive data category is located, providing a data basis for the data transmission process. In actual situations, most of the time, all data is backed up to ensure the security and accurate arrival of the data. However, for massive data, backing up all data will not only waste space, but also cause problems such as low data transmission efficiency. Especially for data that requires high timeliness, if the data is found to be damaged after transmission, a second transmission will make the data provision untimely, affecting the use of the data. Based on this, the present invention calculates the data category sensitive parameters through the data packet arrival value and the decoding characterization value, and then locates the sensitive data category, performs targeted backup and synchronous transmission, and improves the accuracy and efficiency of data transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A schematic diagram of the steps of a database management method based on data analysis according to an embodiment of the invention;
[0057] Figure 2 A logic block diagram of a pre-transmission experiment based on the identity verification result according to an embodiment of the present invention;
[0058] Figure 3 This is a logic block diagram of controlling data transmission of the data source in response to the division result of data packet dependency tendency according to an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0060] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0061] See also Figure 1 , Figure 1 Schematic diagram of the method steps of the database management method based on data analysis according to an embodiment of the invention. The database management method based on data analysis of the present invention comprises:
[0062] Step S1: Obtain a number of data packets corresponding to several historical transmission processes of the data source, extract the data category of each data packet, construct a high-frequency data category sequence, record the number of occurrences of the same high-frequency data category sequence to analyze the historical transmission pattern of the data source, and verify the identity of the data source;
[0063] Step S2: In response to the data source needing to transmit data, a pre-transmission experiment is conducted based on the identity verification result to randomly extract a predetermined proportion of data packets of different data categories transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time domain dependence characterization parameter based on the decoding results of the adjacent data packet groups;
[0064] Step S3, decoding the data packets of the pre-transmission experiment, determining the number of undecodable data packets, and using this to determine the combination dependency characterization parameter;
[0065] Calculating a data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter, and classifying the data packet dependency tendency;
[0066] Step S4, in response to the result of the division of the data packet dependency tendency, controlling the data transmission of the data source, including:
[0067] Monitor the pre-transmission experiment, analyze the data packet arrival value and decoding characterization value of the data packet corresponding to each data category based on the results of the pre-transmission experiment of the data source, calculate the data category sensitive parameters to locate the sensitive data category, select the protection processing data packet based on the positioning result, and perform protection transmission on the data packet.
[0068] Specifically, there is no limitation on specific data categories. For example, in implementation, data categories include but are not limited to application layer data, transport layer data, and network layer data. It can be understood that the purpose of classifying data is to determine the sensitive data category and then protect the transmission. Therefore, those skilled in the art can also divide the data categories according to other basis as long as it is reasonable, which will not be elaborated here.
[0069] Specifically, the arrival time deviation represents the difference between the arrival time difference of a data packet and an adjacent data packet and the specified arrival time difference. In implementation, by obtaining several successfully transmitted high-frequency data category sequences, recording the historical arrival time differences of the corresponding data packets, and determining the average of each historical arrival time difference as the specified arrival time difference, those skilled in the art may also use other methods to determine the specified arrival time difference, which will not be repeated here.
[0070] Specifically, the process of analyzing the historical transmission patterns of the data source and verifying the identity of the data source includes:
[0071] Record the number of similar transmission behaviors of the data source during several historical transmissions;
[0072] Calculate the ratio of the number of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio;
[0073] If the similar behavior transmission ratio is greater than a predetermined similar behavior transmission ratio threshold, it is determined that the data sources are identical;
[0074] The similar transmission behavior must satisfy the requirement that the high-frequency data category sequence corresponding to each data packet transmitted in the historical transmission process is the same as the high-frequency data category sequence corresponding to each data packet in any other historical transmission process.
[0075] In implementation, it is necessary to determine the arrangement of data packets, assign data packet category labels based on the data categories of the data packets, and then obtain the arrangement of labels, and determine the arrangement of labels as the total sequence of data categories.
[0076] The fragments in the total sequence of data categories are determined as data category sequences. The data category sequence contains at least 3 labels. Then, the probability of occurrence of each data category sequence in the total sequence of data categories is determined. If the probability of occurrence of a data category sequence is greater than a predetermined probability threshold, the data category sequence is determined to be a high-frequency data category sequence.
[0077] In implementation, the probability threshold is predetermined, and the occurrence probability of each data category sequence corresponding to several historical transmission processes is pre-stated, the average occurrence probability is solved, and the probability threshold is set to 1.15 times the average occurrence probability.
[0078] Specifically, the predetermined similar behavior transmission ratio threshold represents the minimum occurrence frequency of similar transmission behaviors corresponding to the same data source, so the predetermined similar behavior transmission ratio threshold is set to be selected within the interval [0.3, 0.5].
[0079] Specifically, by verifying the identity of the data source, a data basis is provided for determining whether to conduct a pre-transmission experiment. In actual situations, most of the data transmission status of the data source is judged by obtaining data from the data transmission operation process. However, for massive data, the data of the data transmission operation process is numerous. Directly analyzing it will waste a lot of computing power, increase the computing load, and even cause errors in the analysis. Based on this, the present invention first verifies the identity of the data source, and conducts a pre-transmission experiment on the data source with the same identity, providing a classification theoretical basis for the subsequent analysis of the degree of dependence of the data packet, thereby improving the accuracy and efficiency of data transmission.
[0080] See also Figure 2 , Figure 2 This is a logic block diagram of a pre-transmission experiment based on the identity verification result of an embodiment of the invention. Specifically, based on the identity verification result, a pre-transmission experiment is performed, wherein:
[0081] If the data sources are identical, a preliminary transmission experiment is required to randomly select a predetermined proportion of different data category data packets transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time domain dependence characterization parameters based on the decoding results of the adjacent data packet groups;
[0082] If the data sources are not identical, no pre-transmission experiment is required.
[0083] In implementation, the predetermined ratio is set to 10%. For example, there are ten data packets of the application layer data category and ten data packets of the transport layer data category. According to the predetermined ratio, one data packet of the application layer data category and one data packet of the transport layer data category are extracted. It can be understood that the number of extracted data packets is rounded up based on the predetermined ratio. Of course, those skilled in the art can increase or decrease it according to the number of data packets as long as it is reasonable. This will not be repeated here.
[0084] Specifically, the process of screening adjacent data packet groups and determining the time-domain dependency characterization parameters based on the decoding results of the adjacent data packet groups includes:
[0085] Determine the arrival time deviation corresponding to each data packet;
[0086] If there is a data packet that meets the deviation condition, combining the data packet and the adjacent data packet into an adjacent data packet group;
[0087] determining whether each data packet in the adjacent data packet group can be completely decoded to determine the time-domain dependent data packet group;
[0088] determining a ratio of the number of time-domain dependent data packet groups to the total number of data packets as a time-domain dependence characterization parameter;
[0089] The deviation condition is that the deviation time corresponding to the data packet is greater than a predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.
[0090] Specifically, the adjacent data packet group that cannot be completely decoded is determined as a time domain dependent data packet group. It can be understood that if each data packet in the adjacent data packet group can be completely decoded, it means that the adjacent data packet group is not affected by the time domain and it is considered that there is no time domain dependence.
[0091] It can be understood that the adjacent data packet group consists of two data packets or three data packets. For the first data packet, it and the adjacent subsequent data packet form an adjacent data packet group. For the last data packet, it and the adjacent previous data packet form an adjacent data packet group.
[0092] Specifically, there is no limitation on the specific method of decoding. Protocol analysis tools such as Wireshark and tcpdump can be used for decoding. The corresponding commands can be entered in the command line of the tool to capture the data packet and output it to the terminal. A programming language library can also be used for decoding. The sniff function in scapy is used to capture the data packet, and then the parsing function provided by it is used to parse the data packet. The header information and payload content of the data packet can be parsed layer by layer. Those skilled in the art can select the corresponding decoding tool from the existing data according to their needs, and this will not be repeated.
[0093] Specifically, the predetermined time deviation is predetermined, and the arrival time deviations of each data packet in several historical transmission processes are pre-recorded, and the average value of the arrival time deviations corresponding to each data packet in each historical transmission process is determined as the predetermined time deviation.
[0094] Specifically, by determining the time domain dependency characterization parameters, a data basis is provided for calculating the packet dependency characteristic values. During data transmission, if a slow transmission problem is encountered during the transmission process, the transmission process of the packet will change, resulting in the arrival time of the packet being inconsistent with the specified time. Adjacent packet groups with a high degree of dependence may not be successfully decoded due to the large difference in arrival time, which in turn causes the transmitted data to be unable to be recognized or started. Based on this, the present invention considers determining the time domain dependency characterization parameters according to the results of the preliminary transmission experiment, providing a classification theoretical basis for the subsequent analysis of the dependency degree of the packet, thereby improving the accuracy and efficiency of data transmission.
[0095] Specifically, the process of determining the combination dependency characterization parameters includes:
[0096] determining the difference between the number of undecodable packets and the number of corrupted packets;
[0097] The ratio of the difference to the number of undecodable data packets is determined as a combination dependency characterization parameter.
[0098] Specifically, by determining the combined dependency characterization parameters to provide a data basis for calculating the packet dependency characteristic values, during data transmission, if transmission loss or damage occurs during the transmission process, some data packets may not be successfully decoded due to failure to match with other data packets, thereby causing the transmitted data to be unable to be identified or started. Based on this, the difference between the number of undecodable data packets and the number of damaged data packets is calculated to obtain potential situations where decoding failure may occur due to the need to combine with other data packets for decoding, and then calculate the combined dependency characterization parameters, providing a classification theoretical basis for subsequent analysis of the degree of dependency of data packets, thereby improving the accuracy and efficiency of data transmission.
[0099] Specifically, the process of calculating the data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter includes:
[0100] Determine the ratio of the time-domain dependency characterization parameter to the benchmark time-domain dependency characterization parameter as the time-domain dependency impact factor;
[0101] Determine the ratio of the combination dependency characterization parameter to the benchmark combination dependency characterization parameter as the combination dependency impact factor;
[0102] Determine the weighted sum of the time domain dependency impact factor and the combined dependency impact factor as a data packet dependency characteristic value.
[0103] Specifically, the benchmark time-domain dependency characterization parameter is pre-calculated data, and the corresponding time-domain dependency characterization parameters during several successful data transmission processes are obtained in advance, and the benchmark time-domain dependency characterization parameter is determined to be selected within 1.08 to 1.21 times the average value of each time-domain dependency characterization parameter.
[0104] Specifically, the benchmark combination dependency characterization parameters are pre-calculated data, and the corresponding combination dependency characterization parameters during several successful data transmission processes are obtained in advance, and the benchmark combination dependency characterization parameters are determined to be selected within 1.11 to 1.23 times the average value of each combination dependency characterization parameter.
[0105] Specifically, the weight coefficients of the time-domain dependency impact factor and the combination-dependence impact factor are 1, the weight coefficient of the time-domain dependency impact factor is 0.43, and the weight coefficient of the combination-dependence impact factor is 0.57.
[0106] Please participate Figure 3 , Figure 3 This is a logic block diagram of controlling the data transmission of the data source in response to the result of the division of the data packet dependency tendency according to an embodiment of the invention. Specifically, in response to the result of the division of the data packet dependency tendency, controlling the data transmission of the data source, wherein:
[0107] If the data packet dependency characteristic value is greater than the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a strong dependency tendency, the pre-transmission experiment is monitored, and based on the result of the pre-transmission experiment of the data source, the data packet arrival value and the decoding representation value of the data packet corresponding to each data category are analyzed, and the data category sensitive parameter is calculated to locate the sensitive data category, and a protection processing data packet is selected based on the positioning result, and the data packet is protected and transmitted;
[0108] If the data packet dependency characteristic value is less than or equal to the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a weak dependency tendency.
[0109] Specifically, the data packet dependency characteristic value threshold represents the highest dependency degree of the data packet in a decodable state, so the dependency characteristic value threshold is determined to be selected within the interval [0.65, 0.83].
[0110] Specifically, the process of analyzing the packet arrival value and the decoding representation value of the data packets corresponding to each data category based on the results of the pre-transmission experiment of the data source includes:
[0111] Determine the number of decodable data packets corresponding to each data category as a data packet arrival value;
[0112] The ratio of the number of decodable data packets to the total number of data packets of the same data category is determined as a decoding characterization value.
[0113] Specifically, the process of calculating sensitive parameters of data categories includes:
[0114] Determine the ratio of the baseline data packet arrival value to the data packet arrival value corresponding to the data category as the arrival value impact factor;
[0115] Determine the ratio of the benchmark decoding representation value corresponding to the data category to the decoding representation value as the decoding representation value influencing factor;
[0116] A weighted sum of the arrival value impact factor and the decoding representation value impact factor is determined as the data category sensitive parameter.
[0117] Specifically, the benchmark data packet arrival value is calculated through historical data, and the number of data packets corresponding to the same data category of several successful transmission data sources is obtained in advance to determine the average number of transmission data packets as the benchmark data packet arrival value.
[0118] Specifically, the benchmark decoding characterization value is calculated through historical data, and several successfully transmitted decoding characterization values are obtained in advance to determine an average value of the several decoding characterization values as the benchmark decoding characterization value.
[0119] Specifically, the sum of the weight coefficients of the arrival value influence factor and the decoding representation value influence factor is 1, the weight coefficient of the arrival value influence factor is 0.51, and the weight coefficient of the decoding representation value influence factor is 0.49.
[0120] Specifically, by calculating the data category sensitive parameters, the sensitive data category is located to provide a data basis for the data transmission process. In actual situations, most of the time, all data is backed up to ensure the security and accurate arrival of the data. However, for massive data, backing up all data will not only waste space, but also cause problems such as low data transmission efficiency. Especially for data that requires high timeliness, if the data is found to be damaged after transmission, a second transmission will make the data provision untimely, affecting the use of the data. Based on this, the present invention calculates the data category sensitive parameters through the data packet arrival value and the decoding characterization value, and then locates the sensitive data category, performs targeted backup and synchronous transmission, and improves the accuracy and efficiency of data transmission.
[0121] Specifically, the process of locating the sensitive data category, selecting a protection processing data packet based on the location result, and performing protection transmission on the data packet includes:
[0122] If the data category sensitive parameter is greater than a preset sensitive parameter threshold, the data category is determined to be a sensitive data category;
[0123] Selecting data packets corresponding to sensitive data categories from untransmitted data packets as protection processing data packets;
[0124] Performing protection transmission on the protection-processed data packet;
[0125] The protection transmission is to construct the same data packet based on the protection processing data packet for synchronous transmission.
[0126] Specifically, the preset sensitive parameter threshold represents the minimum decodable number of data packets during data packet transmission, so the preset sensitive parameter threshold is set to be selected within the interval [0.32, 0.45].
[0127] Specifically, by determining the category of sensitive data, a theoretical basis is provided for selecting protection processing data packets. During data transmission, the probability of sensitive data being damaged will increase. In actual situations, most of the time, data damage is determined after the transmission is completed and retransmission is initiated. Especially for data sources that need to transmit data in real time, this will not only lead to low transmission efficiency, but also affect the efficiency of data utilization. Based on this, the present invention determines the sensitive data category and selects protection processing data packets for targeted protection to improve the accuracy and efficiency of data transmission.
[0128] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0129] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A database management method based on data analysis, characterized in that: include: Obtaining several data packets corresponding to several historical transmission processes of the data source, extracting the data category of each data packet, constructing a high-frequency data category sequence, recording the number of occurrences of the same high-frequency data category sequence to analyze the historical transmission pattern of the data source, and verifying the identity of the data source; In response to a data source requiring data transmission, a pre-transmission experiment is conducted based on an identity verification result to randomly extract a predetermined proportion of data packets of different data categories transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine an arrival time deviation, screen adjacent data packet groups, and determine a time-domain dependency characterization parameter based on decoding results of the adjacent data packet groups; Decoding the data packets of the pre-transmission experiment to determine the number of undecodable data packets to determine the combination dependency characterization parameter; Calculating a data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter, and classifying the data packet dependency tendency; In response to the division result of the data packet dependency tendency, controlling the data transmission of the data source, including, monitoring the pre-transmission experiment, analyzing the packet arrival value and decoding representation value of the data packets corresponding to each data category based on the results of the pre-transmission experiment of the corresponding data source, calculating the data category sensitive parameter to locate the sensitive data category, selecting a protection processing data packet based on the location result, and performing protection transmission on the data packet; The process of analyzing the historical transmission patterns of the data source and verifying the identity of the data source includes: Record the number of similar transmission behaviors of the data source during several historical transmissions; Calculate the ratio of the number of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio; If the similar behavior transmission ratio is greater than a predetermined similar behavior transmission ratio threshold, it is determined that the data sources are identical; The similar transmission behavior must satisfy the requirement that the high-frequency data category sequence corresponding to each data packet transmitted during the historical transmission process is the same as the high-frequency data category sequence corresponding to each data packet during any other historical transmission process; The process of determining the combination dependency characterization parameter includes: determining the difference between the number of undecodable packets and the number of corrupted packets; The ratio of the difference to the number of undecodable data packets is determined as a combination dependency characterization parameter.
2. The database management method based on data analysis according to claim 1, characterized in that: Based on the identity verification result, a pre-transmission experiment is performed, wherein: If the data sources are identical, a preliminary transmission experiment is required to randomly select a predetermined proportion of different data category data packets transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time domain dependence characterization parameters based on the decoding results of the adjacent data packet groups; If the data sources are not identical, no pre-transmission experiment is required.
3. The database management method based on data analysis according to claim 1, characterized in that: The process of screening adjacent data packet groups and determining time domain dependency characterization parameters based on decoding results of the adjacent data packet groups includes: Determine the arrival time deviation corresponding to each data packet; If there is a data packet that meets the deviation condition, combining the data packet and the adjacent data packet into an adjacent data packet group; determining whether each data packet in the adjacent data packet group can be completely decoded to determine the time-domain dependent data packet group; determining a ratio of the number of time-domain dependent data packet groups to the total number of data packets as a time-domain dependence characterization parameter; The deviation condition is that the deviation time corresponding to the data packet is greater than a predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.
4. The database management method based on data analysis according to claim 1, characterized in that: The process of calculating the data packet dependency characteristic value based on the time domain dependency characterization parameter and the combined dependency characterization parameter includes: Determine the ratio of the time-domain dependency characterization parameter to the benchmark time-domain dependency characterization parameter as the time-domain dependency impact factor; Determine the ratio of the combination dependency characterization parameter to the benchmark combination dependency characterization parameter as the combination dependency impact factor; Determine the weighted sum of the time domain dependency impact factor and the combined dependency impact factor as a data packet dependency characteristic value.
5. The database management method based on data analysis according to claim 1, characterized in that: The data transmission of the data source is controlled in response to the division result of the data packet dependency tendency, wherein If the data packet dependency characteristic value is greater than the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a strong dependency tendency, the pre-transmission experiment is monitored, and based on the result of the pre-transmission experiment of the corresponding data source, the data packet arrival value and the decoding representation value of the data packet corresponding to each data category are analyzed, and the data category sensitive parameter is calculated to locate the sensitive data category, and a protection processing data packet is selected based on the positioning result, and the data packet is protected and transmitted; If the data packet dependency characteristic value is less than or equal to the data packet dependency characteristic value threshold, the data packet dependency tendency is classified as a weak dependency tendency.
6. The database management method based on data analysis according to claim 1, characterized in that: The process of analyzing the data packet arrival value and decoding representation value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the corresponding data source includes: Determine the number of decodable data packets corresponding to each data category as a data packet arrival value; The ratio of the number of decodable data packets to the total number of data packets of the same data category is determined as a decoding characterization value.
7. The database management method based on data analysis according to claim 1, characterized in that: The process of calculating the data category sensitive parameters includes: Determine the ratio of the baseline data packet arrival value to the data packet arrival value corresponding to the data category as the arrival value impact factor; Determine the ratio of the benchmark decoding representation value corresponding to the data category to the decoding representation value as the decoding representation value influencing factor; A weighted sum of the arrival value impact factor and the decoding representation value impact factor is determined as the data category sensitive parameter.
8. The database management method based on data analysis according to claim 1, characterized in that: The process of locating the sensitive data category, selecting a protection processing data packet based on the locating result, and performing protection transmission on the data packet includes: If the data category sensitive parameter is greater than a preset sensitive parameter threshold, the data category is determined to be a sensitive data category; Selecting data packets corresponding to sensitive data categories from untransmitted data packets as protection processing data packets; Performing protection transmission on the protection-processed data packet; The protection transmission is to construct the same data packet based on the protection processing data packet for synchronous transmission.
Citation Information
Patent Citations
NLP-based big data analysis management system and method
CN114138933A
Data management method and system for multi-modal database
CN119149783A
Transmission analysis method and device for data packet with timestamp
CN119232622A
Data classification management and control method based on artificial intelligence
CN119249493A