Database management method based on data analysis

By performing identity verification and pre-transmission experiments on the data source, screening adjacent data packet groups, determining dependency characterization parameters, dividing data packet dependency tendencies, and performing protection transmission, the problem of database overload in data transmission is solved, and the accuracy and efficiency of data transmission is improved.

CN120296007AActive Publication Date: 2025-07-11北京科杰科技有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510766288.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-11
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

During the data transmission process, all data sources are analyzed and encrypted or backed up indiscriminately, resulting in overloaded databases, reducing the speed and efficiency of data transmission, and affecting the accuracy of data transmission.

Method used

By obtaining data packets in the historical transmission process of data source, building high-frequency data category sequences, performing identity verification, performing pre-transmission experiments, filtering adjacent data packet groups, determining time domain and combination dependency characterization parameters, dividing packet dependency tendencies, monitoring data category sensitive parameters, and performing protection transmission.

Benefits of technology

It improves the accuracy and efficiency of data transmission, reduces the computing load, ensures timely backup and synchronous transmission of sensitive data, and avoids database overload operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296007A_ABST
    Figure CN120296007A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data analysis, in particular to a database management method based on data analysis, and the method comprises the steps: obtaining a plurality of data packets, and determining the data type of each data packet; constructing a data category sequence relative to the data source, and performing identity verification on the data source to perform a pre-transmission experiment; determining arrival time deviation, screening adjacent data packet groups, and determining time domain dependence characterization parameters; recording the number of decodable data packets to determine a combinatorial dependency characterization parameter; calculating a data packet dependency characteristic value, and dividing a data packet dependency tendency; and controlling data transmission of the data source in response to a division result of the dependency tendency of the data packet, including monitoring the pre-transmission experiment, calculating sensitive parameters of a data category to position the sensitive data category, selecting a protection processing data packet based on a positioning result, and performing protection transmission on the data packet. The sensitive data is protected, and the accuracy of data transmission is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and in particular, to a database management method based on data analysis. Background Art

[0002] With the rapid development of information technology, massive amounts of data are being generated in various industries. The traditional manual data management method can no longer cope with such a large amount of data, and it is urgent to use a database management system to efficiently store and manage this data. In particular, as the value of data becomes increasingly prominent, the accuracy and security of data transmission have received more and more attention. Data transmission has become the cornerstone of the operation of various fields. During the data transmission process, by means of artificial intelligence and machine learning technologies, it is possible to comprehensively and finely optimize and manage the data transmission process in real time, ensuring the efficiency, stability, and security of data during transmission, thereby maximizing the efficiency and quality of data transmission, meeting the growing digital needs, and promoting the digital process of various industries to a new level.

[0003] Chinese Patent Publication No.: CN119149783A, discloses a data management method and system for a multimodal database, including: a multimodal data analysis module, a multimodal data storage module, and a multimodal data update module. This invention extracts deep features by using a multimodal data feature analysis model, improving the accuracy and depth of data recognition; optimizes the data storage structure and improves the query efficiency through feature matching technology for database storage allocation; introduces a feature adaptive threshold factor Z and an improved CNN model based on a knowledge graph to enhance the generalization ability and adaptability of the model; realizes dynamic update of database partitions through a multimodal database management update model, responds to changes in user operations, maintains the timeliness and relevance of the database, and enhances the user experience.

[0004] Chinese Patent Publication No.: CN114138933A, discloses a big data analysis management system and method based on NLP, including an analysis management system, a cloud server, and a display unit. The analysis management system is connected to the display unit. The analysis management system includes an exclusive database, a data collection module, a data query module, a data processing module, a graphics processing module, an interactive processing module, and a search engine simulation training module. This invention belongs to the technical field of big data analysis management. Specifically, it provides a big data analysis management system and method based on NLP that can simultaneously meet real-time data display, real-time variable dimension data display, full-text retrieval based on keywords, key phrases, and instructions, and information data query for selecting data presentation indicators. Through massive data retrieval and a large amount of data rendering, a good experience effect is achieved. At the same time, the exclusive corpus can be updated to continuously improve the accuracy of data collection and reduce the dependence on human errors.

[0005] However, the following problems still exist in the prior art: In actual situations, when data is transmitted, most of the data sources are analyzed without discrimination, and data sources that may have risks such as loss or delay are encrypted or backed up without discrimination. This will cause the load of the database to exceed its own bearing capacity, resulting in the database operating overloaded, thereby reducing the speed and efficiency of data transmission and affecting the accuracy of data transmission. Summary of the Invention

[0006] For this reason, the present invention provides a database management method based on data analysis to solve the problem that in actual situations, when data is transmitted, most of the data sources are analyzed without discrimination, and data sources that may have risks such as loss or delay are encrypted or backed up without discrimination. This will cause the load of the database to exceed its own bearing capacity, resulting in the database operating overloaded, thereby reducing the speed and efficiency of data transmission and affecting the accuracy of data transmission.

[0007] To achieve the above object, the present invention provides a database management method based on data analysis, which includes: Obtain a number of data packets corresponding to a number of historical transmission processes of the data source, extract the data categories of each data packet, construct a high-frequency data category sequence, record the occurrence times of the same high-frequency data category sequence to analyze the historical transmission law of the data source, and perform identity verification on the data source; In response to the need for the data source to transmit data, based on the identity verification result, perform a pre-transmission experiment to randomly extract a predetermined proportion of data packets of different data categories from the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen the adjacent data packet groups, and determine the time-domain dependence characterization parameter based on the decoding result of the adjacent data packet groups; Decode the data packets for which the pre-transmission experiment is performed, determine the number of undecodable data packets to determine the combined dependence characterization parameter; Calculate the data packet dependence eigenvalue based on the time-domain dependence characterization parameter and the combined dependence characterization parameter, and divide the data packet dependence tendency; In response to the division result of the data packet dependence tendency, control the data transmission of the data source, including, Monitor the pre-transmission experiment, analyze the data packet arrival value and the decoding characterization value of each data category corresponding to the data source based on the result of the pre-transmission experiment of the data source, calculate the data category sensitivity parameter to locate the sensitive data category, select the protected data packet based on the positioning result, and perform protected transmission on the data packet.

[0008] Further, the process of analyzing the historical transmission law of the data source and performing identity verification on the data source includes, Record the number of times of similar transmission behaviors of the data source during several historical transmissions; Calculate the ratio of the number of times of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio; If the similar behavior transmission ratio is greater than a predetermined similar behavior transmission ratio threshold, it is determined that the data sources have identity; Among them, the similar transmission behaviors need to satisfy that the high-frequency data category sequences corresponding to each data packet transmitted during the historical transmission process are the same as the high-frequency data category sequences corresponding to each data packet during any other historical transmission process.

[0009] Further, based on the identity verification result, a pre-transmission experiment is carried out, where If the data sources have identity, a pre-transmission experiment needs to be carried out. Randomly select data packets of a predetermined proportion of different data categories transmitted by the data source, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time domain dependence characterization parameter based on the decoding results of the adjacent data packet groups; If the data sources do not have identity, no pre-transmission experiment is required.

[0010] Further, the process of screening adjacent data packet groups and determining the time domain dependence characterization parameter based on the decoding results of the adjacent data packet groups includes Determine the arrival time deviation corresponding to each data packet; If there is a data packet that meets the deviation condition, combine the data packet with adjacent data packets into an adjacent data packet group; Determine whether each data packet in the adjacent data packet group can be completely decoded to determine the time domain dependence data packet group; Determine the ratio of the number of time domain dependence data packet groups to the total number of data packets as the time domain dependence characterization parameter; Among them, the deviation condition is that the deviation time corresponding to the data packet is greater than the predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.

[0011] Further, the process of determining the combined dependence characterization parameter includes Determine the difference between the number of undecodable data packets and the number of damaged data packets; Determine the ratio of the difference to the number of undecodable data packets as the combined dependence characterization parameter.

[0012] Further, the process of calculating the data packet dependence eigenvalue based on the time domain dependence characterization parameter and the combined dependence characterization parameter includes Determine the ratio of the time domain dependence characterization parameter to the reference time domain dependence characterization parameter as the time domain dependence influence factor; Determine the ratio of the combined dependency characterization parameter to the benchmark combined dependency characterization parameter as the combined dependency impact factor; Determine the weighted sum value of the time-domain dependency impact factor and the combined dependency impact factor as the data packet dependency eigenvalue.

[0013] Further, in response to the division result of the data packet dependency tendency, control the data transmission of the data source, where If the data packet dependency eigenvalue is greater than the data packet dependency eigenvalue threshold, divide the data packet dependency tendency into a strong dependency tendency, monitor the pre-transmission experiment, analyze the data packet arrival value and the decoded characterization value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source, calculate the data category sensitivity parameter to locate the sensitive data category, select the protected processing data packet based on the location result, and perform protected transmission on the data packet; If the data packet dependency eigenvalue is less than or equal to the data packet dependency eigenvalue threshold, divide the data packet dependency tendency into a weak dependency tendency.

[0014] Further, the process of analyzing the data packet arrival value and the decoded characterization value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source includes Determine the number of decodable data packets corresponding to each data category as the data packet arrival value; Determine the ratio of the number of decodable data packets to the total number of data packets in the same data category as the decoded characterization value.

[0015] Further, the process of calculating the data category sensitivity parameter includes Determine the ratio of the benchmark data packet arrival value corresponding to the data category to the data packet arrival value as the arrival value impact factor; Determine the ratio of the benchmark decoded characterization value corresponding to the data category to the decoded characterization value as the decoded characterization value impact factor; Determine the weighted sum value of the arrival value impact factor and the decoded characterization value impact factor as the data category sensitivity parameter.

[0016] Further, the process of locating the sensitive data category, selecting the protected processing data packet based on the location result, and performing protected transmission on the data packet includes If the data category sensitivity parameter is greater than the preset sensitivity parameter threshold, determine that the data category is a sensitive data category; Select the data packets corresponding to the sensitive data category from the untransmitted data packets as the protected processing data packets; Perform protected transmission on the protected processing data packets; Wherein, the protected transmission is to construct the same data packets based on the protected processing data packets for synchronous transmission.

[0017] Compared with the prior art, the present invention obtains a number of data packets, extracts the data categories of each of the data packets, constructs a high-frequency data category sequence, performs identity verification on the data source to conduct a pre-transmission experiment; determines the arrival time deviation, screens adjacent data packet groups, and determines the time-domain dependence characterization parameters; determines the number of undecodable data packets to determine the combined dependence characterization parameters; calculates the data packet dependence eigenvalue, divides the data packet dependence tendency; in response to the division result of the data packet dependence tendency, controls the data transmission of the data source, monitors the pre-transmission experiment, calculates the data category sensitivity parameter to locate the sensitive data category, selects the protected processing data packet based on the positioning result, and performs protected transmission on the data packet. The present invention analyzes the data packet, locates the sensitive data category, protects the sensitive data, and improves the accuracy of data transmission.

[0018] In particular, by performing identity verification on the data source, it provides a data basis for determining whether to conduct a pre-transmission experiment. In actual situations, most often, the data transmission situation of the data source is judged by obtaining the data during the data transmission operation process. However, for a large amount of data, the data during the data transmission operation process is numerous. Directly analyzing it will waste a large amount of computing power, increase the calculation load, and even lead to analysis errors. Based on this, the present invention first performs identity verification on the data source and conducts a pre-transmission experiment on the data source with identity, providing a classification theoretical basis for subsequent analysis of the dependence degree of data packets, and improving the accuracy and transmission efficiency of data transmission.

[0019] In particular, by determining the time-domain dependence characterization parameters, it provides a data basis for calculating the data packet dependence eigenvalue. When data is transmitted, if there is a problem of slow transmission during the transmission process, the transmission process of the data packet will change, resulting in the arrival time of the data packet being inconsistent with the specified time. Adjacent data packet groups with a high degree of dependence may not be successfully decoded due to too large a difference in arrival time, and thus the transmitted data cannot be recognized or started. Based on this, the present invention considers determining the time-domain dependence characterization parameters according to the results of the pre-transmission experiment, providing a classification theoretical basis for subsequent analysis of the dependence degree of data packets, and improving the accuracy and transmission efficiency of data transmission.

[0020] In particular, by determining the combined dependence characterization parameter, a data basis is provided for calculating the data packet dependence eigenvalue. When data is transmitted, if data loss or damage occurs during the transmission process, some data packets may fail to be successfully decoded due to the failure of combined matching with other data packets, resulting in the transmitted data being unrecognizable or unable to be started. Based on this, the difference between the number of undecodable data packets and the number of damaged data packets is calculated to obtain the possible potential situation where decoding fails because combination with other data packets is required for decoding, and then the combined dependence characterization parameter is calculated, providing a classification theoretical basis for subsequent analysis of the dependence degree of data packets, and improving the accuracy and transmission efficiency of data transmission.

[0021] In particular, by calculating the data category sensitivity parameter to locate sensitive data categories, a data basis is provided for the data transmission process. In actual situations, most data is backed up in its entirety to ensure the safe and accurate arrival of data. However, for a large amount of data, backing up all data not only wastes space but also causes problems such as low data transmission efficiency. Especially for data that requires high timeliness, if data damage is found after transmission and retransmission is carried out, the data will not be provided in a timely manner, affecting the use of the data. Based on this, the present invention calculates the data category sensitivity parameter through the data packet arrival value and the decoding characterization value, and then locates the sensitive data categories for targeted backup and synchronous transmission, improving the accuracy and transmission efficiency of data transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of the steps of the database management method based on data analysis according to an embodiment of the invention; Figure 2 Logic block diagram of the pre - transmission experiment based on the identity verification result according to an embodiment of the invention; Figure 3 Logic block diagram for controlling the data transmission of the data source in response to the division result of the data packet dependence tendency according to an embodiment of the invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the present invention.

[0024] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0025] Please refer to Figure 1 , Figure 1Schematic diagram of the method steps of the database management method based on data analysis according to the invention embodiment. The database management method based on data analysis according to the invention includes: Step S1: Obtain a number of data packets corresponding to several historical transmission processes of the data source, extract the data categories of each data packet, construct a high-frequency data category sequence, record the occurrence times of the same high-frequency data category sequence to analyze the historical transmission law of the data source, and perform identity verification on the data source; Step S2: In response to the need for data transmission of the data source, based on the identity verification result, conduct a pre-transmission experiment, randomly select data packets of different data categories accounting for a predetermined proportion of the data source for transmission, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time-domain dependence characterization parameter based on the decoding result of the adjacent data packet groups; Step S3: Decode the data packets in the pre-transmission experiment to determine the number of undecodable data packets, so as to determine the combined dependence characterization parameter; Calculate the data packet dependence eigenvalue based on the time-domain dependence characterization parameter and the combined dependence characterization parameter, and divide the data packet dependence tendency; Step S4: In response to the division result of the data packet dependence tendency, control the data transmission of the data source, including, Monitor the pre-transmission experiment, analyze the data packet arrival value and the decoding characterization value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source, calculate the data category sensitivity parameter to locate the sensitive data category, select the protected processing data packets based on the location result, and perform protected transmission on the data packets.

[0026] Specifically, no specific limitation is imposed on the specific data category. For example, in implementation, the data category includes but is not limited to application layer data, transport layer data, and network layer data. It can be understood that the purpose of classifying data is to determine the sensitive data category and then perform protected transmission. Therefore, those skilled in the art can also divide the data category according to other bases as long as it is reasonable, and this will not be elaborated here.

[0027] Specifically, the arrival time deviation represents the difference between the arrival time difference between a data packet and its adjacent data packet and the specified arrival time difference. In implementation, by obtaining a number of high-frequency data category sequences of successful transmissions, recording the historical arrival time differences of the corresponding data packets, and determining the average value of each historical arrival time difference as the specified arrival time difference. Those skilled in the art can also use other methods to determine the specified arrival time difference, and this will not be elaborated here.

[0028] Specifically, the process of analyzing the historical transmission law of the data source and performing identity verification on the data source includes, Record the number of times of similar transmission behaviors during several historical transmissions of the data source; Calculate the ratio of the number of times of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio; If the similar behavior transmission ratio is greater than the predetermined similar behavior transmission ratio threshold, it is determined that the data sources have identity; Among them, the similar transmission behaviors need to satisfy that the high-frequency data category sequences corresponding to each data packet transmitted during the historical transmission process are the same as the high-frequency data category sequences corresponding to each data packet in any other historical transmission process.

[0029] In implementation, it is necessary to determine the data packet arrangement, assign data packet category labels to the data packets based on the data categories of the data packets, and then obtain the arrangement of the labels, and determine the arrangement of the labels as the total data category sequence.

[0030] Determine the segments in the total data category sequence as data category sequences. Each data category sequence contains at least 3 labels. Furthermore, determine the occurrence probability of each data category sequence in the total data category sequence. If there is a data category sequence whose occurrence probability is greater than the predetermined probability threshold, it is determined that the data category sequence is a high-frequency data category sequence.

[0031] In implementation, the probability threshold is determined in advance. Statistically analyze the occurrence probabilities of each data category sequence corresponding to several historical transmission processes in advance, solve the average occurrence probability, and set the probability threshold to 1.15 times the average occurrence probability.

[0032] Specifically, the predetermined similar behavior transmission ratio threshold characterizes the minimum occurrence frequency of the corresponding similar transmission behaviors in the case where the data sources have identity. Therefore, the predetermined similar behavior transmission ratio threshold is set to be selected within the interval [0.3, 0.5].

[0033] Specifically, by performing identity verification on the data sources, it provides a data basis for determining whether to conduct a pre-transmission experiment. In actual situations, most of the time, the data transmission situation of the data sources is judged by obtaining the data during the data transmission operation process. However, for a large amount of data, there is a large amount of data during the data transmission operation process. Directly analyzing it will waste a large amount of computing power, increase the computing load, and even lead to analysis errors. Based on this, the present invention first performs identity verification on the data sources, conducts a pre-transmission experiment on the data sources with identity, provides a classification theoretical basis for analyzing the dependency degree of data packets subsequently, and improves the accuracy and transmission efficiency of data transmission.

[0034] Please refer to Figure 2 , Figure 2 which is the logic block diagram of the pre-transmission experiment based on the identity verification result of the invention embodiment. Specifically, based on the identity verification result, a pre-transmission experiment is conducted, where If the data sources have identity, it is necessary to conduct a pre - transmission experiment. Randomly extract data packets of different data categories with a predetermined proportion from the data sources, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time - domain dependence characterization parameters based on the decoding results of the adjacent data packet groups. If the data sources do not have identity, there is no need to conduct a pre - transmission experiment.

[0035] In implementation, set the predetermined proportion to 10%. For example, there are ten data packets of application - layer data categories and ten data packets of transport - layer data categories. According to the predetermined proportion, one data packet of application - layer data category and one data packet of transport - layer data category are extracted. It can be understood that the number of data packets extracted during the extraction is rounded up based on the predetermined proportion. Of course, those skilled in the art can increase or decrease it according to the number of data packets as long as it is reasonable, and this will not be elaborated here.

[0036] Specifically, the process of screening adjacent data packet groups and determining the time - domain dependence characterization parameters based on the decoding results of the adjacent data packet groups includes: Determine the arrival time deviation corresponding to each data packet; If there is a data packet that meets the deviation condition, combine the data packet with adjacent data packets into an adjacent data packet group; Determine whether each data packet in the adjacent data packet group can be completely decoded to determine the time - domain dependence data packet group; Determine the ratio of the number of time - domain dependence data packet groups to the total number of data packets as the time - domain dependence characterization parameter; Among them, the deviation condition is that the deviation time corresponding to the data packet is greater than the predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.

[0037] Specifically, determine the adjacent data packet groups that cannot be completely decoded as the time - domain dependence data packet groups. It can be understood that if each data packet in the adjacent data packet group can be completely decoded, it means that the adjacent data packet group is not affected by the time domain and is considered to have no time - domain dependence.

[0038] It can be understood that the composition of the adjacent data packet group includes two or three data packets. For the first data packet, it forms an adjacent data packet group with the adjacent subsequent data packet. For the last data packet, it forms an adjacent data packet group with the adjacent previous data packet.

[0039] Specifically, there is no limitation on the specific decoding method. Protocol analysis tools such as Wireshark and tcpdump can be used for decoding. Enter the corresponding commands in the command line of the tool to capture data packets and output them to the terminal. Programming language libraries can also be used for decoding. The sniff function in Scapy is used to capture data packets, and then the provided parsing functions are used to parse the data packets. The header information and payload content of the data packets can be parsed layer by layer. Those skilled in the art can select the corresponding decoding tool from the existing data according to their needs, which will not be elaborated here.

[0040] Specifically, the predetermined time deviation is determined in advance. The arrival time deviations of each data packet in several historical transmission processes are recorded in advance, and the average value of the arrival time deviations corresponding to each data packet in each historical transmission process is determined as the predetermined time deviation.

[0041] Specifically, determining the time-domain dependence characterization parameter provides a data basis for calculating the data packet dependence eigenvalue. When data transmission is carried out, if there is a problem of slow transmission during the transmission process, the transmission process of the data packet will change, resulting in the arrival time of the data packet being inconsistent with the specified time. The adjacent data packet group with a high degree of dependence may not be successfully decoded due to the too large difference in arrival time, and then the transmitted data cannot be recognized or started. Based on this, the present invention considers determining the time-domain dependence characterization parameter according to the results of pre-transmission experiments, providing a classification theoretical basis for subsequent analysis of the dependence degree of data packets, and improving the accuracy and transmission efficiency of data transmission.

[0042] Specifically, the process of determining the combined dependence characterization parameter includes determining the difference between the number of undecodable data packets and the number of damaged data packets; determining the ratio of the difference to the number of undecodable data packets as the combined dependence characterization parameter.

[0043] Specifically, determining the combined dependence characterization parameter provides a data basis for calculating the data packet dependence eigenvalue. When data transmission is carried out, if there is data loss or damage during the transmission process, some data packets may not be successfully decoded due to the failure of combination matching with other data packets, and then the transmitted data cannot be recognized or started. Based on this, calculate the difference between the number of undecodable data packets and the number of damaged data packets, obtain the possible situations where decoding fails because combination with other data packets is required for decoding, and then calculate the combined dependence characterization parameter, providing a classification theoretical basis for subsequent analysis of the dependence degree of data packets, and improving the accuracy and transmission efficiency of data transmission.

[0044] Specifically, the process of calculating the data packet dependence eigenvalue based on the time-domain dependence characterization parameter and the combined dependence characterization parameter includes Determine the ratio of the time-domain dependence characterization parameter to the reference time-domain dependence characterization parameter as the time-domain dependence influence factor; Determine the ratio of the combined dependence characterization parameter to the reference combined dependence characterization parameter as the combined dependence influence factor; Determine the weighted sum value of the time-domain dependence influence factor and the combined dependence influence factor as the packet dependence eigenvalue.

[0045] Specifically, the reference time-domain dependence characterization parameter is pre-calculated data. Obtain the corresponding time-domain dependence characterization parameters during several successful data transmission processes in advance, and determine that the reference time-domain dependence characterization parameter is selected within 1.08 to 1.21 times the average value of each time-domain dependence characterization parameter.

[0046] Specifically, the reference combined dependence characterization parameter is pre-calculated data. Obtain the corresponding combined dependence characterization parameters during several successful data transmission processes in advance, and determine that the reference combined dependence characterization parameter is selected within 1.11 to 1.23 times the average value of each combined dependence characterization parameter.

[0047] Specifically, the weight coefficient of the time-domain dependence influence factor and the combined dependence influence factor is 1, the weight coefficient of the time-domain dependence influence factor is 0.43, and the weight coefficient of the combined dependence influence factor is 0.57.

[0048] Please participate Figure 3 , Figure 3 It is a logic block diagram for controlling the data transmission of the data source in response to the classification result of the packet dependence tendency in the invention embodiment. Specifically, in response to the classification result of the packet dependence tendency, control the data transmission of the data source, where, If the packet dependence eigenvalue is greater than the packet dependence eigenvalue threshold, classify the packet dependence tendency as a strong dependence tendency, monitor the pre-transmission experiment, analyze the packet arrival value and the decoding characterization value of the packets corresponding to each data category based on the result of the pre-transmission experiment of the data source, calculate the data category sensitivity parameter to locate the sensitive data category, select the protected processing packet based on the location result, and perform protected transmission on the packet; If the packet dependence eigenvalue is less than or equal to the packet dependence eigenvalue threshold, classify the packet dependence tendency as a weak dependence tendency.

[0049] Specifically, the packet dependence eigenvalue threshold characterizes the highest dependence degree of the packet in the decodable state, so it is determined that the dependence eigenvalue threshold is selected within the interval [0.65, 0.83].

[0050] Specifically, the process of analyzing the packet arrival value and the decoding characterization value of the packets corresponding to each data category based on the result of the pre-transmission experiment of the data source includes, Determine that the number of decodable data packets corresponding to each data category is the data packet arrival value; Determine that the ratio of the number of decodable data packets to the total number of data packets of the same data category is the decoding characterization value.

[0051] Specifically, the process of calculating the data category sensitivity parameter includes, Determine that the ratio of the reference data packet arrival value corresponding to the data category to the data packet arrival value is the arrival value influence factor; Determine that the ratio of the reference decoding characterization value corresponding to the data category to the decoding characterization value is the decoding characterization value influence factor; Determine that the weighted sum value of the arrival value influence factor and the decoding characterization value influence factor is the data category sensitivity parameter.

[0052] Specifically, the reference data packet arrival value is calculated from historical data. Obtain in advance the number of data packets of the same data category corresponding to several successful transmissions of the data source, and determine the average value of several transmission data packet numbers as the reference data packet arrival value.

[0053] Specifically, the reference decoding characterization value is calculated from historical data. Obtain in advance the decoding characterization values of several successful transmissions, and determine the average value of several decoding characterization values as the reference decoding characterization value.

[0054] Specifically, the sum of the weight coefficients of the arrival value influence factor and the decoding characterization value influence factor is 1. The weight coefficient of the arrival value influence factor is 0.51, and the weight coefficient of the decoding characterization value influence factor is 0.49.

[0055] Specifically, by calculating the data category sensitivity parameter to locate the sensitive data category, it provides a data basis for the data transmission process. In actual situations, most data is backed up to ensure the safe and accurate arrival of data. However, for massive data, backing up all data will not only waste space but also cause problems such as low data transmission efficiency. Especially for data that requires high timeliness, if data damage is found after transmission and retransmission is carried out, it will make the data provided untimely and affect the data usage situation. Based on this, the present invention calculates the data category sensitivity parameter through the data packet arrival value and the decoding characterization value, and then locates the sensitive data category for targeted backup and synchronous transmission, improving the accuracy and transmission efficiency of data transmission.

[0056] Specifically, to locate the sensitive data category and select the protected processing data packets based on the positioning result, the process of protecting the transmission of the data packets includes, If the data category sensitivity parameter is greater than the preset sensitive parameter threshold, then determine that the data category is a sensitive data category; Select the data packets corresponding to the sensitive data categories in the untransmitted data packets as the protected processing data packets; Perform protected transmission on the protected processing data packets; Among them, the protected transmission is to construct the same data packets based on the protected processing data packets for synchronous transmission.

[0057] Specifically, the preset sensitive parameter threshold characterizes the minimum decodable number of data packets during data packet transmission. Therefore, the preset sensitive parameter threshold is set to be selected within the range [0.32, 0.45].

[0058] Specifically, by determining the sensitive data categories, a theoretical basis is provided for the selected protected processing data packets. When data is transmitted, the probability of damage to sensitive data will increase. In actual situations, data damage is mostly determined after the transmission is completed, and then retransmission is initiated. Especially for data sources that require real-time data transmission, this will not only result in low transmission efficiency but also affect the data usage efficiency. Based on this, the present invention determines the sensitive data categories and selects the protected processing data packets for targeted protection to improve the accuracy and transmission efficiency of data transmission.

[0059] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0060] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A database management method based on data analysis, characterized in that, Including: Obtain a number of data packets corresponding to several historical transmission processes of the data source, extract the data categories of each data packet, construct a high-frequency data category sequence, record the occurrence times of the same high-frequency data category sequence to analyze the historical transmission law of the data source, and perform identity verification on the data source; In response to the need for data transmission by the data source, based on the identity verification result, conduct a pre-transmission experiment to randomly select data packets of different data categories in a predetermined proportion for the data source to transmit, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time-domain dependence characterization parameter based on the decoding result of the adjacent data packet group; Decode the data packets for which the pre-transmission experiment is conducted to determine the number of undecodable data packets, so as to determine the combined dependence characterization parameter; Calculate the data packet dependence eigenvalue based on the time-domain dependence characterization parameter and the combined dependence characterization parameter, and divide the data packet dependence tendency; In response to the division result of the data packet dependence tendency, control the data transmission of the data source, including Monitor the pre-transmission experiment, analyze the data packet arrival value and the decoding characterization value corresponding to each data category based on the result of the pre-transmission experiment of the data source, calculate the data category sensitivity parameter to locate the sensitive data category, select the protected processing data packet based on the positioning result, and perform protected transmission on the data packet.

2. The database management method based on data analysis according to claim 1, characterized in that, The process of analyzing the historical transmission law of the data source and performing identity verification on the data source includes Record the number of times of similar transmission behaviors of the data source in several historical transmission processes; Calculate the ratio of the number of times of similar transmission behaviors to the total number of transmissions to obtain the similar behavior transmission ratio; If the similar behavior transmission ratio is greater than a predetermined similar behavior transmission ratio threshold, it is determined that the data source has identity; Among them, the similar transmission behavior needs to satisfy that the high-frequency data category sequences corresponding to each data packet transmitted in the historical transmission process are the same as the high-frequency data category sequences corresponding to each data packet in any other historical transmission process.

3. The database management method based on data analysis according to claim 2, characterized in that Based on the identity verification result, conduct a pre-transmission experiment, where If the data source has identity, a pre-transmission experiment needs to be conducted to randomly select data packets of different data categories in a predetermined proportion for the data source to transmit, record the arrival time of each data packet corresponding to the data source to determine the arrival time deviation, screen adjacent data packet groups, and determine the time-domain dependence characterization parameter based on the decoding result of the adjacent data packet group; If the data source does not have identity, there is no need to conduct a pre-transmission experiment.

4. The database management method based on data analysis according to claim 1, wherein The process of screening the adjacent data packet groups and determining the time-domain dependence characterization parameter based on the decoding result of the adjacent data packet groups includes Determine the arrival time deviation corresponding to each data packet; If there is a data packet that meets the deviation condition, combine the data packet with the adjacent data packets into an adjacent data packet group; Determine whether each data packet in the adjacent data packet group can be completely decoded to determine the time-domain dependence data packet group; Determine the ratio of the number of time-domain dependence data packet groups to the total number of data packets as the time-domain dependence characterization parameter; Among them, the deviation condition is that the deviation time corresponding to the data packet is greater than the predetermined time deviation, and complete decoding means that each data packet in the adjacent data packet group can be completely decoded.

5. The database management method based on data analysis according to claim 1, wherein The process of determining the combined dependence characterization parameter includes determining the difference between the number of undecodable data packets and the number of damaged data packets; determining the ratio of the difference to the number of undecodable data packets as the combined dependence characterization parameter.

6. The database management method based on data analysis according to claim 1, wherein, The process of calculating the data packet dependence eigenvalue based on the time-domain dependence characterization parameter and the combined dependence characterization parameter includes determining the ratio of the time-domain dependence characterization parameter to the reference time-domain dependence characterization parameter as the time-domain dependence influence factor; determining the ratio of the combined dependence characterization parameter to the reference combined dependence characterization parameter as the combined dependence influence factor; determining the weighted sum value of the time-domain dependence influence factor and the combined dependence influence factor as the data packet dependence eigenvalue.

7. The database management method based on data analysis according to claim 1, characterized in that In response to the classification result of the data packet dependence tendency, controlling the data transmission of the data source, where if the data packet dependence eigenvalue is greater than the data packet dependence eigenvalue threshold, classifying the data packet dependence tendency as a strong dependence tendency, monitoring the pre-transmission experiment, analyzing the data packet arrival value and the decoding characterization value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source, calculating the data category sensitivity parameter to locate the sensitive data category, selecting the protected processing data packet based on the positioning result, and performing protected transmission on the data packet; if the data packet dependence eigenvalue is less than or equal to the data packet dependence eigenvalue threshold, classifying the data packet dependence tendency as a weak dependence tendency.

8. The database management method based on data analysis according to claim 1, characterized in that, The process of analyzing the data packet arrival value and the decoding characterization value of the data packets corresponding to each data category based on the result of the pre-transmission experiment of the data source includes determining the number of decodable data packets corresponding to each data category as the data packet arrival value; determining the ratio of the number of decodable data packets to the total number of data packets in the same data category as the decoding characterization value.

9. The database management method based on data analysis according to claim 1, wherein The process of calculating the data category sensitivity parameter includes determining the ratio of the reference data packet arrival value corresponding to the data category to the data packet arrival value as the arrival value influence factor; determining the ratio of the reference decoding characterization value corresponding to the data category to the decoding characterization value as the decoding characterization value influence factor; determining the weighted sum value of the arrival value influence factor and the decoding characterization value influence factor as the data category sensitivity parameter.

10. The database management method based on data analysis according to claim 1, characterized in that, The process of locating the sensitive data category, selecting the protected processing data packet based on the positioning result, and performing protected transmission on the data packet includes if the data category sensitivity parameter is greater than the preset sensitivity parameter threshold, determining that the data category is a sensitive data category; selecting the data packet corresponding to the sensitive data category from the untransmitted data packets as the protected processing data packet; performing protected transmission on the protected processing data packet; wherein, the protected transmission is to construct the same data packet for synchronous transmission with the protected processing data packet as the reference.

Citation Information

Patent Citations

  • Transmission analysis method and device for data packet with timestamp

    CN119232622A

  • Data classification management and control method based on artificial intelligence

    CN119249493A

  • Method and system for identifying and analyzing power sensitive data

    CN119484017A

  • Data transmission protection method

    CN119922007A

  • Information data secure transmission method and system

    CN120017388A