A method for cleaning communication data based on a data link
By generating data quality coefficients and cleaning priority values, and combining them with the data storage model to select the best cleaning scheme, the problem of low cleaning efficiency caused by the failure to consider the data writing status in existing technologies is solved, thus achieving efficient and orderly data cleaning and quality improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI UNIV
- Filing Date
- 2023-11-24
- Publication Date
- 2026-05-29
AI Technical Summary
Existing communication data cleaning methods fail to effectively consider the data writing status when processing data within the storage area. This results in low cleaning efficiency when the amount of data written is large and unevenly distributed, affecting the reading efficiency and data quality of the storage area.
Data quality coefficients are generated by identifying data quality characteristics, and cleaning priority values are generated by combining data write and read states. The trained data storage model is used for simulation analysis to select the best cleaning scheme and predict and remind data quality, ensuring that the cleaning process is carried out in an orderly manner.
It improves the efficiency and quality of data cleaning, reduces the interference of the cleaning process on the storage area, extends the working life of sub-areas, and predicts data quality changes in a timely manner to ensure the normal use of the storage area.
Smart Images

Figure CN117472894B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method for cleaning communication data based on a data link. Background Technology
[0002] A data link is a communication link that provides point-to-point data transmission between any two adjacent nodes through one or more communication channels between two or more Data Transfer Equipment (DTEs). A data link is also called a link, and it corresponds to one or more physical links. In the data link layer, each link must run not only the data link layer protocol but also the physical layer protocol used by that link. When the amount of data contained in the data link is too large or the type is too complex, timely data cleaning is required to improve data transmission and communication efficiency.
[0003] Chinese invention patent CN1115065676B discloses a cross-database data cleaning method, apparatus, computer equipment, and storage medium. The method includes: if it is determined that the target dataset to be cleaned is stored in multiple target databases, then establishing communication connections with each target database respectively; after the communication connections are established, retrieving dependent datasets with dependencies from the target datasets stored in each target database and storing them in a cache; and performing data cleaning operations on the dependent datasets in the cache and the independent datasets stored in each target database.
[0004] The above application caches data with dependencies in multiple databases separately, enabling cross-database data cleaning without backing up data between databases. This achieves the technical effect of not requiring any database to be shut down during the cross-database data cleaning process, greatly reducing operation time and costs.
[0005] However, existing communication data cleaning methods do not take the data writing status into account when cleaning data in various databases, that is, in the storage area. They usually pause or slow down the writing and reading of data in the entire storage area before cleaning the data. This will greatly reduce the current data reading efficiency of the storage area. If the writing and reading of data are not paused, and the current data writing volume is large and the data distribution is not uniform, if the selected cleaning scheme is not targeted enough, the efficiency of data cleaning may be difficult to achieve the expected results.
[0006] Therefore, the present invention provides a method for cleaning communication data based on data links. Summary of the Invention
[0007] (a) Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides a data link-based communication data cleaning method. The method involves identifying and acquiring data quality characteristics of each data group, generating data quality coefficients, and issuing cleaning instructions if the coefficients exceed a quality threshold. It also involves generating cleaning priority values for data processing within each sub-region from a set of data conditions, ranking these priority values, and using the obtained ranking as the cleaning order. Furthermore, it involves identifying and acquiring cleaning characteristics of data within each sub-region, matching several corresponding candidate solutions from a pre-built trial cleaning scheme library, and selecting a recommended solution from these candidate solutions. Finally, it involves continuously acquiring several data quality coefficients, predicting the time required for each data quality coefficient to exceed the corresponding quality threshold, and issuing a reminder instruction if the time is shorter than expected. The method evaluates the busy level of each sub-region and adjusts the working status of each sub-region based on the evaluation results to extend its service life, thereby solving the technical problems mentioned in the background art.
[0009] (II) Technical Solution
[0010] To achieve the above objectives, the present invention is implemented through the following technical solution: a method for cleaning communication data based on a data link, comprising the following: if data is continuously written to the storage area, a data receiving state set is established from the data writing status data, and a data receiving state coefficient Ct(s,n) is generated from the data receiving state set; if the data receiving state coefficient Ct(s,n) exceeds the state threshold, an early warning command is issued to the outside.
[0011] Upon receiving the warning instruction, the data group in the sub-region is preprocessed, the data quality characteristics of each data group are identified and obtained, and the data quality set in the storage area is generated after aggregation. The data quality coefficient Tq(s,q) is generated from the data quality set. If the obtained data quality coefficient Tq(s,q) exceeds the quality threshold, a cleaning instruction is issued.
[0012] Based on the data write and read status of each sub-region, a data condition set is established. This data condition set generates the cleaning priority value Yp(p,s) for data processing in each sub-region. Specifically, the data replacement ratio Qp and read count Qs are linearly normalized, mapping the corresponding data values to the interval [0, 1]. Then, the following formula is used:
[0013]
[0014] Weighting coefficients: 0≤ζ≤1, 0≤ψ≤1, and ζ+ψ=1. Sort the cleaning priority values Yp(p,s) in each sub-region and use the obtained sorting as the cleaning order.
[0015] Identify the cleaning characteristics of the data acquired in each sub-region, and match several corresponding candidate solutions from the pre-built trial cleaning solution library. Use the trained data storage model to simulate and analyze the availability of the candidate solutions, and select the recommended solution from the candidate solutions based on the analysis results.
[0016] After implementing the recommendation scheme, the trained data storage model is used to predict the quality parameters of the data, and a data quality set is established. Several data quality coefficients Tq(s,q) are continuously obtained from the data quality set. The time taken to obtain the data quality coefficients Tq(s,q) exceeding the quality threshold is predicted. If it is shorter than expected, an alert instruction is issued.
[0017] Furthermore, the data reception status is monitored to obtain the data reception status within the reception cycle, the amount of data received in each reception cycle is obtained, and the data reception amount Ss is generated; the storage area is divided into several sub-regions, the write ratio P in each reception cycle is obtained, and then the uniformity Un of the data distribution within the storage area is calculated.
[0018] After continuously acquiring several received values Ss and summarizing the uniformity Un, a data reception state set is generated; the data reception state coefficient Ct(s,n) is generated from the data reception state set; if the data reception state coefficient Ct(s,n) exceeds the state threshold, an early warning command is issued.
[0019] Furthermore, the data reception state coefficients Ct(s,n) are generated from the data reception state set, specifically as follows: After linearly normalizing the received quantity Ss and the uniformity Un, the corresponding data values are mapped to the interval [0, 1], and then the following formula is used:
[0020]
[0021] in, Ss represents the historical average of the received data. i Its current value; Un is the historical average of uniformity. i Its current value; n is a positive integer, i = 1, 2, ..., n, which is the number of detection periods, and the weight coefficients are: 0 ≤ β ≤ 1, 0 ≤ α ≤ 1, and α + β = 1.
[0022] Furthermore, the received data is classified according to its type, resulting in several different data groups. After preprocessing the data within each data group, the data quality characteristics of each data group are identified and obtained, including:
[0023] The data within each group are arranged in order of their acquisition time, and data analysis is performed to obtain the quality parameters of the data within each group, including relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss. After summarizing the above data within each group, a data quality set is established within the storage area.
[0024] Furthermore, the data quality coefficients Tq(s,q) are generated from the data quality set, and the specific method is as follows:
[0025] The relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss within each data group are normalized, and the corresponding data values are mapped to the interval [0, 1]. Then, the data quality value Tp(s,s,s) for each sub-region is generated according to the following formula:
[0026]
[0027] Wherein, the parameter means: n is a positive integer greater than 1, i = 1, 2, ..., n, which is the number of data groups in the sub-region, and the weight coefficient is: 0 ≤ F1 ≤ 1, 0 ≤ F2 ≤ 1, 0 ≤ F3 ≤ 1 and F3 + F2 + F1 = 1. The mean of the relative range. The mean of the skewness coefficients. This represents the mean of the kurtosis coefficients.
[0028] Furthermore, based on data quality values The data quality coefficient Tq(s,q) is generated as follows:
[0029]
[0030] Where i = 1, 2, ..., m, m is the number of subregions, and is a positive integer greater than 1. Q i This represents the median data quality value within the sub-region. The average data quality value of the sub-region is used; if the obtained data quality coefficient Tq(s,q) exceeds the preset quality threshold, a cleaning command is sent to the outside.
[0031] Furthermore, upon receiving the cleaning instruction, within the query period, the data replacement ratio Qp of each data item in each sub-region and the number of times the unreplaced data in the sub-region was read are obtained, and the data condition set is established by summarizing them.
[0032] The data condition set generates a cleaning priority value Yp(p,s) for data processing in each sub-region. After obtaining the cleaning priority value Yp(p,s) for each sub-region, the cleaning priority values Yp(p,s) in each sub-region are sorted, and the obtained sorting is used as the cleaning order.
[0033] Furthermore, the preprocessed data in each sub-region is identified, and the corresponding data features are obtained. The types and quantities of data features in each sub-region are used as cleaning features to obtain several data cleaning schemes. After summarizing, a cleaning scheme library is pre-built. Based on the correspondence between the cleaning features and cleaning schemes in each sub-region, the trained matching model is used to match the corresponding cleaning schemes from the pre-built cleaning scheme library and use the cleaning schemes as candidate schemes.
[0034] Furthermore, an initial model is constructed using a neural convolutional network, trained and tested, and then the trained initial model is output as a data storage model. The usability of one or more candidate solutions is simulated and analyzed using the trained data storage model, and the candidate solution with the best performance is output as the recommended solution. Alternatively, the conditional parameters are adjusted to obtain several adjusted candidate solutions as recommended solutions.
[0035] Furthermore, if the storage area continues to receive and write new data, the trained data storage model is used to predict the storage status of the data. At the end of the prediction period, the corresponding quality parameters are obtained, and the data quality set in the storage area is re-established. Several data quality coefficients Tq(s,q) at the end of the prediction period are continuously obtained from the data quality set. If none of the obtained data quality coefficients Tq(s,q) exceed the current quality threshold, the obtained data quality coefficients Tq(s,q) are arranged in order, and the changing trend of the data quality coefficients Tq(s,q) is predicted based on the smoothing exponential model.
[0036] The time required for the data quality coefficient Tq(s,q) to exceed the quality threshold is taken as the risk time. Based on historical data and the expected management of data quality, a time threshold is set in advance. When the predicted risk time exceeds the time threshold, an external reminder instruction is issued.
[0037] (III) Beneficial Effects
[0038] This invention provides a method for cleaning communication data based on a data link, which has the following beneficial effects:
[0039] 1. Generate a data quality coefficient Tq(s,q) to evaluate the current data quality within the storage area. This helps determine whether data cleaning is necessary to remove invalid and low-quality data, thereby improving the overall data quality within the storage area. By obtaining the data quality coefficient Tq(s,q), the distribution and stability of various data items can be encompassed during data quality analysis. Furthermore, averaging data from each sub-region before comprehensive analysis makes the data quality evaluation process more comprehensive and objective.
[0040] 2. When cleaning data in each sub-area, the cleaning process can be completed in an orderly manner. By cleaning different sub-areas sequentially, the remaining sub-areas can maintain normal storage and writing processes. The degree of mutual interference between different sub-areas can be reduced, and the normal usage status of the storage area can be maintained. The busy level of each sub-area can be evaluated, and the working status of each sub-area can be adjusted based on the evaluation results to extend its working life.
[0041] 3. Upon receiving a cleaning instruction, reduce the time spent preparing a cleaning plan in advance to improve data cleaning efficiency; test and screen the matched cleaning plans, selecting the part with the best cleaning effect to improve the data cleaning quality while improving the cleaning effect; improve the correspondence between the modified plan and the cleaning features by revising the cleaning plan to reduce the time spent modifying the cleaning plan and improve the efficiency of data cleaning.
[0042] 4. Use the trained data storage model to predict the data storage status at each time point in the sub-region, obtain the corresponding data quality coefficient Tq(s,q), predict its changing trend, obtain the corresponding risk time, predict and alert for data quality changes in the sub-region, and facilitate timely early detection and processing during data cleaning. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the data link-based communication data cleaning method of the present invention;
[0044] Figure 2 This is a schematic diagram of the communication data cleaning system based on the data link according to the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] Please see Figure 1 This invention provides a method for cleaning communication data based on a data link, comprising the following:
[0047] Step 1: If data is continuously written to the storage area, a data receiving state set is established based on the data writing status data, and a data receiving state coefficient Ct(s,n) is generated from the data receiving state set. If the data receiving state coefficient Ct(s,n) exceeds the state threshold, an early warning command is issued to the outside.
[0048] Step one includes the following:
[0049] Step 101: If the storage area is still in a state of continuous data reception and writing, set the data reception period, for example, 0.1 hours or 0.3 hours as a reception period, and monitor the data reception status to obtain the data reception status within the reception period. Specifically, this includes the following:
[0050] Obtain the amount of data received in each reception cycle and generate the data reception amount Ss;
[0051] The storage area is divided into several sub-regions, and each sub-region is numbered. The proportion of data written to each sub-region in each receiving cycle is obtained as the total amount of data written, and is marked as the write ratio P. Then, the uniformity of data distribution in the storage area Un is calculated.
[0052] The uniformity Un of data distribution within each sub-region is obtained as follows:
[0053]
[0054] Among them, P i The current value of the write ratio. This represents the average write ratio for each sub-region, where n is the number of sub-regions.
[0055] After several receiving cycles, several received values Ss and uniformity Un are continuously acquired and summarized to generate a data receiving status set.
[0056] Step 102: Generate the data reception state coefficients Ct(s,n) from the data reception state set. The specific method is as follows: After linearly normalizing the received quantity Ss and the uniformity Un, map the corresponding data values to the interval [0, 1], and then follow the following formula:
[0057]
[0058] in, Ss represents the historical average of the received data. i Its current value; Un is the historical average of uniformity. i Its current value;
[0059] n is a positive integer, i = 1, 2, ..., n, which is the number of detection periods. The weight coefficient is: 0 ≤ β ≤ 1, 0 ≤ α ≤ 1, and α + β = 1. The specific value of the weight coefficient can be obtained by software simulation or adjusted by the user according to the actual usage.
[0060] Based on historical data and expectations for data management, state thresholds are pre-set for data writing and receiving. If the data receiving state coefficient Ct(s,n) exceeds the state threshold, it indicates that there is a lot of data being written in the current storage area and the distribution in the various sub-areas of the storage area is not uniform enough. Timely processing is required, and an early warning command needs to be issued to the outside.
[0061] When using it, refer to the content in steps 101 and 102:
[0062] Before cleaning the data in the storage area, it is necessary to determine whether the stored data needs to be cleaned. If the amount of data written in the storage area of the memory is large and the distribution of the written data is very uneven, the data reading efficiency will be relatively low after the data writing is completed. Under this condition, it is necessary to clean and process the data in the storage area to improve the data quality.
[0063] However, existing communication data cleaning methods do not take the data writing status into account when cleaning data in the storage area. They usually pause or slow down the writing and reading of data in the entire storage area before cleaning the data. This will greatly reduce the current data reading efficiency of the storage area. If the writing and reading of data are not paused, and the current data writing volume is large and the data distribution is not uniform, if the selected cleaning scheme is not targeted enough, the efficiency of data cleaning may be difficult to achieve the expected results.
[0064] Step 2: After receiving the warning instruction, preprocess the data groups in the sub-region, identify and obtain the data quality characteristics of each data group, summarize them to generate a data quality set in the storage area, and generate a data quality coefficient Tq(s,q) from the data quality set. If the obtained data quality coefficient Tq(s,q) exceeds the quality threshold, a cleaning instruction is issued.
[0065] Step two includes the following:
[0066] Step 201: After completing a stage of data reception, classify the received data according to data type to obtain several different data groups. After preprocessing the data within each data group, identify and obtain the data quality characteristics within each data group, which specifically includes the following:
[0067] The data in each group are arranged in order of their acquisition time, and data analysis is performed to obtain the quality parameters of the data in each group, including relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss. After summarizing the above data in each group, a data quality set in the storage area is established.
[0068] Step 202: Generate data quality coefficients Tq(s,q) from the data quality set. The specific method is as follows: Normalize the relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss within each data group, map the corresponding data values to the interval [0, 1], and then generate the data quality values Tp(s,s,s) for each sub-region according to the following formula:
[0069]
[0070] Wherein, the parameter means: n is a positive integer greater than 1, i = 1, 2, ..., n, which is the number of data groups in the sub-region, and the weight coefficient is: 0 ≤ F1 ≤ 1, 0 ≤ F2 ≤ 1, 0 ≤ F3 ≤ 1 and F3 + F2 + F1 = 1. The mean of the relative range. The mean of the skewness coefficients. The mean of the kurtosis coefficients;
[0071]
[0072] Where i = 1, 2, ..., m, m is the number of subregions, and is a positive integer greater than 1. Q i This represents the median data quality value within the sub-region. This represents the mean of the data quality values for the sub-region.
[0073] Based on historical data and expectations for data quality management, a quality threshold is preset. If the acquired data quality coefficient Tq(s,q) exceeds the preset quality threshold, it indicates that the data quality in the storage area is relatively low and needs to be processed in a timely manner. A cleaning command is sent to the outside to clean the newly received and written data in the current receiving cycle.
[0074] When using it, refer to steps 201 and 202:
[0075] After dividing the storage area into several sub-regions, the data is preprocessed, the data quality characteristics of each data group are identified, and a data quality coefficient Tq(s,q) is generated. This coefficient is used to evaluate the current data quality within the storage area, so as to determine whether the current data needs to be cleaned to remove invalid and low-quality data, thereby improving the overall data quality within the storage area. By obtaining the data quality coefficient Tq(s,q), the distribution and stability of the current data can be included when performing data quality analysis. Furthermore, the average values are calculated for each sub-region before being combined, making the data quality evaluation process more comprehensive and objective.
[0076] Step 3: Based on the data write and read status of each sub-region, establish a data condition set, generate the cleaning priority value Yp(p,s) for data processing in each sub-region from the data condition set, sort the cleaning priority values Yp(p,s) in each sub-region, and use the obtained sorting as the cleaning order.
[0077] Step three includes the following:
[0078] Step 301: After receiving the cleaning instruction, set the query period so that the length of the query period is the same as the receiving period. Within the query period, obtain the data replacement ratio Qp of each data in each sub-region, that is, the proportion of the amount of new and old data replaced to the total amount of data in the region, and the number of times the unreplaced data in the sub-region is read Qs. After summarizing the data replacement ratio Qp and the number of reads Qs in each sub-region, establish a data condition set.
[0079] Step 302: Generate the cleaning priority value Yp(p,s) for data processing in each sub-region from the data condition set. The specific method is as follows: Perform linear normalization on the data replacement ratio Qp and the number of reads Qs, map the corresponding data values to the interval [0, 1], and then follow the formula below:
[0080]
[0081] Weighting coefficients: 0≤ζ≤1, 0≤ψ≤1, and ζ+ψ=1. The specific values of the weighting coefficients can be obtained by software simulation or adjusted by the user according to actual usage.
[0082] After obtaining the cleaning priority value Yp(p,s) of each sub-region, sort the cleaning priority value Yp(p,s) of each sub-region and use the obtained sorting as the cleaning order. When it is necessary to clean the data in each sub-region, the writing and receiving of data in each sub-region can be paused in an orderly manner according to the cleaning order, and the cleaning can begin.
[0083] When using this method, refer to steps 301 and 302:
[0084] When it is not possible to synchronize all sub-regions, the cleaning priority value Yp(p,s) of each sub-region is obtained, and the sub-regions are sorted according to the cleaning priority value Yp(p,s) to generate a cleaning order. When cleaning data in each sub-region, the cleaning process can be completed in an orderly manner. By cleaning different sub-regions in sequence, the remaining sub-regions can maintain normal storage and writing processes, reducing the degree of mutual interference between different sub-regions and maintaining the normal usage state within the storage area. At the same time, the cleaning priority value Yp(p,s) can also be used to evaluate the busy level of each sub-region, and the working status of each sub-region can be adjusted according to the evaluation results to extend its working life.
[0085] Step 4: Identify the cleaning characteristics of the data in each sub-region, and match several corresponding candidate solutions from the pre-built trial cleaning solution library. Use the trained data storage model to perform simulation analysis on the availability of the candidate solutions, and select the recommended solution from the candidate solutions based on the analysis results.
[0086] Step four includes the following:
[0087] Step 401: When it is necessary to clean the data in each sub-region, identify the pre-processed data in each sub-region and obtain the corresponding data features, such as the integrity features, accuracy features, consistency features and correlation features of the data to be cleaned. Use the types and quantities of data features in the sub-region as cleaning features to mark each sub-region.
[0088] Step 402: Obtain several data cleaning solutions through online linear retrieval or offline pre-collection. After summarizing, build a cleaning solution library in advance. Based on the correspondence between the cleaning features of each sub-region and the cleaning solution, use the trained matching model to match the corresponding cleaning solution from the pre-built cleaning solution library and output the cleaning solution as a candidate solution.
[0089] Step 403: Collect specification data, data type and distribution status data of each sub-region, and process data of reading, writing and deleting data in each sub-region within the storage area. After summarizing, generate a storage area status set. Perform feature recognition on the data within the storage area and obtain the corresponding feature data after recognition. After summarizing several feature data, generate a feature data set and select a portion of the data from the feature data set as the training set and test set.
[0090] An initial model is constructed using a neural convolutional network. After training and testing the initial model, the trained initial model is output as a data storage model. The usability of one or more candidate solutions is simulated and analyzed using the trained data storage model. If all candidate solutions are feasible, the candidate solution with the best performance is output as the recommended solution.
[0091] If none of the above are feasible, after obtaining the conditional parameters of the candidate solutions, the conditional parameters are adjusted to obtain several adjusted candidate solutions; the trained data storage model is used to conduct simulation tests on the adjusted candidate solutions, and the solution with the best cleaning effect is selected as the recommended solution.
[0092] When using this method, refer to steps 401 to 403:
[0093] After completing the data preprocessing, the data in the storage area is identified and the corresponding cleaning features are obtained. Then, the corresponding cleaning scheme is matched for each sub-area from the pre-built cleaning scheme library. After receiving the cleaning instruction, the time for preparing the cleaning scheme in advance can be reduced and the data cleaning efficiency can be improved.
[0094] Meanwhile, by constructing a data storage model, the matched cleaning schemes are tested and screened using the trained data storage model. The part with the best cleaning effect is selected, that is, the part with the highest data quality in the sub-region after cleaning. This improves the cleaning effect and the cleaning quality. As a further effect, by modifying the cleaning scheme, the correspondence between the modified scheme and the cleaning features can be improved, which can also reduce the time for modifying the cleaning scheme and improve the efficiency of data cleaning.
[0095] Step 5: After executing the recommendation scheme, use the trained data storage model to predict the data quality parameters, establish a data quality set, continuously obtain several data quality coefficients Tq(s,q) from the data quality set, predict the time taken to obtain the data quality coefficients Tq(s,q) exceeding the quality threshold, and if it is shorter than expected, issue a reminder instruction.
[0096] Step five includes the following:
[0097] Step 501: After executing the recommendation scheme to clean the data in each sub-region, the storage region continues to receive and write new data, and the trained data storage model is used to predict the storage status of the data. The prediction period is set, and at the end of the prediction period, the corresponding quality parameters are obtained. After adding a timestamp, the data quality set in the storage region is re-established.
[0098] Step 502: Continuously obtain several data quality coefficients Tq(s, q) at the end of the prediction period from the data quality set. If none of the obtained data quality coefficients Tq(s, q) exceed the current quality threshold, arrange the obtained data quality coefficients Tq(s, q) in order, and predict the changing trend of the data quality coefficients Tq(s, q) based on the smoothing exponential model.
[0099] The time required for the data quality coefficient Tq(s, q) to exceed the quality threshold is taken as the risk time. Based on historical data and the expected management of data quality, a time threshold is set in advance. When the predicted risk time exceeds the time threshold, an external reminder instruction is issued.
[0100] When using this method, refer to steps 501 and 502:
[0101] After cleaning the sub-regions sequentially according to the selected cleaning scheme, the corresponding sub-regions are re-entered into their respective usage states. Under the condition of continuously receiving and writing data, the trained data storage model is used to predict the data storage status at each time point within the sub-region. The corresponding data quality coefficient Tq(s, q) is obtained, and its changing trend is predicted in conjunction with the smoothing exponential model to obtain the corresponding risk time. Thus, based on the changes in risk time, changes in data quality within the sub-region can be predicted and alerted, thereby facilitating timely and early detection and processing during data cleaning.
[0102] Please see Figure 2 This invention provides a data link-based communication data cleaning system, comprising:
[0103] The early warning unit, if data is continuously written to the storage area, establishes a data receiving status set from the data writing status data, and generates a data receiving status coefficient Ct(s,n) from the data receiving status set. If the data receiving status coefficient Ct(s,n) exceeds the status threshold, it issues an early warning command to the outside.
[0104] After receiving the warning instruction, the data processing unit preprocesses the data groups in the sub-region, identifies and obtains the data quality characteristics of each data group, summarizes them to generate a data quality set in the storage region, and generates a data quality coefficient Tq(s,q) from the data quality set. If the obtained data quality coefficient Tq(s,q) exceeds the quality threshold, a cleaning instruction is issued.
[0105] The sorting unit, in conjunction with the data write and read status of each sub-region, establishes a data condition set, generates a cleaning priority value Yp(p, s) for data processing in each sub-region from the data condition set, sorts the cleaning priority values Yp(p, s) in each sub-region, and uses the obtained sorting as the cleaning order.
[0106] The scheme analysis unit identifies the cleaning characteristics of the data acquired in each sub-region, matches several corresponding candidate schemes from the pre-built trial cleaning scheme library, uses the trained data storage model to perform simulation analysis on the availability of the candidate schemes, and selects the recommended scheme from the candidate schemes based on the analysis results.
[0107] After executing the recommendation scheme, the prediction unit uses the trained data storage model to predict the quality parameters of the data, establishes a data quality set, and continuously obtains several data quality coefficients Tq(s,q) from the data quality set. It predicts the time taken to obtain the data quality coefficients Tq(s,q) exceeding the quality threshold. If it is shorter than expected, it issues a reminder instruction.
[0108] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0109] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0110] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0111] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division of a waterway underwater topography change analysis system and method. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0112] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0113] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0114] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0115] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0116] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for cleaning communication data based on a data link, characterized in that: Includes the following: If data is continuously written to the storage area, a data reception state set is established based on the data writing status data, and a data reception state coefficient is generated from the data reception state set. If the data reception status coefficient When the status threshold is exceeded, an early warning command is issued to the outside. Upon receiving the warning command, the data groups within the sub-region are preprocessed, the data quality characteristics of each data group are identified and obtained, and the data quality is aggregated to generate a data quality set within the storage region. Data quality coefficients are then generated from this data quality set. If the quality coefficient of the acquired data If the quality threshold is exceeded, a cleaning command is issued. By combining the data write and read status of each sub-region, a data condition set is established, and the data condition set is used to generate the cleaning priority value for data processing in each sub-region. The specific method is as follows: The data replacement ratio Qp and the number of reads Qs are linearly normalized, and the corresponding data values are mapped to the interval... Then, follow the formula below: Weighting coefficients: , ,and Cleaning priority values for each sub-region Sort the data, and use the obtained sorting order as the cleaning order; Identify the cleaning characteristics of the data acquired in each sub-region, and match several corresponding candidate solutions from the pre-built trial cleaning solution library. Use the trained data storage model to simulate and analyze the availability of the candidate solutions, and select the recommended solution from the candidate solutions based on the analysis results. After implementing the recommendation scheme, the trained data storage model is used to predict the data quality parameters, a data quality set is established, and several data quality coefficients are continuously obtained from the data quality set. Predict the data quality coefficient If the time taken exceeds the quality threshold, but is shorter than expected, a reminder instruction will be issued.
2. The method for cleaning communication data based on a data link according to claim 1, characterized in that: Monitor the data reception status to obtain the data reception status within the reception period, obtain the amount of data received in each reception period, and generate the data reception amount Ss; The storage area is divided into several sub-regions, the write ratio P is obtained in each receiving cycle, and then the uniformity Un of data distribution in the storage area is calculated. After continuously acquiring several received values Ss and summarizing the uniformity Un, a data reception state set is generated; the data reception state coefficients are then generated from the data reception state set. If the data reception status coefficient When the status threshold is exceeded, an early warning command is issued to the outside.
3. The method for cleaning communication data based on a data link according to claim 2, characterized in that: The data reception state coefficients are generated from the data reception state set. The specific method is as follows: After linearly normalizing the received quantity Ss and the uniformity Un, the corresponding data values are mapped to the interval. Then, follow the formula below: in, This represents the historical average of the received data. Its current value; The historical average of uniformity. Its current value; n is a positive integer. n is the number of detection periods, and the weighting coefficient is: , ,and .
4. The method for cleaning communication data based on a data link according to claim 1, characterized in that: The received data is classified according to its type, resulting in several different data groups. After preprocessing the data within each group, the data quality characteristics of each group are identified and extracted, including: The data within each group are arranged in order of their acquisition time, and data analysis is performed to obtain the quality parameters of the data within each group, including relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss. After summarizing the above data within each group, a data quality set is established within the storage area.
5. The method for cleaning communication data based on a data link according to claim 4, characterized in that: Generating data quality coefficients from a data quality set includes, first, generating data quality values from the data quality set, specifically including: The relative range Sxs, skewness coefficient Pxs, and kurtosis coefficient Kss within each data set are normalized, and the corresponding data values are mapped to intervals. The data quality value is calculated using the following formula. : Wherein, the parameter means: n is a positive integer greater than 1, , where is the number of data groups within the sub-region, and the weighting coefficient is: , , and The The mean of the relative range. The mean of the skewness coefficients. This represents the mean of the kurtosis coefficients.
6. The method for cleaning communication data based on a data link according to claim 5, characterized in that: Data quality value Generate data quality coefficient The specific method is as follows: Among them, among them, The number of sub-regions is a positive integer greater than 1; This represents the median data quality value within the sub-region. The mean of the data quality values for the sub-region; if the obtained data quality coefficient If the quality exceeds the preset threshold, a cleaning command is sent to the outside.
7. The method for cleaning communication data based on a data link according to claim 6, characterized in that: Upon receiving the cleaning instruction, within the query period, the data replacement ratio Qp and the number of reads Qs are obtained respectively, and the data condition set is established by summarizing them. Data cleaning priority values for each sub-region are generated from the data condition set. After obtaining the cleaning priority value of each sub-region Then, the cleaning priority value for each sub-region is determined. Sort the data, and use the obtained sorting order as the cleaning order.
8. The method for cleaning communication data based on a data link according to claim 7, characterized in that: The preprocessed data in each sub-region is identified, and the corresponding data features are obtained. The types and quantities of data features in the sub-region are used as cleaning features to obtain several data cleaning schemes. After summarizing, a cleaning scheme library is pre-built. Based on the correspondence between the cleaning features and cleaning schemes in each sub-region, the trained matching model is used to match the corresponding cleaning schemes from the pre-built cleaning scheme library and the cleaning schemes are used as candidate schemes.
9. The method for cleaning communication data based on a data link according to claim 8, characterized in that: An initial model is built using a neural convolutional network. After training and testing, the trained initial model is output as a data storage model. The usability of one or more candidate solutions is simulated and analyzed using the trained data storage model. The candidate solution with the best performance is then output as the recommended solution. Alternatively, the conditional parameters are adjusted to obtain several adjusted candidate solutions as recommended solutions.
10. The method for cleaning communication data based on a data link according to claim 1, characterized in that: If the storage area continues to receive and write new data, the trained data storage model is used to predict the storage status of the data. At the end of the prediction period, the corresponding quality parameters are obtained, and a data quality set within the storage area is re-established. Several data quality coefficients at the end of the prediction period are then continuously obtained from the data quality set. If the quality coefficient of the acquired data None of them exceeded the current quality threshold; the quality coefficients of the acquired data were used. Arrange the data in an ordered manner and apply the smoothing exponential model to the data quality coefficients. To predict the changing trend; Obtaining data quality coefficients The time required for the quality threshold to be exceeded is considered the risk time, based on historical data. Historical data and expectations for data quality management are used to pre-set time thresholds. If the predicted risk period exceeds the time threshold... When the value is reached, a reminder command is sent to the outside.