Accurate duplicate removal method and system based on gynaecology and obstetrics data acquisition
Through the precise deduplication method based on obstetrics and gynecology data acquisition, and using technical means such as data frequency statistics and hierarchical clustering algorithms, a deduplication strategy model is built, which solves the problem of redundancy in obstetrics and gynecology data, and achieves efficient and accurate data deduplication.
Patent Information
- Application Number
- CN202510523285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing obstetrics and gynecology data collection has the problems of redundant records and inaccurate information. The traditional deduplication method is inefficient and has low accuracy.
The precise deduplication strategy model is adopted based on obstetrics and gynecology data acquisition, including data frequency statistics, hierarchical clustering algorithm, sensitivity analysis, similarity threshold definition, dynamic time window analysis and cyclic convolution algorithm, to build a deduplication strategy model to achieve accurate deduplication of data.
Improve data accuracy and consistency, reduce redundant information, and ensure data quality and deduplication accuracy.
Smart Images

Figure CN120386984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a precise duplicate removal method and system based on obstetrics and gynecology data collection. Background Art
[0002] In modern medical practice, obstetrics and gynecology is an important field, involving the diagnosis, treatment, and monitoring of female reproductive health and the perinatal period. To provide accurate medical services and decision-making support, obstetricians and gynecologists usually need to collect a large amount of obstetrics and gynecology data, including medical records, examination results, test reports, etc. However, due to the complexity of hospital systems and a large amount of data input, there are often problems of duplicate records in obstetrics and gynecology data, resulting in redundant and inaccurate information. Traditional data duplicate removal methods often have problems of low duplicate removal efficiency and low accuracy. Therefore, an intelligent precise duplicate removal method and system based on obstetrics and gynecology data collection are needed. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes a precise duplicate removal method and system based on obstetrics and gynecology data collection to solve at least one of the above technical problems.
[0004] To achieve the above object, the present invention provides a precise duplicate removal method based on obstetrics and gynecology data collection, including the following steps:
[0005] Step S1: Obtain obstetrics and gynecology data; perform data frequency statistics on the obstetrics and gynecology data to generate frequency data; perform duplicate attribute analysis on the frequency data to generate duplicate attribute data;
[0006] Step S2: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; perform sensitivity analysis on the duplicate clustering data to generate sensitivity data; calculate the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data;
[0007] Step S3: Define a similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data; calculate the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data;
[0008] Step S4: Perform dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data; define the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data; Step S5: Perform duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data; optimize the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy;
[0009] Step S6: Use the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy, and construct a duplicate removal strategy model to perform precise data duplicate removal.
[0010] By obtaining obstetrics and gynecology data, the present invention can obtain information related to this field, providing a basis for subsequent data analysis and processing. Frequency statistics on obstetrics and gynecology data can understand the occurrence frequencies of various attribute values, reveal the data distribution, and provide a reference basis for subsequent analysis. By performing duplicate attribute analysis on the frequency data, it can be determined which attribute values appear repeatedly in the data, which helps to discover redundant information and duplicate records in the data, providing a basis for data duplicate removal and optimization. Using the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data can group data with similar characteristics, thereby understanding the internal structure and correlation of the data. Performing sensitivity analysis on the duplicate clustering data can evaluate the stability and reliability of the data clustering results, and understand the sensitivity of the clustering results to changes in the input data. Based on the sensitivity data, calculating the importance weights for the duplicate attribute data can determine the importance of each attribute in the data duplicate removal process, providing a basis and trade-off for subsequent duplicate removal decisions. Based on the importance weight data, a similarity threshold can be defined to measure the similarity between data records, further guiding the formulation of duplicate removal decisions. Based on the similarity threshold data, calculating the similarity of the duplicate attribute data can determine the similarity degree between data records, providing data support for subsequent duplicate removal decisions. Performing dynamic time window analysis on the duplicate attribute data can capture the temporal characteristics and change trends of data records, further understanding the dynamics of the data. Based on the dynamic time window data, defining the time granularity for the duplicate attribute data can determine the granularity size of data records in the time dimension, providing a basis for subsequent duplicate removal decisions and time-related analysis. According to the time granularity data, performing duplicate removal decision analysis on the duplicate attribute data, considering the duplicate situations under different time granularities, and formulating a reasonable duplicate removal strategy to reduce redundant data and improve data quality. Optimizing the strategy for the duplicate removal decision data can further improve the effect and performance of the duplicate removal strategy, and improve the accuracy and reliability of the duplicate removal results. Using the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy can construct a duplicate removal strategy model that can comprehensively consider factors such as the importance, similarity, and time granularity of each attribute to achieve more precise data duplicate removal. By executing the duplicate removal strategy model, precise duplicate removal operations can be performed on the data, removing duplicate records and redundant information, and improving the consistency and accuracy of the data.
[0011] In this specification, a precise duplicate removal system based on obstetrics and gynecology data collection is provided, including:
[0012] An information collection module that acquires obstetrics and gynecology data, performs data frequency statistics on the obstetrics and gynecology data to generate frequency data, and performs duplicate attribute analysis on the frequency data to generate duplicate attribute data;
[0013] A weight analysis module that uses the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data, performs sensitivity analysis on the duplicate clustering data to generate sensitivity data, and calculates the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data;
[0014] A similarity calculation module that defines a similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data, and calculates the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data;
[0015] A time window module that performs dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data, and defines the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data;
[0016] A duplicate removal decision module that performs duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data, and optimizes the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy;
[0017] A strategy model module that uses the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy to construct a duplicate removal strategy model for performing precise data duplicate removal.
[0018] The present invention obtains obstetrics and gynecology data through an information collection module, including medical records, medical images, laboratory results, etc. These data can be used for subsequent data analysis and processing. By performing data frequency statistics on the obstetrics and gynecology data, the frequency of each data item can be calculated. The frequency data reflects the importance and universality of the data item in the sample set. By analyzing the frequency data, attributes with repeated occurrences can be identified. The repeated attribute data shows which attributes have similar values or patterns in different data items. Using the hierarchical clustering algorithm to perform clustering analysis on the repeated attribute data, data items with similar characteristics can be grouped into the same category. The repeated clustering data provides the clustering results between data items, helping to identify groups of data items with similar characteristics. By performing sensitivity analysis on the repeated clustering data, the sensitivity of different data items to the clustering results can be evaluated. The sensitivity data shows the contribution degree of each data item to the clustering result, helping to determine the attributes that mainly affect the clustering result. Based on the sensitivity data, importance weights are calculated for the repeated attribute data. The importance weight data reflects the contribution degree of each attribute to the similarity of data items, helping to determine the importance ranking of attributes. Based on the importance weight data, a similarity threshold is defined for the repeated attribute data. The similarity threshold data determines which data items are considered similar and groups them as candidates for duplicate data. Based on the similarity threshold data, similarity calculations are performed on the repeated attribute data. The similarity data quantifies the similarity degree between different data items, helping to identify true duplicate data items. By performing dynamic time window analysis on the repeated attribute data, the temporal correlation of data items can be determined. The dynamic time window data shows the temporal relationship and temporal sequence pattern between data items. Based on the dynamic time window data, the time granularity is defined for the repeated attribute data. The time granularity data determines the granularity size of data items in time, helping to determine the time span for duplicate removal decisions. By performing duplicate removal decision analysis on the repeated attribute data, the duplicate removal decision data determines which duplicate data items should be retained or deleted, helping to clean up duplicate items in the dataset. By optimizing the strategy for the duplicate removal decision data, more factors such as data quality and business requirements can be considered to generate a more reasonable and effective duplicate removal optimization strategy. The optimized strategy can more accurately identify and process duplicate data items. Using the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy, a duplicate removal strategy model is constructed. This model can perform precise data duplicate removal operations according to the input data characteristics and the duplicate removal optimization strategy, effectively eliminating duplicate data items. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the step flow of the precise duplicate removal method and system based on obstetrics and gynecology data collection according to the present invention;
[0020] Figure 2 It is a schematic diagram of the detailed implementation step flow of step S1;
[0021] Figure 3 It is a schematic diagram of the detailed implementation steps of step S2;
[0022] Figure 4 It is a schematic diagram of the detailed implementation steps of step S3. Specific implementation manner
[0023] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] The embodiments of the present application provide a precise duplicate removal method and system based on obstetrics and gynecology data collection. The execution subjects of the precise duplicate removal method and system based on obstetrics and gynecology data collection include, but are not limited to, the following general computing nodes that carry this system: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. The data processing platform includes, but is not limited to, at least one of an audio and image management system, an information management system, and a cloud data management system.
[0025] Please refer to Figures 1 to 4 , the present invention provides a precise duplicate removal method based on obstetrics and gynecology data collection. The method includes the following steps:
[0026] Step S1: Obtain obstetrics and gynecology data; perform data frequency statistics on the obstetrics and gynecology data to generate frequency data; perform duplicate attribute analysis on the frequency data to generate duplicate attribute data;
[0027] Step S2: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; perform sensitivity analysis on the duplicate clustering data to generate sensitivity data; calculate the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data;
[0028] Step S3: Define a similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data; calculate the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data;
[0029] Step S4: Perform dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data; define the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data; Step S5: Perform duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data; optimize the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy;
[0030] Step S6: Use the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy to construct a duplicate removal strategy model to perform precise data duplicate removal.
[0031] By obtaining obstetrics and gynecology data, the present invention can acquire information related to this field, providing a basis for subsequent data analysis and processing. Conducting frequency statistics on obstetrics and gynecology data can understand the occurrence frequencies of various attribute values, revealing the data distribution and providing a reference basis for subsequent analysis. By performing duplicate attribute analysis on the frequency data, it can be determined which attribute values appear repeatedly in the data, which helps to discover redundant information and duplicate records in the data, providing a basis for data deduplication and optimization. Using the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data can group data with similar characteristics, thereby understanding the internal structure and correlation of the data. Conducting sensitivity analysis on the duplicate clustering data can evaluate the stability and reliability of the data clustering results, and understand the sensitivity of the clustering results to changes in the input data. Based on the sensitivity data, calculating the importance weights for the duplicate attribute data can determine the importance of each attribute in the data deduplication process, providing a basis and trade-off for subsequent deduplication decisions. Based on the importance weight data, a similarity threshold can be defined to measure the similarity between data records, further guiding the formulation of deduplication decisions. Based on the similarity threshold data, calculating the similarity of the duplicate attribute data can determine the similarity degree between data records, providing data support for subsequent deduplication decisions. Conducting dynamic time window analysis on the duplicate attribute data can capture the temporal characteristics and change trends of data records, further understanding the dynamics of the data. Based on the dynamic time window data, defining the time granularity for the duplicate attribute data can determine the granularity size of data records in the time dimension, providing a basis for subsequent deduplication decisions and time-related analysis. According to the time granularity data, conducting deduplication decision analysis on the duplicate attribute data, considering the duplicate situations under different time granularities, and formulating a reasonable deduplication strategy to reduce redundant data and improve data quality. Optimizing the strategy for the deduplication decision data can further improve the effect and performance of the deduplication strategy, enhancing the accuracy and reliability of the deduplication results. Using the circular convolution algorithm to perform dilated convolution on the deduplication optimization strategy can construct a deduplication strategy model that can comprehensively consider factors such as the importance, similarity, and time granularity of each attribute to achieve more accurate data deduplication. By executing the deduplication strategy model, precise deduplication operations can be performed on the data, removing duplicate records and redundant information, and improving data consistency and accuracy.
[0032] In the embodiment of the present invention, with reference to Figure 1 as described, it is a schematic diagram of the step flow of a precise deduplication method and system based on obstetrics and gynecology data collection according to the present invention. In this example, the steps of the precise deduplication method based on obstetrics and gynecology data collection include:
[0033] Step S1: Obtain obstetrics and gynecology data; conduct data frequency statistics on the obstetrics and gynecology data to generate frequency data; perform duplicate attribute analysis on the frequency data to generate duplicate attribute data;
[0034] In this embodiment, a dataset related to obstetrics and gynecology is collected or obtained. These data may include patients' personal profiles, medical records, medical indicators, etc., covering information in different aspects of the field of obstetrics and gynecology. Frequency statistical analysis is performed on the obtained obstetrics and gynecology data to understand the distribution of each attribute in the dataset, including the frequency of occurrence of each attribute value and statistical indicators such as mean and variance. According to the results of the frequency statistics, frequency data is generated. These data can be data tables or datasets that record each attribute and its corresponding frequency, mean, variance and other statistical information. Duplicate attribute analysis is performed on the frequency data. This step aims to identify and process duplicate information that appears in multiple attributes, such as the same patient personal profile or duplicate medical indicator records. Based on the results of the duplicate attribute analysis, duplicate attribute data is generated. These data can be a data table or dataset that records duplicate attributes and their related information for further analysis such as subsequent clustering, correlation analysis, and missing value statistics.
[0035] Step S2: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; perform sensitivity analysis on the duplicate clustering data to generate sensitivity data; calculate the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data;
[0036] In this embodiment, a suitable hierarchical clustering algorithm (such as agglomerative hierarchical clustering) is used to perform clustering analysis on the duplicate attribute data. Before clustering, it is necessary to first determine a similarity measurement method (such as Euclidean distance, cosine similarity, etc.) to measure the similarity degree between the duplicate attribute data. According to the results of the clustering analysis, the duplicate attribute data is divided into different clustering clusters. Each cluster contains duplicate attribute data with similar attribute patterns. The generated duplicate clustering data can be a data structure representing the clustering clusters and their members, facilitating subsequent analysis and processing. For the duplicate clustering data, a sensitivity analysis is performed. This step aims to evaluate the sensitivity of different clusters, that is, the degree of response to changes or perturbations in the input data. For example, sensitivity analysis methods (such as Monte Carlo simulation, parameter sensitivity analysis, etc.) can be used to evaluate the stability and consistency of the duplicate clustering data. Based on the results of the sensitivity analysis, sensitivity data is generated. These data can represent the sensitivity degree of each clustering cluster, which can be a numerical value or a related indicator, used to measure the stability and reliability of the clustering results. Using the sensitivity data, the importance weights of the duplicate attribute data can be calculated. The goal of this step is to determine the degree of importance contribution of each attribute or clustering cluster to the overall data set. Various methods can be used, such as optimization algorithms based on gradient descent, linear regression, etc., to calculate the importance weights according to the sensitivity data. According to the results of the importance weight calculation, importance weight data is generated. These data can be a data structure representing the importance weight values of each attribute or clustering cluster, facilitating subsequent data analysis and decision-making.
[0037] Step S3: Define a similarity threshold for the duplicate attribute data based on the importance weight data, generating similarity threshold data; calculate the similarity of the duplicate attribute data based on the similarity threshold data, generating similarity data;
[0038] In this embodiment, the importance weight data is used to define the similarity threshold. The similarity threshold can be determined according to the relative magnitudes of the importance weights. The similarity threshold between attributes with higher weights can be set smaller, indicating a higher similarity requirement. According to the definition of the similarity threshold, similarity threshold data is generated. These data can be a data structure representing the similarity thresholds between attributes, which can be used for subsequent similarity calculations. Using the similarity threshold data, the similarity of the duplicate attribute data is calculated. Common similarity calculation methods include cosine similarity, Jaccard similarity, etc. For each pair of duplicate attribute data, its similarity is calculated and compared with the similarity threshold. According to the results of the similarity calculation, similarity data is generated. These data can represent the similarity degree between each pair of duplicate attribute data, which can be a numerical value (such as a similarity score) or a related indicator, used to measure the similarity degree between the data.
[0039] Step S4: Perform dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data; define the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data;
[0040] In this embodiment, dynamic time window analysis is performed using the similarity data, which can help identify the clustering patterns in time of duplicate data with similar attributes. The dynamic time window is a time window with a variable size, used to determine which duplicate data are close enough within a period of time to be grouped together. By comparing the similarity data, the appropriate size of the time window can be determined. Based on the results of the dynamic time window analysis, dynamic time window data is generated. These data can represent the clustering patterns of the duplicate data in each time window. Generally, the dynamic time window data can be a data structure that records the identifiers or other relevant information of the duplicate data in each time window. The time granularity is defined using the dynamic time window data. The time granularity refers to dividing time into different intervals or segments to better understand and analyze the time characteristics of the duplicate data. Based on the clustering patterns of the dynamic time window data, the start and end times of each time granularity, as well as the identifiers or other relevant information of the duplicate data included in this time granularity, can be determined. According to the time granularity definition, time granularity data is generated. These data can represent the time characteristics of the duplicate data within different time granularities. The time granularity data can be a data structure that records the start and end times of each time granularity, as well as the identifiers or other relevant information of the duplicate data included in this time granularity.
[0041] Step S5: Perform duplicate removal decision analysis on the duplicate attribute data based on the time granularity data to generate duplicate removal decision data; optimize the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy;
[0042] In this embodiment, based on the time granularity data, the duplicate data within each time granularity is analyzed. Considering the attributes, time characteristics, and business requirements of the duplicate data, an appropriate deduplication strategy is determined. The deduplication decision can include retaining the latest record, retaining the earliest record, selecting records based on a certain attribute rule, etc. According to the results of the deduplication decision analysis, deduplication decision data is generated. The deduplication decision data can be a data structure that records the identifiers or other relevant information of the duplicate data selected for retention within each time granularity. The generated deduplication decision data is evaluated and analyzed to determine the direction of the optimization strategy. According to the analysis results, the following optimization strategies can be considered: adjusting the weights or rules of the deduplication strategy to better meet business requirements, adopting machine learning or data mining techniques to automatically learn and optimize the deduplication strategy, combining artificial intelligence algorithms to automatically identify and handle special cases or boundary cases, considering system performance and resource utilization, and proposing targeted strategy optimization suggestions. Based on the strategy optimization results, a deduplication optimization strategy is generated. The deduplication optimization strategy can be a guiding document or suggestion, including updates or improvements to the deduplication decision strategy.
[0043] Step S6: Use the circular convolution algorithm to perform dilated convolution on the deduplication optimization strategy to construct a deduplication strategy model for performing precise data deduplication.
[0044] In this embodiment, the circular convolution algorithm is an algorithm for processing time series data. Applying the circular convolution algorithm in the deduplication optimization strategy can achieve precise data deduplication. The circular convolution operation convolves the strategy in the deduplication strategy model with the data to produce an output sequence with a similar sliding window effect. According to the previously generated deduplication optimization strategy, a deduplication strategy model is constructed. This model can be a neural network model for representing and executing the deduplication strategy. The optimized strategy is used as the input of the model. A dilation operation is performed on the data to be deduplicated. Data dilation refers to copying or amplifying the original data so that the sliding window can be used to process the data during the circular convolution process. The dilated data can increase the perception range of the model and improve the accuracy of deduplication. Using the circular convolution algorithm, the dilated data is convolved with the deduplication strategy model. During the circular convolution process, the model matches and determines the data within the window with the strategy, and determines whether to retain or remove the data according to the rules of the strategy. According to the results of the circular convolution operation, precise data deduplication is performed on the data. According to the output of the strategy model, it is decided whether to retain or remove specific data records. According to business requirements, the deduplicated data can be saved to a new dataset or overwrite the original dataset.
[0045] In this embodiment, refer to Figure 2 As described, it is a schematic diagram of the detailed implementation step process of step S1. In this embodiment, the detailed implementation steps of step S1 include:
[0046] Step S11: Obtain obstetrics and gynecology data, which includes patient information data, case information, diagnosis results, drug information, and pregnancy and childbirth information;
[0047] Step S12: Extract features from the obstetrics and gynecology data to generate feature data;
[0048] Step S13: Conduct data frequency statistics on the feature data to generate frequency data;
[0049] Step S14: Conduct attribute comparison on the frequency data to generate attribute comparison data;
[0050] Step S15: Conduct frequency difference analysis on the attribute comparison data to generate frequency difference data;
[0051] Step S16: Conduct repeated attribute analysis on the frequency difference data to generate repeated attribute data.
[0052] By obtaining obstetrics and gynecology data, the present invention can acquire rich information related to this field, including patient information, case information, diagnosis results, drug information, and pregnancy and childbirth information. This provides a comprehensive data basis for subsequent analysis and processing. The acquisition of obstetrics and gynecology data ensures the integrity of the dataset, enabling a comprehensive understanding of relevant information in the obstetrics and gynecology field from multiple perspectives, which helps with comprehensive analysis and comprehensive decision-making. By performing feature extraction on obstetrics and gynecology data, representative and important features can be extracted from the original data. These features can better describe the characteristics and attributes of the data, providing a more accurate data basis for subsequent analysis and modeling. Feature extraction makes the original data more concise and easier to understand, reducing data redundancy and complexity, and helping to improve the efficiency and accuracy of data processing. By performing frequency statistics on the feature data, the distribution of different feature values in the data can be understood, including the occurrence frequency and proportion of each feature value, which helps to have a clearer understanding of the overall characteristics of the data. Frequency data can reveal common features and patterns in the data, discover important attributes and trends in the data, and provide a reference for subsequent analysis and decision-making. By performing attribute comparison on the frequency data, the relationships and differences between different attributes can be compared, understanding their mutual influence and correlation, and comprehensively understanding the data features. Attribute comparison data can reveal the correlation and interaction between attributes, helping to discover hidden association relationships and potential rules in the data, and providing guidance for further data analysis and mining. By performing frequency difference analysis on the attribute comparison data, the degree of difference between different attributes can be quantified, thereby evaluating their importance and discrimination in the data, and providing a basis for subsequent analysis and decision-making. Frequency difference data can determine the attributes with significant differences in the data, and more specifically focus on and process these attributes to obtain more valuable information. By performing duplicate attribute analysis on the frequency difference data, redundant attributes existing in the data can be identified, that is, those attributes that are highly similar or highly correlated with other attributes, reducing data redundancy and complexity, and improving the efficiency and accuracy of data processing. The generation of duplicate attribute data can streamline and optimize the data, removing unnecessary duplicate information, making the data more compact and efficient, and providing a better data basis for subsequent analysis and application.
[0053] In this embodiment, the ways to obtain obstetrics and gynecology data can be the hospital's electronic medical record system, scientific research databases, or other legal data sources. Clearly define the scope of obstetrics and gynecology data, including relevant data such as patient information data, case information, diagnosis results, drug information, and pregnancy and childbirth information. According to the data sources and the defined scope, obtain obstetrics and gynecology data through appropriate means to ensure the integrity and accuracy of the data. According to the requirements and analysis objectives, select features related to specific tasks from the obstetrics and gynecology data. Feature selection can be carried out using domain knowledge or data analysis methods. Use appropriate methods to extract the selected features from the original data, which may involve operations such as data transformation, normalization, or encoding to obtain feature data that can be used for further analysis. For each feature, calculate its frequency in the dataset, that is, the number of times a specific value appears in the entire dataset. According to the type and distribution of the features, select appropriate statistical methods, and indicators such as frequency, proportion, or percentage can be used to represent the frequency of the features. Organize the frequency statistics results of each feature into frequency data to construct a frequency table or frequency vector. Determine the attributes or features that need to be compared, which can be different time periods, different disease types, or different patient groups, etc. According to the comparison objectives and attribute types, select appropriate comparison methods, such as comparing frequency differences, calculating relative risks, or conducting correlation analysis, etc. Conduct a comparative analysis of the selected attributes to obtain information on the differences or correlations between the attributes and generate attribute comparison data. According to the type and distribution of the attribute comparison data, select appropriate frequency difference analysis methods, such as chi-square test, t-test, or analysis of variance, etc. Use the selected frequency difference analysis method to calculate the frequency difference indicators in the attribute comparison data, such as p-value, significance level, or effect size, etc. Organize the frequency difference indicators into frequency difference data to record the difference information between the attribute comparisons. Determine the concept and judgment criteria of duplicate attributes, which can be that the attribute values are the same, similar, or conform to certain rules, etc. According to the definition of duplicate attributes, select appropriate analysis methods, such as finding the intersection based on set theory, using similarity measures, or designing matching algorithms, etc. Conduct duplicate attribute analysis on the frequency difference data to find the data records with duplicate attributes, organize them into duplicate attribute data, and record the relevant information of the duplicate attributes.
[0054] In this embodiment, step S16 includes the following steps:
[0055] Step S161: Perform duplicate attribute identification on the frequency difference data to generate attribute identification data;
[0056] Step S162: Perform outlier detection on the frequency data based on the attribute identification data to generate outlier data;
[0057] Step S163: Perform duplicate attribute analysis on the frequency data according to the outlier data to generate duplicate attribute data.
[0058] By performing repeated attribute identification on frequency difference data, the present invention can accurately determine which attributes have duplicate occurrences in the data, i.e., have the same or highly similar characteristics. This helps to identify and process redundant information in the data, simplify the data set, and improve data quality. The attribute-identified data provides important guidance for subsequent data processing and analysis, enabling more effective data cleaning and organization, removing duplicate attributes or merging them into one attribute to better utilize the data and avoid redundancy. By performing outlier detection based on the attribute-identified data, values with abnormal characteristics in the frequency data can be identified and marked, which helps to discover potential data anomalies or errors, ensuring data quality and the accuracy of analysis. The generation of outlier data can clean and correct the data, removing outliers or performing appropriate processing, thereby improving the quality and reliability of the data and providing a more accurate data foundation for subsequent analysis and applications. By performing repeated attribute analysis based on the outlier data, redundant attributes in the data can be further identified and confirmed. The outlier data can reveal the similarities and correlations between attributes, determining which attributes have similar characteristics, thus further optimizing and simplifying the data set. The generation of repeated attribute data enables more precise identification and processing of duplicate information in the data, removing redundant attributes or merging them to achieve the purpose of data optimization and streamlining. This helps to improve data processing efficiency and accuracy and reduce the consumption of data storage and computing resources.
[0059] In this embodiment, the concept and determination criteria of duplicate attributes are determined. For example, attributes with exactly the same attribute values or those that conform to specific rules are considered duplicates. Traverse the records in the frequency difference data one by one. For each record, check if there are other records in the frequency difference data with the same attributes. If so, identify the duplicate attributes of the record and record the relevant information. Organize the records with identified duplicate attributes into attribute identification data, recording the duplicate attribute information for each record. According to the duplicate attribute information in the attribute identification data, perform a filtering operation on the frequency data, only retaining the records with duplicate attributes. According to the requirements and data characteristics, select an appropriate outlier detection method, such as a method based on statistical analysis (such as the 3σ principle or outlier determination), a machine learning algorithm, or rule definition, etc. Use the selected outlier detection method to perform outlier detection on the filtered frequency data, and organize the detected outlier records into outlier data, recording the relevant information for each outlier. According to the outlier data, perform a filtering operation on the frequency data, only retaining the records with outliers. According to the data characteristics and analysis objectives, select a suitable duplicate attribute analysis method, such as intersection finding based on set theory, similarity measurement, or matching algorithms, etc. According to the data characteristics and analysis objectives, select a suitable duplicate attribute analysis method, such as intersection finding based on set theory, similarity measurement, or matching algorithms, etc. Organize the records with duplicate attribute patterns into duplicate attribute data, recording the relevant information for each duplicate attribute.
[0060] In this embodiment, refer to Figure 3 As described, it is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include:
[0061] Step S21: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data;
[0062] Step S22: Perform correlation analysis on the duplicate clustering data to generate correlation data;
[0063] Step S23: Construct a matrix for the correlation data to generate a correlation matrix;
[0064] Step S24: Perform missing value statistics on the duplicate attribute data based on the correlation data to generate missing value data;
[0065] Step S25: Perform omission analysis on the correlation matrix based on the missing value data to generate omission data;
[0066] Step S26: Perform sensitivity analysis on the duplicate attribute data based on the omission data to generate sensitivity data;
[0067] Step S27: Based on the sensitive data, use the repeated attribute importance weight calculation formula to calculate the importance weights of the repeated attribute data and generate importance weight data.
[0068] Through the hierarchical clustering algorithm, the present invention can cluster repeated attributes with similar characteristics and group them together, which helps to understand the similarities and correlations between attributes in the data and provides a more accurate and meaningful data basis for subsequent analysis. The generation of repeated clustered data enables better organization and management of repeated attribute data, classifying and labeling them according to the clustering results. In addition, by visualizing the clustering results, the relationships between attributes and the clustering structure can be more intuitively displayed, deeply understanding the internal patterns and structures of the data. Through the correlation analysis of repeated clustered data, the correlation degree between different attributes can be calculated and evaluated, which helps to discover the associations and dependencies between attributes, understand their interactions and influences, and provide key information for subsequent data processing and decision-making. The generation of correlation data enables better understanding of the relationships and mutual influences between attributes. By analyzing the correlations, it can be determined which attributes have strong correlations and which have weak correlations, so as to select and process the data, identify key attributes and important factors. By constructing a matrix based on the correlation data, the correlation information can be organized and stored in the form of a matrix. The correlation matrix provides a clear structured representation, enabling more convenient viewing and analysis of the correlation relationships between attributes. The generation of the correlation matrix allows the visualization of the associations between attributes in the form of a matrix. By observing the patterns and structures of the matrix, the relationships between attributes can be analyzed more deeply, discovering potential patterns and rules, and providing useful information for subsequent data processing and decision-making. By performing missing value statistics based on the correlation data, the missing value situation in the repeated attribute data can be determined, which helps to identify and mark the missing values in the data, understand the integrity and availability of the data, and provide a basis for subsequent data processing and decision-making. The generation of missing value data enables data cleaning and filling operations, filling in the missing values or adopting appropriate processing strategies, which helps to improve the integrity and accuracy of the data and avoid adverse effects caused by missing values in subsequent analysis and applications. By performing omission analysis on the correlation matrix based on the missing value data, the omission data situation in the correlation matrix can be determined, which helps to discover possible missing association relationships in the correlation matrix, understand the integrity and accuracy of the data. The generation of omission data enables the identification and repair of omission problems in the correlation matrix, supplementing or correcting possibly missing correlation information, which helps to improve the integrity and accuracy of the data and provide a more reliable data basis for subsequent analysis and decision-making. By performing sensitivity analysis on the repeated attribute data based on the omission data, the sensitivity of the attributes in the data to the omission data can be evaluated, which helps to determine which attributes have important impacts on the integrity and accuracy of the data, focusing on and processing sensitive attributes. The generation of sensitivity data can identify and solve sensitivity problems in the repeated attribute data, taking corresponding measures to improve the quality and integrity of the data.By performing data correction, completion, or rectification on sensitive attributes, the sensitivity issues in the data can be reduced, and the reliability and accuracy of the data can be improved. Based on the sensitive data and the importance weight calculation formula for duplicate attributes, the importance weights of duplicate attribute data can be calculated, which helps to determine the importance degree of each attribute in data analysis and decision-making, providing a weight reference for subsequent data processing and decision-making. The generation of importance weight data enables the analysis and understanding of the weight distribution and importance ranking of different attributes in the data. By analyzing the importance weight data, it can be determined which attributes have a relatively high importance for data analysis and decision-making, so as to conduct subsequent data processing and decision-making more targeted.
[0069] In this embodiment, an appropriate hierarchical clustering algorithm is selected, such as the Agglomerative Hierarchical Clustering algorithm or the Divisive Hierarchical Clustering algorithm, etc. The selected hierarchical clustering algorithm is used to perform clustering analysis on the duplicate attribute data. According to the requirements of the algorithm, an appropriate clustering metric method and linkage criterion are selected. The clustering results are organized into duplicate clustering data, and the duplicate attribute data in each clustering cluster is recorded. According to the missing value data and the correlation matrix, the possible omission situations in the correlation matrix are analyzed. Check whether there are unfilled missing values or uncalculated correlation items in the correlation matrix. The omission analysis results are organized into omission data, and the omission situations in the correlation matrix are recorded. According to the omission data, the sensitivity in the duplicate attribute data is analyzed. Check whether there are sensitive attributes corresponding to the omission data in the duplicate attribute data, that is, some attributes are important in the correlation matrix but are not calculated or filled. The sensitivity analysis results are organized into sensitivity data, and the relevant information of each sensitive attribute is recorded. According to the requirements and analysis objectives, an appropriate importance weight calculation formula for duplicate attributes is selected, such as the weight calculation formula based on correlation and missing values. According to the selected calculation formula, the importance weights are calculated for the sensitive attributes. The importance weight calculation results are organized into importance weight data, and the importance weight values of each duplicate attribute are recorded.
[0070] In this embodiment, the importance weight calculation formula for duplicate attributes in step S24 is specifically:
[0071]
[0072] Where, W is the importance weight value of the duplicate attribute, i is the i-th duplicate attribute data, n is the total number of duplicate attribute data, x is the frequency difference value of the duplicate attribute data, c is the sensitivity of the duplicate attribute data, b is the correlation coefficient of the duplicate attribute data, f is the importance weight adjustment factor, and g is the frequency value of the duplicate attribute data.
[0073] The present invention can smoothly adjust the frequency difference value x. When the frequency difference value is small, the ln function will amplify its influence and make it more significant. When the frequency difference value is large, the ln function will suppress its influence to ensure that its impact on the calculation of the importance weight is not too significant. Calculating the logarithm value can narrow the range of data items, making the weight calculation result more stable. By taking the logarithm, larger values can be shrunk and smaller values can be amplified, thus better expressing the importance of attributes. By calculating the second derivative of the attribute importance function with respect to the frequency difference value, the change in the importance of the attribute can be further reflected. The second derivative can indicate the rate of change and curvature of the attribute importance function, which helps to determine the degree of importance of the attribute. Dividing the second derivative by the importance weight adjustment factor and the frequency value can normalize and adjust the attribute importance. This can balance the importance of different attributes and ensure the comparability of the calculation results. Summing and squaring the normalized and adjusted attribute importance weights can obtain the final repeated attribute importance weight value. This value comprehensively considers the importance of all attributes and can help determine the ranking and relative importance of the attributes.
[0074] In this embodiment, the specific steps of step S26 are as follows:
[0075] Step S261: Perform data density detection on the repeated attribute data based on the missing data to generate data density data;
[0076] Step S262: Perform natural semantic analysis on the repeated attribute data based on the data density data to generate natural semantic data;
[0077] Step S263: Identify the data content of the repeated attribute data through the natural semantic data to generate a data content weight value;
[0078] Step S264: Conduct an analysis of the influence of the repeated range on the data content weight value to generate range influence data;
[0079] Step S265: Use the range influence data to evaluate the data consistency of the repeated attribute data to generate consistency measurement data;
[0080] Step S266: Conduct an analysis of the deviation from business requirements on the consistency measurement data to generate requirement deviation data;
[0081] Step S267: Perform a sensitivity measurement on the repeated attribute data based on the requirement deviation data to generate a sensitivity measurement value;
[0082] Step S268: Conduct a sensitivity analysis on the repeated attribute data according to the sensitivity measurement value to generate sensitivity data.
[0083] The present invention detects the data density of duplicate attribute data by analyzing the omission situation in the data. By understanding the omission situation of the data, it can be determined which attributes have duplicate data, and then data density data is generated. This helps to identify duplicate patterns in the data and provides basic information for subsequent analysis and processing. The data density of duplicate attribute data is detected by analyzing the omission situation in the data. By understanding the omission situation of the data, it can be determined which attributes have duplicate data, and then data density data is generated. This helps to identify duplicate patterns in the data and provides basic information for subsequent analysis and processing. By using natural semantic data, the data content of duplicate attribute data is identified and a weight value is assigned to it. By deeply analyzing the content of duplicate attribute data, its importance and impact degree can be determined, and then a data content weight value is generated. This helps to distinguish key information from secondary information in duplicate attribute data and provides a weight index for subsequent impact analysis. By analyzing the weight value of the data content, the influence range of duplicate attribute data on other data or related businesses can be determined. This helps to evaluate the propagation degree and impact degree of duplicate attribute data in the entire data set and provides a basis for subsequent consistency evaluation and sensitivity analysis. By analyzing the range influence of duplicate attribute data, its influence degree on data consistency can be evaluated. This helps to identify problems of duplicate attribute data in terms of data consistency and generate corresponding metrics for evaluating the level of data consistency. By analyzing the range influence of duplicate attribute data, its influence degree on data consistency can be evaluated. This helps to identify problems of duplicate attribute data in terms of data consistency and generate corresponding metrics for evaluating the level of data consistency. By analyzing the requirement deviation data, the sensitivity of duplicate attribute data to business decisions or analysis results can be evaluated. This helps to determine the importance degree of duplicate attribute data to decision results or analysis conclusions and generate corresponding sensitivity metrics. By analyzing the sensitivity metrics, the sensitivity degree of duplicate attribute data to business decisions, analysis or results can be determined. This helps to identify the influence degree of duplicate attribute data on key business indicators or decisions and generate corresponding sensitivity data, providing a basis and reference for data processing, decision-making, etc.
[0084] In this embodiment, by analyzing the data, the attributes of data duplication are determined. Then, it is checked whether there is missing data in the data, that is, some attribute values are missing in some data records. By statistically calculating the data density of each attribute, that is, the occurrence frequency of the attribute value in the entire dataset, data density data is generated. Using the data density data, natural language processing technology is applied to the duplicate attribute data. By analyzing and understanding the semantics of each attribute value, data with natural semantics can be generated. For example, for endometrial blood data, techniques such as part-of-speech tagging and entity recognition are used to determine its meaning and generate corresponding natural semantic descriptions. By analyzing and statistically calculating the content of each attribute value, its importance and weight can be determined. This can be achieved by using techniques such as text mining and machine learning. The generated data content weight value represents the contribution degree of each attribute value to the overall data content. For example, in endometrial blood data, "endometrium" should have a higher contribution degree to the overall data name, while "blood data" has a lower contribution degree to the overall data name because the impact of "blood data" on the overall data content is smaller, and "endometrium" can basically define the type and attributes of a data. For different duplicate attribute data, according to the weight value of its data content, its importance and influence scope in the overall data are evaluated. The generated scope influence data represents the influence degree of each duplicate attribute data on the overall data. Using the scope influence data, a consistency evaluation is performed on the duplicate attribute data. By comparing the scope influence degrees between different duplicate attribute data, the consistency degree between the data can be determined. The generated consistency metric data represents the consistency degree of each group of duplicate attribute data. For example, in endometrial blood data, the consistency degree of "endometrium" is higher than that of "blood data". The absence of the words "endometrium" will cause a serious change in the semantics of the entire data, while the absence of any word in "blood data" will have a certain impact on the semantics of the entire data, but it is far less than the impact caused by the absence of "endometrium". A business requirement deviation analysis is performed on the consistency metric data. By comparing the consistency metric data with the set business requirements, the deviation degree of the data is determined. The generated requirement deviation data represents the difference degree between each group of duplicate attribute data and the business requirements. Using the requirement deviation data, a sensitivity measurement is performed on the duplicate attribute data. According to the deviation degree between the data and the business requirements, the sensitivity of the data is evaluated. The generated sensitivity measurement value represents the sensitivity degree of each group of duplicate attribute data. For example, in endometrial blood data, the sensitivity of "endometrium" is higher than that of "blood data" because "blood data" may appear in multiple data names, such as "patient blood data" and "ovarian cyst blood data", while the appearance of the data name "endometrium" basically defines the semantics of the entire data and is restricted to the endometrial part. According to the sensitivity degree of the data, the corresponding sensitivity data is determined.This data can be used for further data processing and the decision-making process.
[0085] In this embodiment, referring to Figure 4 as described, it is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include:
[0086] Step S31: Perform threshold impact analysis on the duplicate attribute data based on the importance weight data to generate impact analysis data;
[0087] Step S32: Perform threshold analysis on the duplicate attribute data based on the impact analysis data to generate threshold data;
[0088] Step S33: Define the similarity threshold for the threshold data to generate similarity threshold data;
[0089] Step S34: Calculate the similarity of the duplicate attribute data using the duplicate attribute data similarity calculation formula based on the similarity threshold data to generate similarity data.
[0090] Through threshold impact analysis based on importance weight data, the present invention can evaluate the impact degree of different thresholds on repeated attribute data, which helps to determine appropriate threshold settings to control the sensitivity and accuracy in the data processing and decision-making processes. Through threshold impact analysis, impact analysis data can be generated, which records the changes in attribute data under different threshold conditions. These data can be used for subsequent threshold analysis and decision-making, and to understand the impact of different threshold settings on data results. Through threshold analysis based on the impact analysis data, an appropriate threshold range or specific threshold can be determined to control the sensitivity and accuracy in the data processing and decision-making processes. Threshold analysis determines when the value of a certain attribute is regarded as important or unimportant, thus affecting subsequent data processing and decision results. The result of threshold analysis is to generate threshold data, which records the threshold settings of different attributes and the corresponding data changes. The threshold data can be used for subsequent similarity threshold definition and similarity calculation to determine the discrimination criteria and threshold settings for similarity. Through similarity threshold definition for the threshold data, the threshold settings for determining the similarity of repeated attribute data can be determined. Similarity threshold definition determines what degree of difference in attribute values is considered similar, thus affecting subsequent similarity calculation and data analysis. The result of similarity threshold definition is to generate similarity threshold data, which records the similarity threshold settings between different attributes. The similarity threshold data can be used for subsequent similarity calculation to determine the discrimination criteria and threshold settings for similarity. Based on the similarity threshold data and the similarity calculation formula for repeated attribute data, the similarity of repeated attribute data can be calculated. Similarity calculation measures the similarity degree between different attributes, thus identifying similar data instances or attribute patterns. The result of similarity calculation is to generate similarity data, which records the similarity values between different attributes. The similarity data can be used for subsequent data analysis and decision-making to identify and utilize similarity information for tasks such as data classification, clustering, or recommendation.
[0091] In this embodiment, for each piece of duplicate attribute data, threshold impact analysis is performed on the dataset. This may involve statistical analysis of the duplicate attribute data and other associated attributes, such as calculating the correlation between attributes, importance scores, etc. Based on the analysis results, impact analysis data is generated, which may include the relationships between each piece of duplicate attribute data and other attributes, the degree of impact of importance weights, etc. Using the impact analysis data, threshold analysis is performed on the duplicate attribute data. This can determine the setting of the threshold based on the relationships between different attributes and the impact of importance weights. Using the similarity threshold data, a suitable similarity calculation formula is selected to calculate the similarity of the duplicate attribute data. According to the calculation results, similarity data is generated, which includes the similarity values of each pair of duplicate attribute data. According to the threshold data, the similarity threshold of the duplicate attribute data is defined. This can be determined according to business requirements and data characteristics. The definition of the similarity threshold may involve weight settings between different attributes, the selection of similarity measurement methods, etc. According to the defined similarity threshold, similarity threshold data is generated, which includes the similarity threshold set for each piece of duplicate attribute data. Using the similarity threshold data, a suitable similarity calculation formula is selected to calculate the similarity of the duplicate attribute data. According to the calculation results, similarity data is generated, which includes the similarity values of each pair of duplicate attribute data.
[0092] In this embodiment, the similarity calculation formula for the duplicate attribute data in step S34 is specifically as follows:
[0093]
[0094] Where S is the similarity of the duplicate attribute data, λ is the similarity threshold, i is the i-th piece of duplicate attribute data, n is the total number of duplicate attribute data, w i is the weight of the i-th piece of duplicate attribute data, p is the degree of threshold impact, ∈ is the average similarity of the duplicate attribute data, c is the sensitivity of the duplicate attribute data, γ is the similarity calculation adjustment factor, g is the frequency value of the duplicate attribute data, and μ is the difference of the duplicate attribute data.
[0095] The present invention By summing the weights and logarithms of each piece of duplicate attribute data, the importance and similarity of each data item can be comprehensively considered. The logarithmic function can adjust the range of the data, making the calculation results more stable and reliable. Dividing the sum of the weights and logarithms by the total number of duplicate attribute data and the sensitivity can obtain a measure of similarity. This value comprehensively considers the importance, similarity, and sensitivity of each data item, which helps to evaluate the overall similarity of the duplicate attribute data. Through the calculation of the exponential function, the measured value of similarity can be mapped to a range from 0 to 1. The exponential function maps the exponential value of a negative number to a value between (0, 1), making the calculation result more in line with the definition of similarity. γ·(g - μ) 2 Calculating the square of the difference between the frequency value and the average value and multiplying it by the similarity calculation adjustment factor can further adjust the similarity. This term takes into account the influence of frequency differences on similarity, making the calculation result of similarity more accurate and comprehensive. The formula considers importance and similarity, adjusts the similarity value, and takes into account frequency differences, etc., and can accurately evaluate the similarity of duplicate attribute data, providing an important reference basis for duplicate removal decisions and data analysis.
[0096] In this embodiment, step S4 includes the following steps:
[0097] Step S41: Perform time series analysis on the duplicate attribute data based on the similarity data to generate time series analysis data;
[0098] Step S42: Perform time slicing processing on the time series analysis data to generate time series slice data;
[0099] Step S43: Perform dynamic time window analysis on the time series slice data to generate dynamic time window data;
[0100] Step S44: Perform periodic analysis on the duplicate attribute data based on the dynamic time window data to generate periodic data;
[0101] Step S45: Define the time granularity for the periodic data to generate time granularity data.
[0102] Through the time series analysis of repeated attribute data, the present invention can understand the changing trends and patterns of the data over time. Time series analysis helps to reveal characteristics such as the periodicity, trend, and seasonality of the data, thereby understanding the evolution and variation laws of the data. The result of time series analysis is the generation of time series analysis data, which records the relevant information of the repeated attribute data changing over time. The time series analysis data can be used for subsequent time slicing processing and dynamic time window analysis to further explore the time-dependent relationship and periodic characteristics of the data. By performing time slicing processing on the time series analysis data, continuous time series data can be divided into discrete time segments, which helps to transform complex time series data into time segment data that is easy to process and analyze, facilitating subsequent dynamic time window analysis and periodic analysis. The result of time slicing processing is the generation of time slice data, which records the divided time segments and their corresponding attribute data. The time slice data can be used for subsequent dynamic time window analysis and periodic analysis to further study the time correlation and periodicity of the data. By performing dynamic time window analysis on the time slice data, the characteristics and behaviors of the data can be observed and compared within different time ranges. Dynamic time window analysis helps to discover local patterns, trends, and anomalies in the data, thereby better understanding the dynamic changes of the data. The result of dynamic time window analysis is the generation of dynamic time window data, which records the attribute data situation under different time windows. The dynamic time window data can be used for subsequent periodic analysis and time granularity definition to reveal the periodic characteristics and time scales of the data. By performing periodic analysis based on the dynamic time window data, periodic patterns and cycle lengths in the data can be identified and extracted. Periodic analysis helps to reveal the repetitive behavior and periodic characteristics of the data, thereby understanding the periodic changes and trends of the data. The result of periodic analysis is the generation of periodic data, which records the periodic patterns and cycle lengths of different attributes. The periodic data can be used for subsequent time granularity definition and analysis, identifying and utilizing the periodic information of the data to perform tasks such as periodic prediction, adjustment, and optimization. By performing time granularity definition on the periodic data, the time granularity for data analysis and processing can be determined. The time granularity definition determines the time scale of the data, thereby affecting subsequent data analysis, prediction, and decision-making. Different time granularities can provide data views at different levels and precisions, better understanding the evolution and variation laws of the data. The result of time granularity definition is the generation of time granularity data, which records the attribute data situation under different time granularities. The time granularity data can be used for subsequent data analysis, visualization, and decision-making, and select an appropriate time scale for data interpretation and application as needed.
[0103] In this embodiment, using the similarity data, time series analysis is performed on the duplicate attribute data. Time series analysis may involve the change trend of attribute values over time, periodicity, etc. According to the analysis results, time series analysis data is generated, including the attribute values of each duplicate attribute data at different time points. The time series analysis data is processed by time slicing, dividing the entire time range into several consecutive time segments. Each time segment can be a fixed time interval or set according to the characteristics and requirements of the data. For each time segment, the time series analysis data therein is recorded, that is, the attribute values of each duplicate attribute data within the time segment. The time series slice data is subjected to dynamic time window analysis. A dynamic time window refers to analyzing the data within a given time range according to a certain window size and sliding step. Appropriate window size and sliding step are selected, and according to these parameters, the time series slice data is divided into multiple time windows. Within each time window, the duplicate attribute data is analyzed, such as calculating the average value, maximum value, and minimum value of the attribute. Using the dynamic time window data, periodic analysis is performed on the duplicate attribute data. Periodic analysis can reveal the periodic changes existing in the data, and methods of time series analysis, such as Fourier transform, autocorrelation function, etc., can be applied to detect and determine the periodicity. According to the analysis results, periodic data is generated, including information such as the start time, end time, and period length of the period. The time granularity of the periodic data is defined, that is, according to the business requirements and data characteristics, the time unit of the periodic data is determined, such as hours, days, weeks, months, etc. The periodic data is subjected to time granularity conversion so that the periodic data is recorded according to the defined time unit, generating time granularity data, including the time granularity unit and the corresponding periodic data.
[0104] In this embodiment, step S5 includes the following steps:
[0105] Step S51: Perform duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data;
[0106] Step S52: Perform misjudgment detection on the duplicate removal decision data to generate misjudgment detection data;
[0107] Step S53: Perform duplicate removal accuracy analysis on the duplicate removal decision data based on the misjudgment detection data to generate accuracy data;
[0108] Step S54: Use the accuracy data to optimize the strategy of the duplicate removal decision data to generate a duplicate removal optimization strategy.
[0109] Based on the time-granularity data, the present invention can perform duplicate removal decision analysis on duplicate attribute data. The goal of the duplicate removal decision analysis is to determine which duplicate attribute data should be retained and which should be removed to ensure the accuracy and consistency of the data. The result of the duplicate removal decision analysis is to generate duplicate removal decision data, which records the duplicate removal decision results of each duplicate attribute data item. The duplicate removal decision data can be used for subsequent misjudgment detection and accuracy analysis to evaluate the effectiveness and accuracy of the duplicate removal decision. By performing misjudgment detection on the duplicate removal decision data, possible misjudgment situations can be identified, that is, data items that are wrongly determined to be duplicates or data items that are wrongly determined not to be duplicates. Misjudgment detection helps to evaluate the accuracy and reliability of the duplicate removal decision and to discover possible misjudgment problems. The result of the misjudgment detection is to generate misjudgment detection data, which records the misjudgment situations of each duplicate attribute data item. The misjudgment detection data can be used for subsequent duplicate removal accuracy analysis and strategy optimization to better understand the accuracy and performance of the duplicate removal decision and to make corresponding improvements and adjustments. Based on the misjudgment detection data, duplicate removal accuracy analysis can be performed on the duplicate removal decision data. The goal of the duplicate removal accuracy analysis is to evaluate the accuracy and performance of the duplicate removal decision, determine indicators such as the misjudgment rate, accuracy rate, and recall rate of the duplicate removal decision, and discover possible accuracy problems and improvement spaces. The result of the duplicate removal accuracy analysis is to generate accuracy data, which records the accuracy indicators and evaluation results of the duplicate removal decision. The accuracy data can be used for subsequent strategy optimization and decision-making to understand the performance status of the duplicate removal decision and to make corresponding adjustments and improvements according to requirements. Based on the accuracy data, statistical and analytical methods can be used to optimize the duplicate removal decision data. The goal of the strategy optimization is to improve the accuracy and performance of the duplicate removal decision, reduce the misjudgment rate, and increase the accuracy rate. By analyzing the accuracy data, problems can be discovered, improvement directions can be determined, and corresponding optimization strategies and rules can be formulated. The result of the strategy optimization is to generate duplicate removal optimization strategies, which include improved duplicate removal decision rules, adjusted parameter settings, optimized algorithms, etc. The duplicate removal optimization strategies can be applied to the actual data duplicate removal process to improve the effectiveness and accuracy of duplicate removal and reduce misjudgments and errors.
[0110] In this embodiment, time granularity data is used to perform duplicate removal decision analysis on duplicate attribute data. The goal of duplicate removal decision analysis is to determine which one should be selected as a representative from each set of duplicate attribute data for duplicate removal operations. According to the time granularity data and business requirements, some decision-making bases can be considered, such as selecting the data with the earliest timestamp, selecting the latest data, selecting based on attribute similarity, etc. According to the analysis results, duplicate removal decision data is generated, that is, the representative data selected from each set of duplicate attribute data is determined. Misjudgment detection is performed on the duplicate removal decision data, aiming to evaluate the accuracy of the duplicate removal decision and discover possible misjudgment situations. By comparing the differences between the selected representative data and other duplicate attribute data, misjudgment detection can be carried out. The differences can include changes in attribute values, outliers, etc. According to the detection results, misjudgment detection data is generated to identify possible misjudgment situations. Using the misjudgment detection data, duplicate removal accuracy analysis is performed on the duplicate removal decision data to evaluate the accuracy of the duplicate removal decision. According to the misjudgment situations identified in the misjudgment detection data, accuracy indicators of the duplicate removal decision are calculated, such as accuracy rate, recall rate, etc., and accuracy data is generated to record the accuracy indicators and related information of the duplicate removal decision. According to the accuracy data, the duplicate removal decision strategy is optimized. According to the accuracy indicators, problems existing in the duplicate removal decision are discovered and improvement strategies are proposed. Ways such as adjusting the basis of the duplicate removal decision, considering more attribute factors, and introducing machine learning algorithms can be tried for optimization. According to the optimization results, a duplicate removal optimization strategy is generated, including updated decision rules, adjusted weight parameters, etc.
[0111] In this embodiment, step S6 includes the following steps:
[0112] Step S61: Use the circular convolution algorithm to perform convolution preprocessing on the duplicate removal optimization strategy to generate a convolution sample set;
[0113] Step S62: Perform convolution data cutting on the convolution sample set to generate a convolution sequence;
[0114] Step S63: Perform dilated convolution on the convolution sequence to generate a duplicate removal convolution network;
[0115] Step S64: Perform pooling multi-layer sampling on the duplicate removal convolution network to generate a convolution feature map;
[0116] Step S65: Perform data mining modeling on the convolutional feature map to construct a duplicate removal policy model for precise data duplicate removal. In the present invention, the duplicate removal optimization policy is preprocessed by a circular convolution algorithm. The optimization policy can be applied to different parts of the dataset in a sliding window manner, which can extract local features in the dataset and capture the correlation and similarity between data items. The result of the convolutional preprocessing is to generate a convolutional sample set, which contains the optimized policy samples after convolutional processing. The convolutional sample set can be used for subsequent convolutional data cutting and dilated convolution, providing input data for constructing the duplicate removal policy model. Performing convolutional data cutting on the convolutional sample set can divide the sample set into multiple consecutive sequences, which can preserve the order relationship between data items and provide input sequences for subsequent dilated convolution. The result of the convolutional data cutting is to generate a convolutional sequence, which contains the divided consecutive sequences. The convolutional sequence can be used for dilated convolution operations to extract higher-level features and provide richer information for constructing the duplicate removal policy model. Performing dilated convolution operations on the convolutional sequence can expand the receptive field of the convolutional kernel and capture a wider context. Dilated convolution can extract the long-range dependence relationship between data items and enhance the feature representation ability. The result of the dilated convolution is to generate a duplicate removal convolutional network, which contains the convolutional sequence after dilated convolution operations. The duplicate removal convolutional network can be used for subsequent pooling multi-layer sampling to extract higher-level features and provide a more representative feature representation for constructing the duplicate removal policy model. Performing data mining modeling on the convolutional feature map can analyze the patterns, trends, and correlations in the feature map. By using machine learning, deep learning, or other data mining techniques, useful information can be extracted from the feature map and a duplicate removal policy model can be constructed. Through the process of data mining modeling, a duplicate removal policy model can be constructed. This model can learn the similarity and difference between data items based on the features in the feature map to determine which data is duplicate. By constructing an accurate duplicate removal policy model, precise data duplicate removal can be performed to remove duplicate data items and improve data quality and analysis effects.
[0117] In this embodiment, for the duplicate removal optimization strategy, a circular convolution algorithm is used for convolution preprocessing. Circular convolution is an algorithm that convolves an input sequence with a weight sequence. The duplicate removal optimization strategy is transformed into an input sequence suitable for the circular convolution algorithm, and a corresponding weight sequence is defined. The circular convolution algorithm is used to perform a convolution operation on the input sequence and the weight sequence to obtain a convolution sample set. The convolution data of the convolution sample set is cut, and the data in the sample set is divided into multiple consecutive subsequences according to a certain size. The size of the divided subsequences can be determined according to the actual situation and algorithm requirements. Usually, a suitable window size is selected for cutting, and each subsequence is called a convolution sequence for subsequent dilated convolution operations. The dilated convolution operation is performed on the divided convolution sequences to generate a duplicate removal convolution network. Dilated convolution is a convolution operation that introduces a dilation rate, which can expand the receptive field of the convolution kernel and capture more context information. The dilated convolution is applied to each convolution sequence, and the obtained feature maps are connected to form a duplicate removal convolution network to extract more rich feature information. A multi-layer pooling operation is performed on the duplicate removal convolution network to reduce the dimension of the feature map and extract more representative feature information. Different pooling operations, such as max pooling or average pooling, can be used to downsample the features in the duplicate removal convolution network. Through multiple pooling operations, a convolution feature map with a lower dimension is obtained for subsequent data mining and modeling. The convolution feature map is used for data mining and modeling to construct a duplicate removal strategy model. Machine learning algorithms, such as neural networks and support vector machines, can be used to train and classify the convolution feature map. Based on the constructed duplicate removal strategy model, the data to be de-duplicated is classified and judged, and accurate data de-duplication operations are performed.
[0118] In this embodiment, a precise duplicate removal system based on obstetrics and gynecology data collection is provided, including:
[0119] An information collection module, which acquires obstetrics and gynecology data; performs data frequency statistics on the obstetrics and gynecology data to generate frequency data; performs duplicate attribute analysis on the frequency data to generate duplicate attribute data;
[0120] A weight analysis module, which uses the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; performs sensitivity analysis on the duplicate clustering data to generate sensitivity data; calculates the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data;
[0121] A similarity calculation module, which defines a similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data; calculates the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data;
[0122] The time window module performs dynamic time window analysis on duplicate attribute data based on similarity data to generate dynamic time window data; defines the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data;
[0123] The duplicate removal decision module performs duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data; optimizes the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy;
[0124] The strategy model module performs dilated convolution on the duplicate removal optimization strategy using a circular convolution algorithm to construct a duplicate removal strategy model for performing precise data duplicate removal.
[0125] The present invention obtains obstetrics and gynecology data through an information collection module, including medical records, medical images, laboratory results, etc. These data can be used for subsequent data analysis and processing. By performing data frequency statistics on the obstetrics and gynecology data, the frequency of each data item can be calculated. The frequency data reflects the importance and universality of the data item in the sample set. By analyzing the frequency data, attributes with repeated occurrences can be identified. The repeated attribute data shows which attributes have similar values or patterns in different data items. Using the hierarchical clustering algorithm to perform clustering analysis on the repeated attribute data, data items with similar characteristics can be grouped into the same category. The repeated clustering data provides the clustering results between data items, helping to identify groups of data items with similar characteristics. By performing sensitivity analysis on the repeated clustering data, the sensitivity of different data items to the clustering results can be evaluated. The sensitivity data shows the contribution degree of each data item to the clustering results, helping to determine the attributes that mainly affect the clustering results. Based on the sensitivity data, importance weight calculation is performed on the repeated attribute data. The importance weight data reflects the contribution degree of each attribute to the similarity of data items, helping to determine the importance ranking of attributes. Based on the importance weight data, a similarity threshold is defined for the repeated attribute data. The similarity threshold data determines which data items are considered similar and groups them as candidates for duplicate data. Based on the similarity threshold data, similarity calculation is performed on the repeated attribute data. The similarity data quantifies the similarity degree between different data items, helping to identify true duplicate data items. By performing dynamic time window analysis on the repeated attribute data, the temporal correlation of data items can be determined. The dynamic time window data shows the temporal relationship and temporal sequence pattern between data items. Based on the dynamic time window data, the time granularity is defined for the repeated attribute data. The time granularity data determines the granularity size of data items in time, helping to determine the time span of the duplicate removal decision. By performing duplicate removal decision analysis on the repeated attribute data, the duplicate removal decision data determines which duplicate data items should be retained or deleted, helping to clean up the duplicate items in the dataset. By performing strategy optimization on the duplicate removal decision data, more factors such as data quality and business requirements can be considered to generate a more reasonable and effective duplicate removal optimization strategy. The optimized strategy can more accurately identify and process duplicate data items. Using the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy, a duplicate removal strategy model is constructed. This model can perform precise data duplicate removal operations according to the input data characteristics and the duplicate removal optimization strategy, effectively eliminating duplicate data items.
[0126] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be embraced within the present invention.
[0127] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0128] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A precise deduplication method based on obstetrics and gynecology data collection, characterized in that, It includes the following steps: Step S1: Obtain obstetrics and gynecology data; conduct data frequency statistics on the obstetrics and gynecology data to generate frequency data; Conduct duplicate attribute analysis on the frequency data to generate duplicate attribute data; Step S2: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; Conduct sensitivity analysis on the duplicate clustering data to generate sensitivity data; Calculate the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data; Step S3: Define the similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data; calculate the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data; Step S4: Conduct dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data; Define the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data; Step S5: Conduct duplicate removal decision analysis on the duplicate attribute data according to the time granularity data to generate duplicate removal decision data; optimize the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy; Step S6: Use the circular convolution algorithm to perform dilated convolution on the duplicate removal optimization strategy to construct a duplicate removal strategy model for precise data duplicate removal.
2. The method according to claim 1, characterized in that The specific steps of Step S1 are as follows: Step S11: Obtain obstetrics and gynecology data, where the obstetrics and gynecology data includes patient information data, case information, diagnosis results, drug information, and pregnancy and childbirth information; Step S12: Extract features from the obstetrics and gynecology data to generate feature data; Step S13: Conduct data frequency statistics on the feature data to generate frequency data; Step S14: Conduct attribute comparison on the frequency data to generate attribute comparison data; Step S15: Conduct frequency difference analysis on the attribute comparison data to generate frequency difference data; Step S16: Conduct duplicate attribute analysis on the frequency difference data to generate duplicate attribute data.
3. The method according to claim 1, characterized in that The specific steps of Step S16 are as follows: Step S161: Conduct duplicate attribute identification on the frequency difference data to generate attribute identification data; Step S162: Conduct outlier detection on the frequency data based on the attribute identification data to generate outlier data; Step S163: Conduct duplicate attribute analysis on the frequency data according to the outlier data to generate duplicate attribute data.
4. The method according to claim 1, wherein The specific steps of Step S2 are as follows: Step S21: Use the hierarchical clustering algorithm to perform clustering analysis on the duplicate attribute data to generate duplicate clustering data; Step S22: Conduct correlation analysis on the duplicate clustering data to generate correlation data; Step S23: Construct a matrix for the correlation data to generate a correlation matrix; Step S24: Conduct missing value statistics on the duplicate attribute data based on the correlation data to generate missing value data; Step S25: Conduct omission analysis on the correlation matrix based on the missing value data to generate omission data; Step S26: Conduct sensitivity analysis on the duplicate attribute data based on the omission data to generate sensitivity data; Step S27: Calculate the importance weight of the duplicate attribute data based on the sensitivity data using the duplicate attribute importance weight calculation formula to generate importance weight data; Among them, the specific formula for calculating the importance weight of repeated attributes in step S27 is as follows: Among them, W is the importance weight value of repeated attributes, i is the i-th repeated attribute data, n is the total number of repeated attribute data, x is the frequency difference value of repeated attribute data, c is the sensitivity of repeated attribute data, b is the correlation coefficient of repeated attribute data, f is the importance weight adjustment factor, and g is the frequency value of repeated attribute data.
5. The method according to claim 4, wherein The specific steps of step S26 are as follows: Step S261: Perform data density detection on repeated attribute data based on missing data to generate data density data; Step S262: Perform natural semantic analysis on repeated attribute data based on the data density data to generate natural semantic data; Step S263: Identify the data content of repeated attribute data through natural semantic data to generate a data content weight value; Step S264: Analyze the influence of the repeated range on the data content weight value to generate range influence data; Step S265: Use the range influence data to evaluate the data consistency of repeated attribute data to generate consistency measurement data; Step S266: Analyze the deviation of business requirements from the consistency measurement data to generate requirement deviation data; Step S267: Perform sensitivity measurement on repeated attribute data based on the requirement deviation data to generate a sensitivity measurement value; Step S268: Perform sensitivity analysis on repeated attribute data according to the sensitivity measurement value to generate sensitivity data.
6. The method according to claim 4, characterized in that The specific steps of step S3 are as follows: Step S31: Analyze the threshold influence on repeated attribute data based on the importance weight data to generate influence analysis data; Step S32: Perform threshold analysis on repeated attribute data based on the influence analysis data to generate threshold data; Step S33: Define the similarity threshold for the threshold data to generate similarity threshold data; Step S34: Calculate the similarity of repeated attribute data based on the similarity threshold data using the similarity calculation formula for repeated attribute data to generate similarity data. Among them, the specific formula for calculating the similarity of repeated attribute data in step S34 is as follows: Among them, S is the similarity of duplicate attribute data, λ is the similarity threshold, i is the i-th duplicate attribute data, n is the total number of duplicate attribute data, w i is the weight of the i-th duplicate attribute data, p is the degree of influence of the threshold, ∈ is the average similarity of duplicate attribute data, c is the sensitivity of duplicate attribute data, γ is the similarity calculation adjustment factor, g is the frequency value of duplicate attribute data, and μ is the difference of duplicate attribute data.
7. The method according to claim 1, wherein The specific steps of step S4 are as follows: Step S41: Perform time series analysis on repeated attribute data based on the similarity data to generate time series analysis data; Step S42: Perform time slicing on the time series analysis data to generate time series slice data; Step S43: Perform dynamic time window analysis on the time series slice data to generate dynamic time window data; Step S44: Perform periodic analysis on repeated attribute data based on the dynamic time window data to generate periodic data; Step S45: Define the time granularity for the periodic data to generate time granularity data.
8. The method according to claim 1, wherein The specific steps of step S5 are as follows: Step S51: Perform duplicate removal decision analysis on repeated attribute data according to the time granularity data to generate duplicate removal decision data; Step S52: Perform misjudgment detection on the duplicate removal decision data to generate misjudgment detection data; Step S53: Analyze the accuracy of duplicate removal for the duplicate removal decision data based on the misjudgment detection data to generate accuracy data; Step S54: Optimize the duplicate removal decision data using the accuracy data to generate a duplicate removal optimization strategy.
9. The method according to claim 1, wherein The specific steps of step S6 are as follows: Step S61: Perform convolution preprocessing on the duplicate removal optimization strategy using the circular convolution algorithm to generate a convolution sample set; Step S62: Perform convolution data cutting on the convolution sample set to generate a convolution sequence; Step S63: Perform dilated convolution on the convolution sequence to generate a duplicate removal convolution network; Step S64: Perform pooling multi-sampling on the duplicate removal convolution network to generate a convolution feature map; Step S65: Perform data mining modeling on the convolution feature map to construct a duplicate removal strategy model for performing precise data duplicate removal.
10. A precise duplicate removal system based on obstetrics and gynecology data collection, characterized in that, For performing the precise duplicate removal method based on obstetrics and gynecology data collection as described in claim 1, including: An information collection module that obtains obstetrics and gynecology data; performs data frequency statistics on the obstetrics and gynecology data to generate frequency data; performs duplicate attribute analysis on the frequency data to generate duplicate attribute data; A weight analysis module that performs clustering analysis on the duplicate attribute data using the hierarchical clustering algorithm to generate duplicate clustering data; performs sensitivity analysis on the duplicate clustering data to generate sensitivity data; calculates the importance weight of the duplicate attribute data based on the sensitivity data to generate importance weight data; A similarity calculation module that defines a similarity threshold for the duplicate attribute data based on the importance weight data to generate similarity threshold data; calculates the similarity of the duplicate attribute data based on the similarity threshold data to generate similarity data; A time window module that performs dynamic time window analysis on the duplicate attribute data based on the similarity data to generate dynamic time window data; defines the time granularity for the duplicate attribute data based on the dynamic time window data to generate time granularity data; A duplicate removal decision module that performs duplicate removal decision analysis on the duplicate attribute data based on the time granularity data to generate duplicate removal decision data; optimizes the strategy for the duplicate removal decision data to generate a duplicate removal optimization strategy; A strategy model module that performs dilated convolution on the duplicate removal optimization strategy using the circular convolution algorithm to construct a duplicate removal strategy model for performing precise data duplicate removal.