One-person, one-file and one-attribution cleaning and attribution method for historical repeated archives of crowds in super-large area
By extracting and integrating the historical archives of residents in the super-large area, forming a unique and complete new archives, and allocating management agencies according to the attribution management rules, the problems of duplication of archives and confusion in management are solved, and one person, one file and one territorial management of residents' archives is realized.
Patent Information
- Application Number
- CN202510410137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In the prior art, there are duplication and incomplete historical archives of people in super-large areas, resulting in confusion in archive management and the inability to realize one-person, one-level management of residents' archives.
By extracting matching historical archives from the archives of all districts and counties in a super-large area based on the resident ID, integrating them into new archives, and allocating management agencies and interfaces according to the attribution management rules, ensuring the uniqueness and integrity of archives.
It has realized the deduplication and integration of residents' files in a huge area, ensured that each resident has a unique and complete file, and clarified the archives' attribution management agency, solving the chaotic problem of archive management.
Smart Images

Figure CN120047294A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file cleaning and attribution. More specifically, the present invention relates to a method for cleaning and attributing historical duplicate files of a large-area population, with one file per person and one jurisdiction per file. Background Art
[0002] In traditional residential file management methods, each district and county has independently established residential files for a long time, and each district and county has an independent residential file database. With the flow of the population, the same person may go to different districts and counties for activities such as going to school, living, working, and traveling, resulting in the same person possibly receiving services and establishing a file separately in different districts and counties. In this way, from the perspective of prefecture-level, provincial-level, or even national-level, there are problems of historical file duplication of a large-area population and multiple duplicate filings of a single resident in different places, resulting in duplicate files, fragmented, incomplete, and discontinuous service content records, thus losing the value of the files. In addition, due to the establishment of files for the same person in different districts and counties, there is also a problem of chaotic file attribution management.
[0003] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for cleaning and attributing historical duplicate files of a large-area population, with one file per person and one jurisdiction per file, to solve the problems raised in the above background art by applying different product inspection methods.
[0005] To achieve the above object, the present invention provides the following technical solution. A method for cleaning and attributing historical duplicate files of a large-area population, with one file per person and one jurisdiction per file, includes: extracting historical files matching the resident ID from the file libraries of all districts and counties within the large area based on the resident ID; integrating all the historical files matching each resident ID to obtain a new file corresponding to each resident ID, and storing the new file; assigning a management agency to the new file of each resident according to the attribution management rules; establishing a management interface for the new file of each resident; and the management agency assigned to the new file manages the new file through the established management interface.
[0006] Preferably, the historical archives include the front page of the historical archives and historical event records; integrating all the historical archives matched with each resident ID to obtain a new file corresponding to each resident ID includes: integrating the historical event records in all the historical archives matched with each resident ID in chronological order of the occurrence time of the events to obtain a complete event record corresponding to each resident ID; integrating the front pages of the historical archives of all the historical archives matched with each resident ID, and constructing a unique file front page corresponding to each resident ID according to the integration result; the complete event record corresponding to each resident ID and the corresponding unique file front page form a new file corresponding to each resident ID.
[0007] The beneficial technical effects of the present invention: In order to solve the problems of the uniqueness and integrity of historical archives, and to ensure the continuity and integrity of resident files in a super-large area and achieve full-life cycle records, the first step is to extract historical archives matched with the resident ID from the archives of all districts and counties in the super-large area based on the resident ID to ensure the uniqueness and integrity of the historical archives (since a large amount of manpower, material and financial resources have been invested in building the archives over the years, if they are not cleaned up and utilized, it will cause huge waste and the loss of residents' historical information); the second step is to integrate all the historical archives matched with each resident ID to obtain a new file corresponding to each resident ID, and assign a management agency to the new file of each resident according to the attribution management rules to ensure the uniqueness and integrity of the file in the new service process; the present invention first cleans up and removes duplicate historical archives of the population in the super-large area and retains one file; secondly, integrates the historical service records of different periods and different regions recorded in the duplicate files of residents in multiple places into one file; thirdly, according to the attribution management rules, clarify the unique management agency of the file. The present invention can clean up duplicate historical archives in super-large areas such as prefecture-level, provincial or national levels, obtain a complete and unique new file for each resident by integrating the historical archives of each resident, and set a unique attribution management agency for each new file, realizing the one-person-one-file-one-territory management of resident files. Brief Description of the Drawings
[0008] Figure 1 It is a flowchart of a method for cleaning up and attributing duplicate historical archives of the population in a super-large area of the present invention.
[0009] Figure 2 It is a flowchart of the redundant data removal process of the file information platform of the present invention. Detailed Embodiments
[0010] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0011] Embodiment 1
[0012] Please refer to Figure 1 , a method for cleaning and attributing historical duplicate files of people in a super-large area with one file per person and one file per jurisdiction, and the specific operation process is as follows:
[0013] Step A1, extract historical files matching the resident ID from the file libraries of all districts and counties within the super-large area based on the resident ID. The resident ID is preferably but not limited to the ID card number or the birth medical certificate number (for newborns). The super-large area represents more than one prefecture-level area or more than one provincial jurisdiction area or a national jurisdiction area.
[0014] Step A2, integrate all historical files matching each resident ID to obtain a new file corresponding to each resident ID, and store the new file. The integration process is not limited to being quickly completed in the following manner to obtain a complete and unique new file corresponding to each resident ID. The historical files include the historical file front page and historical event records. For example, if the resident file is an electronic health record, the historical events include physical examination events, treatment events, inspection events, vaccination events, and diagnosis events, etc., and each event has an event occurrence time. Step A2 includes:
[0015] Step A21, integrate the historical event records in all historical files matching each resident ID in the order of the event occurrence time to obtain a complete event record corresponding to each resident ID.
[0016] Step A22, integrate the historical file front pages of all historical files matching each resident ID, and construct a unique file front page corresponding to each resident ID according to the integration result. The integration process of the historical file front pages of all historical files matching each resident ID mainly includes: integrating all fields with different field names included in all historical file front pages, updating the content of each field to the latest record data of the field, such as setting the residence address field to the residence address in the historical file with the most recent time. Write the content of all fields with different field names after updating for each resident ID into a preset file front page template to obtain a unique file front page corresponding to each resident ID.
[0017] Step A23, the complete event record corresponding to each resident ID and the corresponding unique file front page form the new file corresponding to each resident ID.
[0018] Step A3: Assign a management agency to the new file of each resident according to the attribution management rule. The attribution management rule is to use the file management agency in the district or county where the resident's permanent residence or household registration is located as the management agency for the new file of each resident.
[0019] Step A4: Establish a management interface for the new file of each resident. The management interface is preferably but not limited to an application interface or an Internet link.
[0020] Step A5: The management agency assigned to the new file manages the new file through the established management interface, such as updating, modifying, etc.
[0021] In a preferred implementation manner of this example, the new files in the file information platform are stored in a manner divided into more than one data packet. As the running time of the new file increases, it is continuously updated. With the accumulation of updates, it will inevitably bring a lot of redundant data to the file information platform, reducing the efficiency of file data retrieval. The existing method for processing redundant data is to detect the number of repetitions. When the number of repetitions exceeds the preset retention threshold, the redundant data is screened out, and there is a possibility that key medical information will be missing, reducing the problem of information retrieval efficiency. Therefore, please refer to Figure 2 , this method also includes the process of removing redundant data from the file information platform, including:
[0022] Prioritize obtaining existing fields. Among them, the existing fields of file information include electronic medical records, imaging materials, and the patient's historical health data, etc. The identification method of redundant data is obtained based on the above-mentioned existing technology and will not be elaborated here;
[0023] S1: Retrieve the redundant data of the existing fields and mark the redundant data. After summarizing the marked data, perform a deduplication operation, detect the marked data, count the proportion of the atomic values of the existing fields, and then perform a joint record on the marked data to obtain the joint call frequency of the marked data;
[0024] Among them, the identification method of redundant data can be based on existing technologies, such as duplicate data detection based on hash value comparison, data similarity analysis, or machine learning models. On the basis of redundant data identification, this invention adopts a multi-level marking mechanism to ensure the integrity and traceability of data. The specific marking method is as follows:
[0025] For each piece of electronic file data (i.e., electronic medical records, imaging materials, and the patient's historical health data), the system generates a unique identifier UID based on the data content + timestamp + resident ID to distinguish different versions of data. Through the UID, it can be quickly determined whether a certain piece of data is redundant data and ensure the uniqueness of the data;
[0026] Optionally, since medical data may be continuously updated due to factors such as patient visits and doctor modifications, the system will set a version number for each piece of data to record changes in different versions. Combining blockchain technology, hash evidence preservation can be performed on important data versions to ensure data traceability and avoid accidental deletion of key information;
[0027] After marking the redundant data, the marked data (i.e., the redundant data is the marked data) is obtained;
[0028] After the redundant data is identified, the system will mark the redundant data. The marked data is summarized according to the classification based on the same UID for subsequent deduplication processing. The specific summarization process is as follows:
[0029] Since each electronic file data will generate a unique identifier based on the data content + timestamp + resident ID, the system will classify the redundant data according to the unique identifier. Data with the same unique identifier is grouped into the same set to form a redundant data set, S = {D 1 , D 2 , …, D n}, where D i represents the i-th version of the redundant data under the same unique identifier, i = 1, 2, …, n, and n represents the total number of versions of the redundant data, and n is a positive integer;
[0030] After the marked data is summarized, the system will perform a deduplication operation to ensure that only necessary data is retained in the database. The specific deduplication steps are as follows:
[0031] In each dataset corresponding to a unique identifier, the latest version is preferentially retained, and the old version data with exactly the same data (i.e., the content and metadata are consistent) is deleted;
[0032] Furthermore, if the old version data has the value of evidence preservation (such as medical history records), it can be stored in the long-term archive database or blockchain without affecting the storage efficiency of the main database;
[0033] Detect the summarized and marked data removal, and count the proportion of existing field atomic values;
[0034] Among them, the atomic value in the proportion of existing field atomic values refers to the smallest data unit that cannot be further divided in the database, such as patient name, date of birth, medical record number, etc. The acquisition logic of the proportion of existing field atomic values is to extract the data of all fields from the marked data, count the number of occurrences of each atomic value in each field, calculate the total number of atomic values of the existing fields, including duplicate values, and then calculate the number of unique atomic values in this field, removing duplicate values, and calculate the ratio of the number of unique atomic values to the total number of atomic values to obtain the proportion of existing field atomic values;
[0035] It should be noted that the corresponding marked data is included in the proportion of atomic values of existing fields. Therefore, the number of parameters of the proportion of atomic values of existing fields obtained is the same as the number of marked data;
[0036] Among them, when the proportion of atomic values of existing fields is higher, the degree of duplication of marked data is lower. A high proportion of atomic values of existing fields means that in a certain field, the proportion of the number of unique atomic values to the total number of atomic values is relatively high, indicating that most records are unique and there is less duplicate data. A low proportion of atomic values of existing fields means that there are more duplicate values in this field, the number of unique atomic values is relatively small, and there is more redundant data;
[0037] Then, the combined records of the marked data are made. Specifically, the combined record refers to recording the number of times of the combined call mechanism for the marked data. The specific combined call mechanism operation is as follows:
[0038] Randomly select a time period for analysis. This time period will be divided into N unit times (for example, seconds, minutes, etc.) for detailed observation. Within each unit time, capture the call status of the marked data and the call status of the remaining data, and judge the combined call status. Specifically: if the marked data and other data are called simultaneously within the current unit time, the current unit time is recorded as 1; if only the marked data is called within the current unit time and other data is not called, the current unit time is recorded as 0. Summarize the status of each unit time to form a call record data set.
[0039] This data set will be used for subsequent analysis;
[0040] It should be noted that there may be multiple schedules within a unit time. However, even if a unit time contains multiple schedules of marked data, if it is not scheduled simultaneously with the remaining data, the current unit time is marked as 0. At the same time, if a unit time contains multiple schedules of marked data and the remaining data, the current unit time is also marked as 1;
[0041] Thus, a set marked with 0 or 1 is obtained, and the total number of marks in the set is the same as the total number of unit times;
[0042] The acquisition logic of the combined call frequency of the marked data is to record through the combined scheduling mechanism within a predetermined time period, generate the status marks of each unit time, summarize the status marks of all unit times to form a mark set, corresponding to each unit time respectively, count the number of values of 1 in the mark set and the number of values of 0 in the mark set, and perform a ratio calculation to obtain the combined call frequency of the marked data;
[0043] It should be noted that a status mark of 1 means that the marked data and other data are called simultaneously within this unit time, and a status mark of 0 means that the marked data is called alone within this unit time;
[0044] Among them, the higher the combined call frequency of the marked data indicates that the marked data is called more times simultaneously with other data. When the marked data is widely used in a specific context, it is less likely to delete the marked data as the screened data;
[0045] S2: Obtain the proportion of atomic values of existing fields and the combined call frequency of marked data, perform standardization processing, and substitute them into the logistic regression model to obtain the data screening coefficient;
[0046] Obtain the proportion of atomic values of existing fields and the combined call frequency of marked data for standardization processing. Specifically, mark the proportion of atomic values of existing fields as Yz, and mark the combined call frequency of marked data as Lh;
[0047] Standardization usually uses Min - Max standardization, and the specific formula is expressed as:
[0048]
[0049] In the formula, Yz is the proportion of atomic values of existing fields, Yz norm is the proportion of atomic values of existing fields after standardization, Yz min is the minimum value of the proportion of atomic values of existing fields, Yz max is the maximum value of the proportion of atomic values of existing fields;
[0050] Among them, the combined call frequency of marked data also uses the above formula for standardization processing, which will not be elaborated here;
[0051] Through standardization processing, it is convenient to compare the proportions of atomic values of different fields and the combined call frequency of marked data, eliminate the influence caused by magnitude differences, and further analyze the standardized results to identify potential correlations and patterns;
[0052] Substitute the proportion of atomic values of existing fields and the combined call frequency of marked data into the logistic regression calculation. The specific formula is expressed as follows:
[0053]
[0054] In the formula, L is the result of logistic regression calculation, that is, the data screening coefficient, e is the natural base, y is the linear combination term of the logistic regression model. Specifically, y can be set as:
[0055] y = β 0 +β 1 ·Yz + β 2 ·Lh;
[0056] In the formula, β 0 is the bias term, β 1 and β2 They are the regression coefficients of the proportion of existing field atomic values and the joint call frequency of labeled data;
[0057] Furthermore, the number of data screening coefficients is consistent with the number of labeled data, which will not be elaborated here;
[0058] It should be noted that the larger the data screening coefficient is, the greater the proportion of existing field atomic values and the frequency of joint calls of labeled data are, which means that the labeled data corresponding to the data screening coefficient does not need to be deleted as redundant data. At the same time, it means that the importance of the corresponding labeled data in the data set is increased, and the important information carried by the data is mined, revealing the internal connection between the data;
[0059] S3: Obtain a data screening coefficient and compare it with a preset screening threshold to obtain standard data and screened data. Count the memory usage of the screened data, obtain the deletion ratio based on the screening mechanism, screen out the screened data corresponding to the deletion ratio, and obtain the remaining screened data.
[0060] The logic for obtaining the screening threshold is to collect the data screening coefficient set corresponding to the historical marked data, divide the data set into a training set and a test set, set the evaluation index and clustering algorithm, and in each round of cross-validation, train the model on the training set and evaluate the model performance on the test set. Then, adjust the screening threshold according to the performance of the validation set. Therefore, the screening threshold is constantly updated.
[0061] In the present invention, clustering algorithm is a kind of unsupervised learning algorithm, which is used to divide the labeled data in the data set into groups or clusters with labels; the common one is K-means clustering, which divides the labeled data in the data set into K clusters, so that the distance between each labeled data and the center point (center of mass) of the cluster to which it belongs is minimized, and finally the possibility of screening out the labeled data is measured by the Euclidean distance, so as to set the screening threshold;
[0062] After obtaining the data screening coefficient, the data screening coefficient is compared and analyzed with the preset iterative screening threshold;
[0063] If the data screening coefficient is greater than or equal to the screening threshold, the current marked data is marked as standard data and an end signal is generated;
[0064] If the data screening coefficient is less than the screening threshold, the current marked data is marked as screened data, and a screening signal is generated;
[0065] For the marked data marked as filtered data, the memory occupancy rate is counted, and the total memory occupancy of the filtered data and the total memory occupancy of the entire set of marked data are calculated, and the memory occupancy rate of the filtered data is obtained by ratio calculation;
[0066] It should be noted that the entire set of marked data contains all the marked data. The occupancy rate in the data memory of the screened data represents the ratio of the occupied screened data. The higher the ratio, the higher the deletion ratio;
[0067] Among them, the screening mechanism refers to performing weighted calculation after normalizing the hardware information of the file information platform and the occupancy rate in the data memory of the screened data;
[0068] The hard disk information of the file information platform includes the total capacity of the hard disk for electronic files;
[0069] Calculate the ratio of the total memory occupancy of the screened data to the total capacity of the hard disk for electronic files to obtain the occupancy ratio of the screened data;
[0070] Perform normalization on the occupancy rate in the data memory of the screened data and the occupancy ratio of the screened data, and substitute it into the weighted calculation to obtain the deletion ratio;
[0071] Specifically, the deletion ratio is kept within the range of 0 - 1;
[0072] According to the deletion ratio, sort the screened data in ascending order of the data screening coefficient values, and screen out the screened data corresponding to the deletion ratio to obtain the remaining screened data;
[0073] For example: when the deletion ratio is 45%, sort the screened data in ascending order of the data screening coefficient values to obtain a set {L 1 , L 2 , …, L z}; where L z is the data screening coefficient corresponding to the z-th screened data, and the numerical value is L 1 < L 2 < … < L z ; then screen out the first 45% of the screened data, and retain the remaining 55% of the screened data as the remaining screened data;
[0074] The present invention marks the redundant data of the existing fields by retrieval, summarizes the marked data and then performs a deduplication operation, detects and jointly records the marked data, obtains the proportion of the atomic values of the existing fields and the joint call frequency of the marked data and performs standardization processing, substitutes it into the logistic regression model, obtains the data screening coefficient, compares it with the screening threshold, obtains the standard data and the screened data, counts the memory occupancy rate of the screened data, obtains the deletion ratio according to the screening mechanism, screens out the screened data corresponding to the deletion ratio, and obtains the remaining screened data, reducing redundant occupancy, ensuring data integrity, avoiding loss of key data, and improving information call efficiency.
[0075] The redundant data removal process of the file information platform provided in this embodiment has the following beneficial technical effects:
[0076] (1) The present invention marks redundant data in existing fields through retrieval, aggregates the marked data and then performs a deduplication operation, detects and jointly records the marked data, obtains the proportion of atomic values of existing fields and the joint call frequency of the marked data and performs standardization processing, substitutes it into a logistic regression model to obtain a data screening coefficient, compares it with a screening threshold to obtain standard data and screened-out data, calculates the memory occupancy rate for the screened-out data, obtains a deletion ratio according to the screening mechanism, screens out the screened-out data corresponding to the deletion ratio to obtain remaining screened-out data, reduces redundant occupancy, ensures data integrity, avoids loss of key data, improves information call efficiency, and thus improves the generation efficiency of the first page of the resident file;
[0077] (2) The present invention obtains and identifies the remaining screened-out data, calls the data packet of the file information platform, obtains the occupancy rate of the remaining screened-out data in the data packet and the change rate of the remaining screened-out data, formulates a set of fuzzy rules for fuzzy reasoning to determine the result of the remaining screened-out data, improves the intelligence level of data screening, re-screens the remaining screened-out data, reduces the storage pressure of the system, and further improves the data call efficiency and the generation efficiency of the first page of the resident file.
[0078] Example 2
[0079] In Example 1 of the present invention, it is mainly illustrated by way of example that redundant data in existing fields are marked through retrieval, the marked data are aggregated and then a deduplication operation is performed, the marked data are detected and jointly recorded, the proportion of atomic values of existing fields and the joint call frequency of the marked data are obtained and standardized processing is performed, substituted into a logistic regression model to obtain a data screening coefficient, compared with a screening threshold to obtain standard data and screened-out data, the memory occupancy rate is calculated for the screened-out data, a deletion ratio is obtained according to the screening mechanism, and the screened-out data corresponding to the deletion ratio are screened out to obtain the remaining screened-out data; however, in Example 1, only one layer of processing is performed on the screened-out data, and whether the remaining screened-out data need to be retained is not mentioned. Obviously, although the information call efficiency can be ensured in a short time under the actual establishment method, in the long-term operation process, untimely deletion of the screened-out data will also lead to a decrease in the information call efficiency and an increase in the memory operation pressure. In view of the above problems, Example 2 of the present invention is further refined;
[0080] S4: Obtain and identify the remaining screened-out data, call the data packet of the file information platform, and obtain the occupancy rate of the remaining screened-out data in the data packet and the change rate of the remaining screened-out data;
[0081] Identify the status of the remaining screened-out data in the data packet of the file information platform, call the data packet of the file information platform, and obtain the occupancy rate of the remaining screened-out data in the data packet and the change rate of the remaining screened-out data;
[0082] The acquisition logic of the remaining screened data in the data packet occupancy rate is to determine the total number of data packets in the archive information platform, analyze the ratio of the remaining screened data in each archive information platform data packet to each archive information platform data packet, and perform cumulative calculation and then ratio calculation with the total number of archive information platform data packets to obtain the occupancy rate of the remaining screened data in the data packet;
[0083] Specifically, the archive information platform data packet refers to the data set stored in the electronic archive information management system, and this data packet usually exists in the form of a folder, a database table or a storage partition. For example, a relational database is used to store archive data;
[0084] The acquisition logic of the change rate of the remaining screened data is to set two identical time periods, analyze the ratio of the number of calls of the remaining screened data in the first time period to the number of calls of the remaining screened data in the second time period, and obtain the change rate of the remaining screened data;
[0085] It should be noted that the set identical time periods may not be adjacent identical time periods, and may be identical time periods with the same duration set in different time dimensions to eliminate the influence of factors such as seasonality and business cycles, making the change rate calculation more accurate;
[0086] S5: Obtain the occupancy rate of the remaining screened data in the data packet and the change rate of the remaining screened data, substitute them into the fuzzy logic for fuzzy inference, and obtain the result of the remaining screened data;
[0087] For example, "High", "Low", "Medium" for the occupancy rate of the remaining screened data in the data packet, and "Fast", "Slow", "Moderate" for the change rate of the remaining screened data;
[0088] Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example:
[0089] Mark the occupancy rate of the remaining screened data in the data packet as X, the change rate of the remaining screened data as U, and the result of the remaining screened data as C_results;
[0090] Then it can be defined as:
[0091] Rule 1: IF (X is Low) AND (U is Slow) THEN (C_results is High)
[0092] Rule 2: IF (X is High) AND (U is Fast) THEN (C_results is Low) ...
[0094] Perform fuzzy inference according to fuzzy rules to determine the results of the remaining data to be screened out;
[0095] It should be noted that the division of fuzzy sets can be adjusted according to the actual situation. For example, although three fuzzy sets are taken as an example in this embodiment, in fact, the remaining data to be screened out can be divided into more than three sets in terms of the packet occupancy rate and the change rate of the remaining data to be screened out, so as to facilitate more accurate adjustment according to different occupancy rates and change rates;
[0096] Furthermore, for the judgment of high, medium, and low of the remaining data to be screened out in terms of the packet occupancy rate and the change rate of the remaining data to be screened out, thresholds can be set for judgment according to the actual situation. For example, when the packet occupancy rate of the remaining data to be screened out exceeds 71%, it is labeled as "High", and when the change rate of the remaining data to be screened out is higher than 67%, it is labeled as "Fast", etc., which will not be elaborated here;
[0097] S6: Archive and store based on the results of the remaining data to be screened out and the standard data, construct the data storage structure of the archive information platform, and provide a data packet download interface;
[0098] Classify the remaining data to be screened out based on the results of the remaining data to be screened out, and classify them into retained data and deleted data. For example: count the results of the remaining data to be screened out labeled as "High" and labeled as "Low", mark the remaining data to be screened out labeled as "High" as retained data for retention, and mark the remaining data to be screened out labeled as "Low" as deleted data for deletion;
[0099] Optionally, partial deletion can also be performed based on the proportion of the number of the remaining data to be screened out labeled as "Low", and the specific number of deletions is not limited, which will not be elaborated here;
[0100] Archive and store the remaining data to be screened out and the standard data. For example:
[0101] Store structured data in a relational database for efficient query, store unstructured data in object storage for convenient access to large files (such as pictures, videos, etc.), encrypt and store sensitive data, and control access permissions;
[0102] It should be noted that the encrypted storage can be determined based on a hash algorithm, or an asymmetric encryption algorithm can be set, and the specific technical operations of encryption are not limited;
[0103] Among them, data attributes refer to the basic characteristics and classification methods of data, and can be based on the data attributes of the classification method, such as retained data and standard data, etc.;
[0104] Furthermore, construct the data storage structure of the archive information platform, including an indexing mechanism, a data access interface, and a storage format. For example, establish a relational database, a NoSQL database, or a distributed storage, etc.;
[0105] Define a data access protocol that allows different systems or users to query, filter, and update the data in the archive information platform. When defining the protocol, access permissions can be set according to the different systems and the permissions of the users. The specific content of the setting is determined by the experimenter based on the actual application data content and will not be elaborated here;
[0106] Write the management rules for standard data and retained data into the metadata management system to ensure that the archive information platform can continuously optimize data storage and management;
[0107] Provide a data packet download interface for external systems or users to access the data in the archive information platform to support business applications;
[0108] The present invention improves the intelligence level of data screening by obtaining and identifying the remaining screened data, calling the data packets of the archive information platform, obtaining the occupancy rate of the remaining screened data in the data packets and the change rate of the remaining screened data, formulating a set of fuzzy rules for fuzzy inference, determining the result of the remaining screened data, re-screening the remaining screened data, reducing the storage pressure of the system, and improving the data call efficiency;
[0109] By using the above technical means, even in the face of a huge amount of historical data to ensure the uniqueness of the data. At the initial stage of establishment, the uniqueness of the data has been verified during the process of marking, summarizing, and de-duplicating the existing fields. At the same time, a distributed storage is established, and distributed identity management can be used to handle the uniqueness of the existing running archives. Finally, the experimenter sets up a mechanism for real-time data verification and change detection and applies it to the specific implementation schemes in the above embodiments to ensure the dynamic update of the archives.
[0110] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for cleaning and assigning historical duplicate files of a large population to one person, one file, and one place, characterized by: include: Based on the resident ID, extract the historical archives matching the resident ID from the archives of all districts and counties in the super large area; Integrate all historical files matching each resident ID, obtain new files corresponding to each resident ID, and store the new files; Assign an administration to each resident's new file in accordance with the attribution administration rules; Establish a management interface for each resident's new profile; The management agency for new archive allocation manages the new archives through the established management interface.
2. The method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location as claimed in claim 1, characterized in that: Historical archives include the front page of historical archives and records of historical events; The integration of all historical files matching each resident ID to obtain a new file corresponding to each resident ID includes: Integrate the historical event records in all historical archives matching each resident ID in chronological order to obtain the complete event record corresponding to each resident ID; Integrate the historical archive homepages of all historical archives matching each resident ID, and construct a unique archive homepage corresponding to each resident ID based on the integration results; The complete event record corresponding to each resident ID and the corresponding unique file homepage constitute a new file corresponding to each resident ID.
3. The method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location as claimed in claim 1, characterized in that: New archives in the archive information platform are stored in more than one data package; The method also includes a redundant data removal process of the archive information platform, including: S1: Retrieve redundant data of existing fields and mark the redundant data, aggregate the marked data and then perform deduplication operations, detect the marked data, count the proportion of atomic values of existing fields, and then jointly record the marked data to obtain the joint call frequency of the marked data; S2: Obtain the proportion of existing field atomic values and the joint call frequency of labeled data and perform normalization, substitute them into the logistic regression model, and obtain the data screening coefficient; S3: Obtain a data screening coefficient and compare it with a preset screening threshold to obtain standard data and screened data. Count the memory usage of the screened data, obtain the deletion ratio based on the screening mechanism, screen out the screened data corresponding to the deletion ratio, and obtain the remaining screened data. S4: Obtain and identify the remaining screened data, call the archive information platform data package, and obtain the remaining screened data in the data package and the remaining screened data change rate; S5: Obtain the occupancy rate of the remaining screened data in the data packet and the change rate of the remaining screened data, substitute them into fuzzy logic for fuzzy reasoning, and obtain the remaining screened data result; S6: Archiving and storing the remaining screened data and standard data, building the data storage structure of the archive information platform, and providing a data package download interface.
4. According to claim 3, a method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location, is characterized by: Existing fields of archival information, including electronic medical records, imaging data, and historical health data of patients, generate unique identifiers based on data content + timestamp + resident ID, and then mark the identified redundant data to obtain marked data, which is then aggregated and deduplicated; Extract data from all fields from the tag data, count the number of occurrences of each atomic value in each field, calculate the total number of atomic values in the existing field, including duplicate values, and then calculate the number of unique atomic values in the field, remove duplicate values, and calculate the ratio of the number of unique atomic values to the total number of atomic values to obtain the ratio of atomic values in the existing field; Then the marked data is jointly recorded. The joint record is to record the number of times the marked data is jointly called. The joint call mechanism randomly selects a time period and divides it into multiple unit times. In each unit time, the call status of the marked data and the call status of the remaining data are captured to judge the joint call status. Recording is performed through the joint scheduling mechanism to generate a status mark for each unit time. The status marks of all unit times are summarized to form a mark set corresponding to each unit time. The number of different mark sets is counted and the ratio is calculated to obtain the joint call frequency of the mark data.
5. According to claim 4, a method for cleaning and assigning historical duplicate files of a large-scale population to one person, one file, and one location, characterized in that: Mark the proportion of existing field atomic values as Yz and the joint call frequency of marked data as Lh for standardization; Substituting the proportion of existing field atomic values and the joint call frequency of labeled data into the logistic regression calculation, the specific formula is expressed as follows: In the formula, L is the result of logistic regression calculation, that is, the data screening coefficient, e is the natural base, and y is the linear combination term of the logistic regression model. Specifically, y can be set as: y=β0+β1·Yz+β2·Lh; Where β0 is the bias term, β1 and β2 are the regression coefficients of the proportion of existing field atomic values and the joint call frequency of labeled data, respectively.
6. The method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location according to claim 5, characterized in that: The number of data screening coefficients is consistent with the number of labeled data; After obtaining the data screening coefficient, the data screening coefficient is compared and analyzed with the preset iterative screening threshold; If the data screening coefficient is greater than or equal to the screening threshold, the current marked data is marked as standard data and an end signal is generated; If the data screening coefficient is less than the screening threshold, the current marked data is marked as screened data and a screening signal is generated.
7. The method for cleaning and assigning historical duplicate files of a large-scale population to one person, one file, and one location according to claim 6, characterized in that: For the marked data marked as filtered data, the memory occupancy rate is counted, and the total memory occupancy of the filtered data and the total memory occupancy of the entire set of marked data are calculated, and the memory occupancy rate of the filtered data is obtained by ratio calculation; The screening mechanism is to calculate the hardware information of the archive information platform and the memory occupancy rate of the screened data and then perform weighted calculation after normalization; The hard disk information of the archive information platform includes the total capacity of the electronic archive hard disk; The total memory usage of the screened data is calculated by ratio to the total capacity of the electronic archive hard disk to obtain the screened data usage ratio; The in-memory occupancy rate of the screened data and the occupancy ratio of the screened data are normalized and substituted into the weighted calculation to obtain the deletion ratio.
8. The method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location according to claim 7, characterized in that: The deletion ratio is maintained within the range of 0-1; according to the deletion ratio, the screened data are sorted in ascending order according to the data screening coefficient value, and the screened data corresponding to the deletion ratio are screened out to obtain the remaining screened data.
9. The method for cleaning and assigning historical duplicate files of a large-area population to one person, one file, and one location according to claim 8, characterized in that: Identify the status of the remaining filtered data in the archive information platform data package, call the archive information platform data package, and obtain the remaining filtered data in the data package occupancy rate and the remaining filtered data change rate; By determining the total number of archive information platform data packets, analyzing the proportion of the remaining filtered data in each archive information platform data packet to each archive information platform data packet, and calculating the ratio with the total number of archive information platform data packets after cumulative calculation, the occupancy rate of the remaining filtered data in the data packet is obtained; By setting two identical time periods, analyzing the ratio of the number of remaining screened data calls in the first period to the number of remaining screened data calls in the second period, the remaining screened data change rate is obtained.
10. The method for cleaning and assigning historical duplicate files of a large-scale population to one person, one file, and one location according to claim 9, characterized in that: The remaining filtered data occupancy rate in the data packet and the remaining filtered data change rate are defined as input variables, and they are divided into different fuzzy sets respectively; The remaining screened data results are defined as output variables and divided into fuzzy sets; Formulate fuzzy rules to describe the influence of the remaining filtered data occupancy rate in the data packet and the remaining filtered data change rate on the remaining filtered data results; Fuzzy reasoning is performed according to fuzzy rules to determine the remaining screened data results.
Citation Information
Patent Citations
Block chain-based resident health record tracing management system
CN112906060A
Cross-regional medical insurance file rapid positioning and sharing system
CN116030926A
Archive filing method and related equipment
CN116069963A
Method and device for realizing electronic archiving of checking business archives
CN116401206A
Automatic filing method and system for personnel archives
CN117827750A