A method for cleaning and attributing one file per person per territory of a historical repeat file of a super large population

By integrating and managing resident files across a vast area, the problems of duplicate filing and chaotic attribution management in traditional file management have been solved, ensuring the uniqueness and integrity of the files and improving the efficiency of information retrieval.

CN120047294BActive Publication Date: 2025-11-07北京啄木鸟云健康科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510410137.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-11-07
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In traditional resident file management, the creation of multiple files for the same person in different districts and counties leads to duplicate filing, fragmented and incomplete service content records, and chaotic management, making it impossible to achieve the uniqueness and integrity of the files.

Method used

By extracting historical archives from archives across a vast region based on resident IDs, integrating and assigning them to a single management agency, the uniqueness and integrity of each resident's archives are ensured. Furthermore, a multi-level tagging mechanism and blockchain technology are employed to ensure the uniqueness and traceability of the data. Combined with redundant data identification and deduplication, redundant data usage is reduced.

Benefits of technology

It ensures the uniqueness and integrity of resident files within a vast area, guarantees the continuity and completeness of the files, reduces redundant data usage, and improves information retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047294B_ABST
    Figure CN120047294B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of archive cleaning and attribution, and discloses a method for cleaning and attributing one-person-one-file-one-territory historical repeated archives of a super-large-area population, which comprises the following steps: extracting historical archives matched with a resident ID from archive databases of all counties in a super-large area based on the resident ID; integrating all historical archives matched with each resident ID to obtain new archives corresponding to each resident ID, and storing the new archives; assigning a management organization to the new archives of each resident according to attribution management rules; establishing a management interface for the new archives of each resident; and managing the new archives through the established management interface by the management organization assigned to the new archives. The application can clean historical repeated archives in a super-large area, obtain complete and unique new archives of each resident by integrating the historical archives of each resident, and set unique attribution management organizations for each new archive, so that one-person-one-file-one-territory management of resident archives is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of archive cleaning and attribution, and more particularly to a method for cleaning and attributing one person one archive one attribution area for historical repeated archives of a super large regional population. BACKGROUND

[0002] In the traditional resident archive management method, each district and county has independently established resident archives for a long time, and each district and county has an independent resident archive database. With the population flow, the same person goes to different districts and counties for activities such as studying, living, working and traveling, which leads to the same person receiving services in different districts and counties and establishing an archive respectively. Thus, from the municipal level, the provincial level or even the national level, there are problems of historical archive repetition of a super large regional population and multiple archives of a single resident, which leads to archive repetition, service content record fragmentation, incompleteness and discontinuity, and thus loses the value of the archives. In addition, since the same person has established archives in different districts and counties, there is also a problem of chaotic archive attribution management.

[0003] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, and thus it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a method for cleaning and attributing one person one archive one attribution area for historical repeated archives of a super large regional population, which solves the problems raised in the above background technology by using different product inspection methods.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical scheme, a method for cleaning and attributing one person one archive one attribution area for historical repeated archives of a super large regional population, comprising: extracting historical archives matched with a resident ID from archive databases of all districts and counties in a super large region based on the resident ID; integrating all historical archives matched with each resident ID to obtain a new archive corresponding to each resident ID, and storing the new archive; assigning a management agency for each resident's new archive according to attribution management rules; establishing a management interface for each resident's new archive; and managing the new archive by the management agency assigned to the new archive through the established management interface.

[0006] Preferably, the historical archives include a historical archive home page and historical event records; the integrating all historical archives matched with each resident ID to obtain a new archive corresponding to each resident ID comprises: integrating the historical event records in all historical archives matched with each resident ID in chronological order to obtain complete event records corresponding to each resident ID; integrating the historical archive home pages of all historical archives matched with each resident ID, and constructing a unique archive home page corresponding to each resident ID according to the integration result; and the complete event records corresponding to each resident ID and the unique archive home page corresponding to each resident ID constitute the new archive corresponding to each resident ID.

[0007] The beneficial technical effects of the present application: the present application solves the problem of uniqueness and integrity of historical archives, in order to ensure the continuity and integrity of the archives of residents in a super large area, and realize the record of the whole life cycle, the first step is to extract the historical archives matched with the resident ID from the archives of all counties in the super large area based on the resident ID, to ensure the uniqueness and integrity of the historical archives (if the archives established by huge manpower, material and financial resources over the years are not cleaned up and utilized, it will cause great waste, and also lead to the loss of historical information of residents); the second step is to integrate all historical archives matched with each resident ID to obtain a new archive corresponding to each resident ID, and assign a management agency for the new archive of each resident according to the attribution management rule, to ensure the uniqueness and integrity of the archives in the new service process; the present application first cleans up and removes the historical duplicate archives of the super large area population, and retains one archive; secondly, the historical service records of the residents in multiple places are integrated and collected into one archive; thirdly, according to the attribution management rule, the unique management agency of the archives is clear. The present application can clean up the historical duplicate archives in a super large area such as a city, a province or a country, obtain a complete and unique new archive for each resident by integrating the historical archives of each resident, and set a unique attribution management agency for each new archive, thereby realizing the one-person-one-file-one-territory management of the archives of residents. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 The flowchart of the method for cleaning up and attributing the historical duplicate archives of the super large area population according to the present application.

[0009] Figure 2 The flowchart of the process of removing redundant data of the archive information platform according to the present application. DETAILED DESCRIPTION

[0010] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work belong to the scope of protection of the present application.

[0011] Embodiment 1

[0012] Please refer to Figure 1 A method for cleaning and attributing a person's historical archives in a super large area population, the specific operation process is as follows:

[0013] Step A1, based on the resident ID, extract the historical archives matched with the resident ID from the archives library of all counties in the super large area. The resident ID is preferably but not limited to the ID card number or the birth medical certificate number (newborn). The super large area represents one or more city-level areas or one or more provincial administrative areas or national administrative areas.

[0014] Step A2, integrate all historical archives matched with each resident ID, obtain a new archive corresponding to each resident ID, and store the new archive. It is not limited to quickly completing the integration process by the following way to obtain the integrity and uniqueness of the new archive corresponding to each resident ID. The historical archives include the historical archive homepage and the historical event record, such as electronic health archives of the resident archives, historical events including physical examination events, treatment events, examination events, vaccination events and diagnosis events, etc., each event has an event occurrence time. Step A2 includes:

[0015] Step A21, integrate the historical event records in all historical archives matched with each resident ID according to the chronological order of the event occurrence time, and obtain the complete event record corresponding to each resident ID.

[0016] Step A22, integrate the historical archive homepage of all historical archives matched with each resident ID, and construct the unique archive homepage corresponding to each resident ID according to the integration result. The integration of the historical archive homepage of all historical archives matched with each resident ID mainly includes: integrating all different field names in all historical archive homepages, updating the content of each field to the latest record data of the field, such as setting the residence address field to the residence address in the historical archive with the latest time. Write the content of all different field names of each resident ID after updating to the preset archive homepage template to obtain the unique archive homepage corresponding to each resident ID.

[0017] Step A23, the complete event record corresponding to each resident ID and the unique archive homepage form a new archive corresponding to each resident ID.

[0018] Step A3, assign a management agency to each resident's new file according to the rules of home management. The rules of home management are to take the file management agency of the resident's permanent residence or the district where the resident's household is located as the management agency of each resident's new file.

[0019] Step A4, establish a management interface for each resident's new file. The management interface is preferably but not limited to an application interface or an Internet link.

[0020] Step A5, the management agency assigned to the new file manages the new file through the established management interface, such as updating, modifying, etc.

[0021] In a preferred embodiment of the present example, the new file in the file information platform is stored in the form of more than one data package. As the running time of the new file increases, it is constantly updated, and as the accumulation of updates increases, it will inevitably bring a lot of redundant data to the file information platform, reducing the efficiency of calling file data. The existing redundant data processing method is to detect the number of repetitions, and when the number of repetitions exceeds the preset retention threshold, the redundant data is screened out, which may cause the loss of key medical information, reducing the efficiency of information calling. Therefore, please refer to Figure 2 The method also includes a redundant data removal process of the file information platform, comprising:

[0022] Priority is given to existing fields, which include electronic medical records, image data, and patient health history data, etc. The identification method of redundant data is based on the existing technology, which is not described here;

[0023] S1: search for redundant data in existing fields and mark the redundant data. After the marked data is summarized and the de-duplication operation is performed, the marked data is detected, the proportion of existing field atomic values is counted, and the marked data is recorded jointly to obtain the joint calling frequency of the marked data;

[0024] Among them, the identification method of redundant data can be based on existing technologies, such as hash value comparison, data similarity analysis or machine learning model for duplicate data detection. Based on the identification of redundant data, the present application adopts a multi-level marking mechanism to ensure the integrity and traceability of the data. The specific marking method is as follows:

[0025] For each electronic file data (i.e. electronic medical records, image data and patient health history data), the system generates a unique identifier UID based on data content + timestamp + resident ID to distinguish different versions of data. Through UID, it can quickly judge whether a certain data is redundant data, and ensure the uniqueness of the data;

[0026] Optionally, since the medical data may be updated due to patient's medical condition, doctor's modification and other factors, the system sets a version number for each piece of data to record the changes of different versions. Combined with the blockchain technology, the important data versions can be stored by hash, ensuring the traceability of data and avoiding the deletion of key information.

[0027] After marking the redundant data, the marked data (i.e. redundant data) is obtained.

[0028] After the identification of redundant data is completed, the system will mark the redundant data. The marked data is summarized according to the classification based on the same UID, so as to facilitate subsequent deduplication processing. The specific summarization process is as follows:

[0029] Since each electronic archive data will generate a unique identifier based on data content + timestamp + resident ID, the system will classify the redundant data according to the unique identifier. The data with the same unique identifier is classified into the same group to form a redundant data set, S = {D1, D2, …, Di}, where Di represents the i-th version of redundant data under the same unique identifier, i = 1, 2, …, n, n represents the total number of versions of redundant data, and n is a positive integer. n i

[0030] After the marked data is summarized, the system will perform a deduplication operation to ensure that only necessary data is retained in the database. The specific deduplication steps are as follows:

[0031] In each data set corresponding to a unique identifier, the latest version is preferentially retained, and the old version data that is completely identical (i.e. consistent in content and metadata) is deleted.

[0032] Further, for the old version data that has archival value (such as medical history records), it can be stored in a long-term archival database or a blockchain, without affecting the storage efficiency of the main database.

[0033] The marked data that has been summarized and removed is detected, and the proportion of existing field atomic values is calculated.

[0034] Among them, the atomic value in the proportion of existing field atomic values refers to the smallest unit of data that cannot be further divided in the database, such as patient name, date of birth, medical record number, etc. The logic of obtaining the proportion of existing field atomic values is to extract all the data of the fields from the marked data, count the number of occurrences of each atomic value in each field, calculate the total number of atomic values of the field, including repeated values, calculate the number of unique atomic values in the field, remove the repeated values, and calculate the ratio of the number of unique atomic values to the total number of atomic values to obtain the proportion of existing field atomic values.

[0035] ​​It should be noted that the existing field atomic value proportion has an existing field corresponding to the marked data, therefore, the number of obtained existing field atomic value proportion parameters is consistent with the number of marked data;

[0036] Wherein, the higher the atomic value proportion of the existing field is, the lower the degree of repetition of the marked data is, the high atomic value proportion of the existing field represents that the number of unique atomic values accounts for a high proportion of the total atomic value number in a certain field, indicating that most records are unique and there is less repeated data, the low atomic value proportion of the existing field indicates that there are more repeated values in the field and the number of unique atomic values is relatively small, and there is more redundant data;

[0037] The marked data is further combined and recorded, and the specific combined recording refers to recording the number of times of combined calling mechanism of the marked data, and the specific combined calling mechanism operation is as follows:

[0038] A time period is randomly selected for analysis, and the time period will be divided into N unit times (for example, seconds, minutes, etc.) for detailed observation, in each unit time, the calling state of the marked data and the calling state of the remaining data are captured, and the combined calling state is judged, specifically: if the marked data and other data are called at the same time in the current unit time, the current unit time is recorded as 1, if only the marked data is called without calling other data in the current unit time, the current unit time is recorded as 0, and the state of each unit time is summarized to form a calling record data set.

[0039] The data set will be used for subsequent analysis;

[0040] It should be noted that there may be multiple dispatches in a unit time, however, even if the unit time contains multiple dispatches of the marked data, if it is not dispatched with the remaining data, it is marked as 0 for the current unit time, at the same time, the unit time contains multiple dispatches of the marked data and the remaining data, which are also marked as 1 for the current unit time;

[0041] Therefore, a set of 0 or 1 is obtained, and the total number of marks in the set is consistent with the total number of unit times;

[0042] The logic of obtaining the combined calling frequency of the marked data is to record the state mark of each unit time by the combined scheduling mechanism in a predetermined time period, summarize the state marks of all unit times to form a mark set, respectively corresponding to each unit time, count the number of values 1 in the mark set and the number of values 0 in the mark set, and perform ratio calculation to obtain the combined calling frequency of the marked data;

[0043] It should be noted that the state mark of 1 indicates that the marked data and other data are called at the same time in the unit time, and the state mark of 0 indicates that the marked data is called alone in the unit time.

[0044] Wherein, the higher the joint call frequency of the marked data, the more times the marked data is called simultaneously with other data, and the more the marked data is widely used in a specific context, the less likely it is to delete the marked data as screening data;

[0045] S2: Obtain the atomic value proportion of the existing field and the joint call frequency of the marked data and perform standardization processing, and substitute into the logistic regression model to obtain the data screening coefficient;

[0046] Obtain the atomic value proportion of the existing field and the joint call frequency of the marked data and perform standardization processing, specifically, mark the atomic value proportion of the existing field as Yz, and mark the joint call frequency of the marked data as Lh;

[0047] Standardization usually uses Min-Max standardization, and the specific formula is expressed as:

[0048]

[0049] In the formula, Yz is the atomic value proportion of the existing field, Yz norm is the standardized atomic value proportion of the existing field, Yz min is the minimum value of the atomic value proportion of the existing field, Yz max is the maximum value of the atomic value proportion of the existing field;

[0050] Wherein, the joint call frequency of the marked data is also standardized by the above formula, which will not be repeated here;

[0051] Through standardization processing, the atomic value proportion of different fields and the joint call frequency of the marked data can be easily compared, the influence of the order of magnitude difference is eliminated, and the standardized results are further analyzed to identify potential relevance and patterns;

[0052] Substitute the atomic value proportion of the existing field and the joint call frequency of the marked data into the logistic regression calculation, and the specific formula is expressed as follows:

[0053]

[0054] In the formula, L is the logistic regression calculation result, that is, the data screening coefficient, e is the natural base, y is the linear combination term of the logistic regression model, and specifically y can be set as:

[0055] y = β0 + β1·Yz + β2·Lh;

[0056] In the formula, β0 is the bias term, and β1 and β2 are the regression coefficients of the atomic value proportion of the existing field and the joint call frequency of the marked data, respectively;

[0057] Further, the number of data screening coefficients is consistent with the number of labeled data, which is not described here;

[0058] It should be noted that the larger the data screening coefficient is, the larger the joint call frequency of the existing field atomic value proportion and the labeled data is, the more the labeled data corresponding to the data screening coefficient is unnecessary to be deleted as redundant data, and at the same time, the importance of the corresponding labeled data in the data set is improved, important information carried by the data is mined, and the internal relationship between the data is revealed;

[0059] S3: Obtain the data screening coefficient, and compare it with the preset screening threshold to obtain standard data and screened data, count the memory occupancy rate of the screened data, obtain the deletion rate according to the screening mechanism, screen the screened data corresponding to the deletion rate, and obtain the remaining screened data;

[0060] The acquisition logic of the screening threshold is to collect a set of data screening coefficients corresponding to the historical labeled data, then divide the data set into a training set and a test set, set an evaluation index and a clustering algorithm, train the model on the training set in each iteration of cross-validation, and evaluate the model performance on the test set, and then adjust the screening threshold according to the performance of the validation set, therefore, the screening threshold is constantly updated;

[0061] In the present application, the clustering algorithm is a kind of unsupervised learning algorithm, which is used to divide the labeled data in the data set into groups or clusters with labels; common K-means clustering divides the labeled data in the data set into K clusters, so that the distance between each labeled data and the center point (centroid) of its belonging cluster is minimized, and finally the Euclidean distance is used to measure the screening possibility of the labeled data, so as to set the screening threshold;

[0062] After obtaining the data screening coefficient, the data screening coefficient is compared and analyzed with the preset constantly updated screening threshold;

[0063] If the data screening coefficient is greater than or equal to the screening threshold, the current labeled data is marked as standard data, and an end signal is generated;

[0064] If the data screening coefficient is less than the screening threshold, the current labeled data is marked as screened data, and a screening signal is generated;

[0065] For the labeled data marked as screened data, the occupancy rate in the memory is counted, the total memory occupancy of the screened data and the total memory occupancy of the entire labeled data set are calculated, the ratio is calculated, and the memory occupancy rate of the screened data is obtained;

[0066] It should be noted that the entire set of marked data contains all the marked data, and the memory occupation rate of the screening data is used to express the ratio of the screening data, the higher the ratio, the higher the deletion rate;

[0067] The screening mechanism refers to the weighted calculation after normalization processing of the file information platform hardware information and the memory occupation rate of the screening data;

[0068] The file information platform hard disk information includes the total capacity of the electronic file hard disk;

[0069] The total memory occupation of the screening data is calculated by the total capacity of the electronic file hard disk, and the screening data occupation ratio is obtained;

[0070] The screening data memory occupation rate and the screening data occupation ratio are normalized and substituted into the weighted calculation to obtain the deletion rate;

[0071] Specifically, the deletion rate is kept in the range of 0-1;

[0072] According to the deletion rate, the screening data is sorted in order from small to large according to the data screening coefficient value, and the screening data corresponding to the deletion rate is screened out to obtain the remaining screening data;

[0073] For example, when the deletion rate is 45%, the screening data is sorted in order from small to large according to the data screening coefficient value to obtain a set {L1, L2, …, L z}; Wherein, L z is the data screening coefficient corresponding to the zth screening data, and the value size is L1<L2<…<L z ; The first 45% of the screening data is screened out, and the remaining 55% of the screening data is reserved as the remaining screening data;

[0074] The present application marks the redundant data by searching the existing fields, and then performs the deduplication operation after the marked data is collected. The marked data is detected and recorded to obtain the atomic value proportion of the existing field and the joint call frequency of the marked data, and then standardized processing is performed. The data screening coefficient is obtained by substituting the logistic regression model, and compared with the screening threshold to obtain the standard data and the screening data. The memory occupation rate of the screening data is calculated, and the deletion rate is obtained according to the screening mechanism. The remaining screening data is obtained by screening the screening data corresponding to the deletion rate, reducing the redundant occupation, ensuring the data integrity, avoiding the loss of key data, and improving the information calling efficiency.

[0075] The redundant data removal processing of the file information platform provided by the embodiment has the following beneficial technical effects:

[0076] (1) The present application marks the redundant data of the existing field by searching, performs the deduplication operation after the marked data is summarized, detects and records the marked data, obtains the atomic value proportion of the existing field and the joint call frequency of the marked data, and performs standardization processing, substitutes into the logistic regression model, obtains the data screening coefficient, and compares with the screening threshold, obtains the standard data and the screened data, and for the screened data, the memory occupation rate is counted, the deletion ratio is obtained according to the screening mechanism, the screened data corresponding to the deletion ratio is screened out, the remaining screened data is obtained, the redundant occupation is reduced, the data integrity is guaranteed, the key data loss is avoided, the information call efficiency is improved, and the resident file first page generation efficiency is improved;

[0077] (2) The present application obtains and identifies the remaining screened data, calls the archive information platform data package, obtains the data package occupation rate of the remaining screened data and the change rate of the remaining screened data, formulates a set of fuzzy rules for fuzzy reasoning, determines the remaining screened data result, improves the intelligent degree of data screening, re-screens the remaining screened data, reduces the system storage pressure, and further improves the data call efficiency and the resident file first page generation efficiency.

[0078] Embodiment 2

[0079] In the embodiment 1 of the present application, the redundant data of the existing field is marked by searching, the marked data is summarized, and then the deduplication operation is performed, the marked data is detected and recorded, the atomic value proportion of the existing field and the joint call frequency of the marked data are obtained, and standardization processing is performed, which is substituted into the logistic regression model, the data screening coefficient is obtained, and compared with the screening threshold, the standard data and the screened data are obtained, and for the screened data, the memory occupation rate is counted, the deletion ratio is obtained according to the screening mechanism, the screened data corresponding to the deletion ratio is screened out, and the remaining screened data is obtained; but in the embodiment 1, only one layer of processing is performed on the screened data, and whether the remaining screened data needs to be retained is not mentioned, obviously, although the information call efficiency can be guaranteed in a short time under the actual establishment method, the information call efficiency is also reduced if the screened data is not deleted in time in the long-term running process, the memory running pressure is increased, and in view of the above problems, the embodiment 2 of the present application is further refined;

[0080] S4: obtaining and identifying the remaining screened data, calling the archive information platform data package, obtaining the data package occupation rate of the remaining screened data and the change rate of the remaining screened data;

[0081] The state of the remaining screened data in the archive information platform data package is identified, the archive information platform data package is called, the data package occupation rate of the remaining screened data and the change rate of the remaining screened data are obtained;

[0082] The acquisition logic of the residual screening data in the data packet occupancy rate is to determine the total number of data packets of the archive information platform, analyze the proportion of the residual screening data in each data packet of the archive information platform and each data packet of the archive information platform, and then perform accumulation calculation and ratio calculation with the total number of data packets of the archive information platform to obtain the residual screening data in the data packet occupancy rate.

[0083] Specifically, the archive information platform data packet refers to a data collection stored in the electronic archive information management system. The data packet usually exists in the form of a folder, a database table or a storage partition, such as a relational database or a storage archive data.

[0084] The acquisition logic of the residual screening data change rate is to set two same time periods, analyze the ratio of the residual screening data call times in the first time period and the residual screening data call times in the second time period, and obtain the residual screening data change rate.

[0085] It should be noted that the set same time period can not be adjacent to the same time period, and can be the same time period set in different time dimensions to eliminate the influence of seasonal factors, business cycles and other factors, so that the change rate calculation is more accurate.

[0086] S5: Obtain the residual screening data in the data packet occupancy rate and the residual screening data change rate, and input the fuzzy logic for fuzzy reasoning to obtain the residual screening data result.

[0087] For example, "High", "Low", "Medium" for the residual screening data in the data packet occupancy rate, "Fast", "Slow", "Moderate" for the residual screening data change rate.

[0088] A set of fuzzy rules are developed to describe the influence of different input variables on the output variable. The definition of the rule can be based on professional knowledge, or obtained through data analysis and experiments. For example:

[0089] The residual screening data in the data packet occupancy rate is marked as X, the residual screening data change rate is marked as U, and the residual screening data result is marked as C_results.

[0090] Then the following can be defined:

[0091] Rule 1: IF (X is Low) AND (U is Slow) THEN (C_results is High)

[0092] Rule 2: IF (X is High) AND (U is Fast) THEN (C_results is Low) ...

[0094] According to the fuzzy rule, the fuzzy inference is carried out to determine the remaining screening data result;

[0095] It should be noted that the division of the fuzzy set can be adjusted according to the actual situation, for example, although three fuzzy sets are taken as examples in the embodiment, the remaining screening data can be divided into more than three sets according to the data packet occupancy rate and the remaining screening data change rate, so as to facilitate better and more accurate adjustment according to different occupancy rates and change rates;

[0096] Further, for the judgment of the high, medium and low of the remaining screening data in the data packet occupancy rate and the remaining screening data change rate, the threshold can be set to judge according to the actual situation, for example, when the remaining screening data in the data packet occupancy rate is more than 71%, it is marked as "High", and when the remaining screening data change rate is higher than 67%, it is marked as "Fast", and the like, which will not be repeated here;

[0097] S6: According to the remaining screening data result and the standard data, the data storage structure of the archive information platform is constructed, and a data packet download interface is provided;

[0098] Based on the remaining screening data result, the remaining screening data is classified into reserved data and deleted data, for example: the remaining screening data results marked as "High" and "Low" are counted, the remaining screening data marked as "High" is marked as reserved data for reservation, and the remaining screening data marked as "Low" is marked as deleted data for deletion;

[0099] Optionally, the number of the remaining screening data marked as "Low" can also be partially deleted based on the number of the remaining screening data marked as "Low", and the specific number of deletion is not limited, which will not be repeated here;

[0100] The remaining screening data and the standard data are archived and stored, for example:

[0101] The structured data is stored in a relational database for efficient query, the unstructured data is stored in an object storage for convenient access to large files (such as pictures, videos, etc.), sensitive data is stored by encryption, and access permission is controlled;

[0102] It should be noted that the encryption storage can be determined based on a hash algorithm, or an asymmetric encryption algorithm can be set, and the specific encryption technical operation is not limited;

[0103] Among them, the data attribute refers to the basic characteristics and classification method of the data, which can be based on the data attribute of the classification method, such as reserved data and standard data, etc.

[0104] Further, the data storage structure of the archive information platform is constructed, including an index mechanism, a data access interface and a storage format, for example, a relational database, a NoSQL database or a distributed storage, etc.

[0105] A data access protocol is defined to allow different systems or users to query, filter and update the data in the archive information platform, and the access permission can be set according to the permissions of different systems and users when the protocol is defined, and the specific setting content is determined by the experimenters based on the actual application data content, which is not described here.

[0106] The management rules of the standard data and the reserved data are written into the metadata management system to ensure that the archive information platform can continuously optimize data storage and management.

[0107] A data package download interface is provided to allow external systems or users to access the data in the archive information platform to support business applications.

[0108] The application improves the intelligent degree of data screening by obtaining and identifying the remaining screening data, calling the archive information platform data package, obtaining the data package occupancy rate of the remaining screening data and the change rate of the remaining screening data, formulating a set of fuzzy rules for fuzzy reasoning, determining the remaining screening data result, re-screening the remaining screening data, reducing the system storage pressure and improving the data calling efficiency.

[0109] By using the above technical means, even if faced with huge historical data to ensure the uniqueness of the data, the uniqueness of the data has been verified in the process of marking and summarizing the existing fields in the initial stage, and a distributed storage is established, and the distributed identity management can cope with the uniqueness of the existing running archives, and finally the experimenters set the real-time data verification and change detection mechanism applied to the specific implementation scheme in the above embodiment to ensure the dynamic update of the archives.

[0110] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for clearing and attributing duplicate historical records of a large population area, characterized by: Comprising: extracting a historical profile matching the resident ID from the profile database of all districts and counties in the super-large region based on the resident ID; integrating all historical profiles matching each resident ID to obtain a new profile corresponding to each resident ID, and storing the new profile; assigning a management agency to the new profile of each resident according to the attribution management rule; establishing a management interface for the new profile of each resident; The management mechanism of the new file allocation manages the new file through the established management interface, wherein the new file in the file information platform is stored in a manner of being divided into more than one data package; the method further includes a redundant data removal processing procedure of the file information platform, including: S1: searching for redundant data of an existing field and marking the redundant data, performing a de-duplication operation after the marked data is summarized, detecting the marked data, counting the proportion of atomic values of the existing field, and jointly recording the marked data to obtain the joint call frequency of the marked data, wherein the existing field of the file information includes an electronic medical record, image data, and patient historical health data, a unique identifier is generated based on data content+timestamp+resident ID, and after the identified redundant data is marked, the marked data is obtained, the marked data is summarized and de-duplicated; the data of all fields in the marked data is extracted, the occurrence times of each atomic value in each field are counted, the total atomic value quantity of the existing field is calculated, including repeated values, the unique atomic value quantity of the field is calculated, the repeated values are removed, the ratio of the unique atomic value quantity to the total atomic value quantity is calculated, and the atomic value proportion of the existing field is obtained; the marked data is jointly recorded, the joint recording is recording the number of times of joint call mechanism of the marked data, the joint call mechanism is randomly selecting a time period to be divided into multiple unit times, capturing the call state of the marked data and the call state of the remaining data in each unit time, and judging the joint call state; the state mark of each unit time is generated by recording through the joint scheduling mechanism, all the state marks of the unit times are summarized to form a mark set corresponding to each unit time, the number of different mark sets is counted and the ratio is calculated to obtain the joint call frequency of the marked data; S2: obtaining the atomic value proportion of the existing field and the joint call frequency of the marked data and performing standardization processing, substituting into a logistic regression model to obtain a data screening coefficient; S3: obtaining the data screening coefficient and comparing it with a preset screening threshold to obtain standard data and screened data, counting the memory occupancy rate of the screened data, obtaining a deletion ratio according to a screening mechanism, screening the screened data corresponding to the deletion ratio to obtain remaining screened data; S4: obtaining and identifying the remaining screened data, calling the file information platform data package, obtaining the data package occupancy rate of the remaining screened data and the change rate of the remaining screened data; S5: obtaining the data package occupancy rate of the remaining screened data and the change rate of the remaining screened data, substituting into fuzzy logic for fuzzy reasoning to obtain the remaining screened data result; S6: archiving and storing the standard data and the remaining screened data result to construct a data storage structure of the file information platform and provide a data package download interface.

2. The method for cleaning and attributing a person to a file to a territory of a historical repetition file of a super large area of people according to claim 1, characterized in that: The historical file includes a historical file home page and historical event records; The method for integrating all the historical files matched with each resident ID to obtain new files corresponding to each resident ID includes: Integrate the historical event records in all historical archives matched with each resident ID according to the chronological order to obtain complete event records corresponding to each resident ID; Integrate the first pages of all historical archives matched with each resident ID, and construct a unique archive first page corresponding to each resident ID according to the integration result; The complete event records corresponding to each resident ID and the unique archive first page corresponding to each resident ID constitute a new archive corresponding to each resident ID.

3. The method of claim 1, wherein the method is characterized by: Mark the proportion of the atomic value of the existing field as Yz, mark the joint calling frequency of the marked data as Lh, and perform standardization processing; The specific formula expression of the existing field atomic value proportion and the joint calling frequency of the marked data in the logistic regression is as follows: In the formula, L is the logistic regression calculation result, that is, the data screening coefficient, e is the natural base, y is the linear combination term of the logistic regression model, and y can be set as: y = β0 + β1·Yz + β2·Lh; In the formula, β0 is a bias term, β1 and β2 are regression coefficients of the existing field atomic value proportion and the joint calling frequency of the marked data, respectively.

4. The method for cleaning and attributing a person's file to a territory according to claim 3, characterized in that: The number of data screening coefficients is consistent with the number of marked data; After obtaining the data screening coefficient, the data screening coefficient is compared and analyzed with the preset continuously iterated screening threshold value; If the data screening coefficient is greater than or equal to the screening threshold value, the current marked data is marked as standard data, and an end signal is generated; If the data screening coefficient is less than the screening threshold value, the current marked data is marked as excluded data, and an exclusion signal is generated.

5. A method for cleaning and attributing a person's file to a territory according to claim 4, characterized in that: For the marked data marked as excluded data, the memory occupancy rate is calculated, the total memory occupancy of the excluded data and the total memory occupancy of the entire marked data set are calculated, and the ratio is calculated to obtain the memory occupancy rate of the excluded data; The exclusion mechanism is to normalize the hardware information of the archive information platform and the memory occupancy rate of the excluded data, and then perform weighted calculation; The hard disk information of the archive information platform includes the total capacity of the electronic archive hard disk; The total memory occupancy of the excluded data and the total capacity of the electronic archive hard disk are calculated to obtain the excluded data occupancy ratio; The memory occupancy rate of the excluded data and the excluded data occupancy ratio are normalized and substituted into the weighted calculation to obtain the deletion ratio.

6. A method for cleaning and attributing a person's file to a territory according to claim 5, characterized in that: The deletion ratio is kept in the range of 0-1; according to the deletion ratio, the excluded data is sorted in order from small to large according to the data screening coefficient value, the excluded data corresponding to the deletion ratio is excluded, and the remaining excluded data is obtained.

7. A method for cleaning and attributing a person's file to a territory according to claim 6, characterized in that: Identify the state of the remaining excluded data in the archive information platform data packet, call the archive information platform data packet, and obtain the data packet occupancy rate of the remaining excluded data and the change rate of the remaining excluded data; By determining the total number of archive information platform data packets, analyzing the proportion of the remaining excluded data in each archive information platform data packet, and performing cumulative calculation and ratio calculation with the total number of archive information platform data packets, the data packet occupancy rate of the remaining excluded data is obtained; By setting two same time periods, the ratio of the number of times of calling the remaining excluded data in the first time period to the number of times of calling the remaining excluded data in the second time period is analyzed to obtain the change rate of the remaining excluded data.

8. The method for cleaning and attributing a person's file to a territory according to claim 7, characterized in that: The remaining screening data is defined as input variable in terms of data packet occupancy and remaining screening data change rate, and is divided into different fuzzy sets respectively; The remaining screening data result is defined as output variable, and is divided into fuzzy sets; Fuzzy rules are made to describe the influence of data packet occupancy and remaining screening data change rate on the remaining screening data result; Fuzzy reasoning is made according to the fuzzy rules to determine the remaining screening data result.

Citation Information

Patent Citations

  • Automatic filing method and system for personnel archives

    CN117827750A

  • Electric power engineering archive knowledge service method and system based on AI

    CN119166747A