A data processing method and apparatus

By storing and naming data to be processed in key-value pairs according to preset rules in the data processing system, and responding to data acquisition requests to determine enhanced data and perform de-identification processing, the problem of data fusion management across multiple business platforms is solved, and secure, accurate and unified data management is achieved.

CN114741421BActive Publication Date: 2026-04-10KANGFUZI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

How to achieve integrated management of data from multiple business platforms while ensuring data accuracy, especially given the unique nature of medical data and the diverse rules for integrating data from multiple business platforms, presents significant challenges in the implementation of existing technologies.

Method used

By employing data processing methods and devices, the data to be processed is stored in the form of key-value pairs, named in accordance with preset rules, and responds to the data acquisition requests of the target object. After determining the enhanced data, it is then de-identified to achieve data security management.

Benefits of technology

It achieves accurate data fusion across multiple data platforms while ensuring security, enabling secure and accurate unified management and control of data from multiple business platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741421B_ABST
    Figure CN114741421B_ABST
Patent Text Reader

Abstract

The application provides a data processing method and device, after obtaining a plurality of to-be-processed data from a plurality of data platforms, the to-be-processed data is stored in the form of key-value pairs, the naming of the keys in the key-value pairs conforms to a preset rule, so that unified storage of the data of the plurality of data platforms can be realized, the naming of the unified keys makes the stored to-be-processed data easy to process, the processing accuracy of the to-be-processed data is improved, in response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group is determined according to the stored plurality of to-be-processed data, so that fusion management of the data of the plurality of data platforms can be realized, after desensitization processing of the enhanced data obtains desensitized data, the desensitized data is provided to the target object, so that security management of the data is realized, therefore, the embodiments of the application can realize accurate fusion of the data of the plurality of data platforms, and security is taken into account, so that unified management and control of the data of the plurality of service platforms is realized in a safe and accurate manner.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data processing, in particular to a data processing method and device. BACKGROUND

[0002] With the advancement of data informatization construction, there is a need to manage the data provided by multiple business platforms. For example, a hospital or other enterprises with medical qualifications as multiple business platforms have created numerous patient-based data systems, each of which has accumulated a large amount of data, and the importance and necessity of data synchronization between data systems have become increasingly prominent.

[0003] How to realize the fusion management of multiple business platforms under the premise of ensuring data accuracy is an important research at present. SUMMARY

[0004] Therefore, the purpose of the present application is to provide a data processing method and device to realize accurate unified management and control of multiple business platform data.

[0005] To achieve the above purpose, the present application has the following technical solutions:

[0006] The present application provides a data processing method, characterized in that it comprises:

[0007] After obtaining multiple to-be-processed data from multiple data platforms, the to-be-processed data is stored in the form of key-value pairs; the naming of the key in the key-value pair conforms to a preset rule;

[0008] In response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group is determined according to the stored multiple to-be-processed data;

[0009] After desensitizing the enhanced data to obtain desensitized data, the desensitized data is provided to the target object.

[0010] Optionally, the to-be-processed data has a data weight corresponding to the data platform to which it belongs, and the determination of the enhanced data belonging to the preset information group according to the stored multiple to-be-processed data comprises:

[0011] The original data whose key matches the preset information group is determined from the to-be-processed data;

[0012] According to the data weight of the original data, multiple to-be-merged data corresponding to the first key in the original data are merged to obtain merged data corresponding to the first key;

[0013] The multiple to-be-merged data in the original data are replaced by the merged data to obtain new original data;

[0014] determining the enhanced data according to the new raw data.

[0015] Optionally, the preset information set includes the first key, and the merging of the plurality of to-be-merged data corresponding to the same key in the raw data according to the data weight of the raw data to obtain the merged data corresponding to the same key includes:

[0016] taking the data with the highest data weight in the plurality of first data as the merged data corresponding to the first key;

[0017] The determining of the enhanced data according to the new raw data includes:

[0018] determining the merged data as the enhanced data corresponding to the first key.

[0019] Optionally, the preset information set includes a second key related to the first key and not existing in the raw data set, and the merging of the plurality of to-be-merged data corresponding to the same key in the raw data according to the data weight of the raw data to obtain the merged data corresponding to the same key includes:

[0020] performing weighted averaging on the plurality of to-be-merged data according to the data weight of the plurality of to-be-merged data to obtain the merged data corresponding to the first key;

[0021] The determining of the enhanced data according to the new raw data includes:

[0022] calculating the enhanced data corresponding to the first key according to the merged data.

[0023] Optionally, the providing of the desensitized data to the target object includes:

[0024] broadcasting the desensitized data, and / or sending the desensitized data to the target object.

[0025] Optionally, before the storing of the to-be-processed data in the form of key-value pairs, the method further includes:

[0026] verifying the to-be-processed data to determine that the value of the to-be-processed data satisfies the value range condition in the field declaration of the to-be-processed data.

[0027] Embodiments of the present application provide a data processing apparatus, which includes:

[0028] a data storage unit configured to store, in the form of key-value pairs, to-be-processed data obtained from a plurality of data platforms; and

[0029] an enhanced data obtaining unit, configured to determine, in response to a data obtaining request containing a preset information group from a target object, enhanced data belonging to the preset information group according to the stored plurality of to-be-processed data;

[0030] a desensitized data obtaining unit, configured to provide desensitized data obtained by desensitizing the enhanced data to the target object.

[0031] Optionally, the to-be-processed data has a data weight corresponding to a data platform to which the to-be-processed data belongs, and the enhanced data obtaining unit comprises:

[0032] an original data searching unit, configured to determine, from the to-be-processed data, original data whose key matches the preset information group;

[0033] a data merging unit, configured to merge a plurality of to-be-merged data corresponding to a first key in the original data according to a data weight of the original data, to obtain merged data corresponding to the first key;

[0034] a data replacing unit, configured to replace the plurality of to-be-merged data in the original data with the merged data, to obtain new original data;

[0035] an enhanced data determining unit, configured to determine the enhanced data according to the new original data.

[0036] Optionally, the preset information group comprises the first key, and the data merging unit is specifically configured to:

[0037] take data with the highest data weight in the plurality of first data as the merged data corresponding to the first key;

[0038] the enhanced data determining unit is specifically configured to:

[0039] determine the merged data as the enhanced data corresponding to the first key.

[0040] Optionally, the preset information group comprises a second key related to the first key and not existing in the original data group, and the data merging unit is specifically configured to:

[0041] perform weighted average on the plurality of to-be-merged data according to data weights of the plurality of to-be-merged data, to obtain the merged data corresponding to the first key;

[0042] the enhanced data determining unit is specifically configured to:

[0043] calculate the enhanced data corresponding to the first key according to the merged data.

[0044] Optionally, the desensitization data acquisition unit is specifically configured to:

[0045] broadcasting the desensitization data, and / or sending the desensitization data to the target object.

[0046] Optionally, the apparatus further comprises:

[0047] The checking unit is configured to check the to-be-processed data before storing the to-be-processed data in the form of key-value pairs, and determine that the value of the to-be-processed data satisfies the value range condition in the field declaration of the to-be-processed data.

[0048] The embodiments of the present application provide a data processing method and apparatus. After obtaining a plurality of to-be-processed data from a plurality of data platforms, the to-be-processed data is stored in the form of key-value pairs, and the naming of the key in the key-value pair conforms to a preset rule. In this way, the unified storage of the data of the plurality of data platforms can be implemented, and the naming of the unified key makes the stored to-be-processed data easy to process, thereby improving the processing accuracy of the to-be-processed data. In response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group is determined according to the stored plurality of to-be-processed data. In this way, the fusion management of the data of the plurality of data platforms can be implemented. After desensitization processing of the enhanced data obtains desensitization data, the desensitization data is provided to the target object, thereby realizing the security management of the data. Therefore, the embodiments of the present application can realize the accurate fusion of the data of the plurality of data platforms, and take into account the security, thereby realizing the unified management and control of the data of the plurality of business platforms in a safe and accurate manner. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without any creative effort.

[0050] Figure 1 A flowchart of a data processing method provided by the embodiments of the present application;

[0051] Figure 2 A structural block diagram of a data processing apparatus provided by the embodiments of the present application. DETAILED DESCRIPTION

[0052] In order to make the above-mentioned objects, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings.

[0053] In the following description, a lot of specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the present application, therefore, the present application is not limited to the specific embodiments disclosed below.

[0054] With the advancement of data informatization construction, there is a need for fusion management of data provided by multiple business platforms. For example, a hospital or other enterprise with medical qualifications as multiple business platforms has created numerous patient-based data systems, each of which has accumulated a large amount of data, and the importance and necessity of data synchronization between data systems have become increasingly prominent.

[0055] Currently, the industry mainly adopts database-based and data hub-based solutions. The database-based data management synchronization method usually adopts a relational database with big data capabilities such as HIVE or a non-relational database such as MongoDB. These databases are usually driven by engineers, and the upper page system passively receives data. When data needs to be returned, the bottom layer data needs to be connected, which leads to great challenges to data security and accuracy (credibility) in actual production process. In the data hub-based solution, strong dependence on ETL (Extract-Transform-Load) extraction, conversion and loading is often required. However, such a solution has high challenges to the particularity of medical data and the diversity of data fusion rules of multiple business platforms. In terms of secondary calculation of medical data, changes of data within a period, and cluster usage, it is difficult to implement the data hub.

[0056] How to realize the fusion management of data of multiple business platforms while ensuring data accuracy is an important research at present.

[0057] Therefore, the embodiments of the present application provide a data processing method and device. After obtaining multiple to-be-processed data from multiple data platforms, the to-be-processed data are stored in the form of key-value pairs, the naming of the keys in the key-value pairs conforms to a preset rule, so that the unified storage of the data of the multiple data platforms can be realized, the naming of the unified keys makes the stored to-be-processed data easy to process, improves the processing accuracy of the to-be-processed data, and in response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group is determined according to the stored multiple to-be-processed data, so that the fusion management of the data of the multiple data platforms can be realized. After the enhanced data is desensitized to obtain desensitized data, the desensitized data is provided to the target object, so that the security management of the data is realized. Therefore, the embodiments of the present application can realize the accurate fusion of the data of multiple data platforms, and take into account the security, so as to realize the unified management and control of the data of multiple business platforms in a safe and accurate manner.

[0058] In order to better understand the technical solutions and technical effects of the present application, specific embodiments will be described in detail below in combination with the drawings.

[0059] The embodiment of the present application provides a data processing method, which is suitable for the data processing system provided by the embodiment of the present application. The data processing system comprises a memory and a controller for executing the data processing method. Referring to Figure 1 The embodiment of the present application provides a data processing method, which is suitable for the data processing system provided by the embodiment of the present application. The data processing system comprises a memory and a controller for executing the data processing method. Referring to

[0060] S101, after obtaining a plurality of to-be-processed data from a plurality of data platforms, the to-be-processed data is stored in the form of key-value pairs.

[0061] S102, in response to a data acquisition instruction containing a preset information group from a target object, determining enhanced data belonging to the preset information group according to the stored to-be-processed data.

[0062] S103, after desensitizing the enhanced data to obtain desensitized data, providing the desensitized data to the target object.

[0063] In the embodiment of the present application, the to-be-processed data can be medical data or other data. The data processing system is a management platform for to-be-processed data, which is used to manage the to-be-processed data. The plurality of data platforms are providing platforms for to-be-processed data, which are used to provide to-be-processed data to the data processing system. Taking the to-be-processed data as medical data as an example, the data platform can obtain different business data related to patients based on different business functions. These business data are to-be-processed data. The data platform can include a health management platform, a health audit platform, a patient self-submission platform, etc. The health management platform can be used by doctors and obtain medical data from doctors, or be used by medical laboratories and obtain medical data from medical laboratories. The health audit platform can be used by auditors and obtain medical data from health auditors. The patient self-submission platform can be used by patients and obtain medical data from patients. Generally, due to the difference in the acquisition approach, the to-be-processed data obtained by the health audit platform is more reliable than the to-be-processed data obtained by the health management platform, and the to-be-processed data obtained by the health management platform is more reliable than the to-be-processed data obtained by the patient self-submission platform.

[0064] In S101, the data processing system provided by the embodiment of the present application can obtain the to-be-processed data from the data platform, and can obtain multiple to-be-processed data from multiple data platforms. One to-be-processed data can be obtained from one data platform, or multiple to-be-processed data can be obtained from one data platform. The to-be-processed data obtained from the same data platform or not obtained through the data platform do not overlap with each other, which is easier to trace back when data processing is performed, and thus better adapts to the variability of medical data. The to-be-processed data from the same data platform can be generated at different times (which is conducive to better adaptation to the variability of medical data), or can be generated at the same time and belong to different patients. The to-be-processed data from different data platforms can be generated at different times, or can be generated at the same time, and can belong to different patients or the same patient. For example, blood pressure can change after taking medicine, so the blood pressure data of the same patient generated at different times can be obtained through the same data platform, which is conducive to better adaptation to the variability of medical data.

[0065] The data platform can be embodied in the form of a service, and relevant personnel (such as doctors, patients, and health auditors) as data uploaders can upload to-be-processed data to the data platform, so that the to-be-processed data is uploaded to the data processing system through the data platform. For example, doctors can upload medical data through a health management platform, health auditors can upload medical data through a health audit platform, and patients can upload medical data through a patient self-submission platform, thereby embodying the multi-source nature of medical data, and comprehensive medical data is conducive to further improvement of medical theory.

[0066] The data platform can provide a registration function, and the data uploader can obtain a pair of service identifier (id) and service key when registering through the data platform. The data platform can also not provide a registration function, and the data uploader can register through the administrator of the data platform to obtain the service identifier and the service key. The data platform can have an authentication module for providing the registration function. The service identifier is used to identify the identity of the data uploader, and can also identify the data platform to which the data uploader belongs. The service key is used to identify the identity of the data uploader. The same data uploader can register multiple service identifiers corresponding to multiple data platforms according to actual business needs. For example, a doctor can log in and upload medical data related to a prescription through the service identifier corresponding to the health management platform, or can log in and upload medical data signed by the patient through the service identifier corresponding to the patient self-upload platform.

[0067] The data platform can provide a login function, and the service identifier and service key obtained by the data uploader through registration can log in to the corresponding data platform, and then upload the to-be-processed data, and the authenticity of the source of the to-be-processed data is ensured through the authentication of the data uploader. The data platform can have an authentication module for providing login function. After determining that the login information of the data uploader is authenticated, the data platform can also provide the data uploader with historical uploaded data information. Specifically, the data platform can provide multiple login paths for data uploaders belonging to different data platforms to log in; the data platform can also provide the same login path, and after the data uploaders of different data platforms log in, the data platform can determine the data platform to which the data uploaders belong according to the service identifier of the data uploaders.

[0068] Since different data sources often have different reliability levels, the data weight of the to-be-processed data of different data sources can be determined. The to-be-processed data with a higher reliability level can have a higher weight, and the to-be-processed data with a lower reliability level can have a lower data weight, thereby greatly ensuring the accuracy and availability of the data. The data weight is automatically generated according to the data source, which has high convenience. For example, the to-be-processed data from the patient self-upload platform can have a first weight, the to-be-processed data from the health management platform can have a second weight, and the to-be-processed data from the health audit platform can have a third weight. As an example, the first weight can be 1, the second weight can be 2, and the third weight can be 3.

[0069] In S101, after obtaining the to-be-processed data from the multiple data platforms, the to-be-processed data can be stored in the form of key-value pairs, thereby realizing unified storage of data from multiple data platforms. Specifically, the to-be-processed data from the multiple data platforms can flow to the data processing system through a message queue or an interface. The data processing system can include an adaptation module for listening to the message queue to better control the uploading of to-be-processed data. The data processing system can include a storage control module for storing to-be-processed data in a storage device.

[0070] The to-be-processed data can be divided into attributes and events, and the type of the to-be-processed data can be represented according to the data type (data_type) field. For example, the data type field is 1, which represents the type of the to-be-processed data as an attribute, and the data type field is 2, which represents the type of the to-be-processed data as an event. Taking medical data as an example, the birth date of a patient can be a to-be-processed data of type attribute, data_type is 1, and the name of the key (also known as the key field) is birth_date, and the value is the date.

[0071] When the data type of the to-be-processed data is an event, the to-be-processed data can be a collection of multiple sub-data of the attribute type, and of course, the sub-data of the attribute type can be an event itself, so that the event and the attribute are nested, and the to-be-processed data can have a value list for indicating the multiple sub-data of the attribute type contained in the to-be-processed data. For example, a blood routine examination can be a to-be-processed data of the event type, the data_type is 2, the name of the key is blood_routine, and the event can include sub-data of the attribute type such as RBC, Hb, WBC, PLT, and the like.

[0072] The key of the to-be-processed data has a field declaration for defining the key of the to-be-processed data, and the field declaration can include the name of the key, the name of the to-be-processed data, a brief description, additional attributes, the data type of the value, and the like, wherein the additional attributes can include a value range condition for limiting the value range, so as to verify the to-be-processed data based on the value range condition, and the additional attributes can limit the data unit of the value, the normal range of the data, and the like, and the format can be FSON format, and the data type of the value defines the string format of the value, and can include string, float, int, date, list, dict, and the like.

[0073] When uploading the to-be-processed data, the data uploading party can query whether the key in the to-be-processed data is defined, if yes, it indicates that the field declaration of the key exists in the data processing system, and the uploading of the to-be-processed data can be directly performed, and if not, the corresponding field declaration needs to be provided, the definition of the key in the data processing system can be generated by the administrator of the data processing system, and the data processing system can have an original data management module for managing the definition of the key. In the data processing system, the naming of the key in the key-value pair conforms to the preset rule, and in actual operation, the keys of the same meaning need to be unified in principle, the field of the key needs to be as self-explanatory as possible, and for the medical field, the international disease classification system (idc10) and the common name of medical personnel can be referred to. For example, the two fields of birth_date and birth_day both represent the date of birth, and the two fields cannot be unified and combined, which does not conform to the design principle, and of course, if the information of the to-be-processed data field declaration corresponding to the two fields is consistent except the name of the key, the two fields can be considered to be of the same meaning, and the field can be adjusted to be unified.

[0074] Specifically, the additional attribute of the to-be-processed data can include a value_map as a value range condition, for example, the key is marital_status, and the additional attribute is {"value_map": {"1": "unmarried and childless", "2": "married and with child", "3": "married and without child", "4": "divorced"}}, which defines the identifiers corresponding to multiple marital statuses, and the value of the to-be-processed data can be one of the identifiers and cannot be other values, can be in an integer or string format, and cannot be in other formats, for example, the identifier corresponding to unmarried and childless is "1". Alternatively, the additional attribute of the to-be-processed data can include a valid_range as a value range condition, for example, the key is heart_rate, and the additional attribute is {"unit": "times / min", "valid_range": [30, 600]}, which defines the numerical range of the heart rate, and the format of the numerical value can be an integer or a string, and cannot be in other formats. Alternatively, the additional attribute of the to-be-processed data can include a value_list as a value range condition, for example, the key is blood_routine, and the type of the value is list, and the additional attribute is {"value_list": ["RBC", "Hb", "WBC", "PLT"]}, which defines the keys of the included sub-data as the values of the to-be-processed data, wherein RBC, Hb, WBC, and PLT are all defined attribute type fields.

[0075] Referring to Table 1, an example of field declaration provided by an embodiment of the present application is shown, taking birthdate and WBC as examples.

[0076] Example of field declaration in Table 1

[0077]

[0078] After obtaining the to-be-processed data, the compliance of the to-be-processed data can be checked to determine whether the to-be-processed data meets the preset rules. If the to-be-processed data does not meet the preset rules, the storage of the to-be-processed data can be skipped, and if the to-be-processed data meets the preset rules, the to-be-processed data can be stored in the form of key-value pairs. The data processing system can have an original data management module for providing a compliance checking function. The checking of the to-be-processed data mainly includes checking the naming of the key in the key-value pair and checking the value in the key-value pair.

[0079] Specifically, if the name of the key of the obtained to-be-processed data is not defined in the data processing system and the to-be-processed data does not contain a field declaration, it is determined that the name of the key of the to-be-processed data does not conform to the preset rule, and the compliance check of the to-be-processed data for the key fails. For example, the system defines birth_date, the to-be-processed data contains the key birth_day, and the key is not defined by the field. It is considered that the to-be-processed data does not conform to the rule. Conversely, if the name of the key of the obtained to-be-processed data is defined in the data processing system or the to-be-processed data contains a field declaration, it is determined that the name of the key of the to-be-processed data conforms to the preset rule, and the compliance check of the to-be-processed data for the key passes. The unified naming of the key makes the stored to-be-processed data easy to process, and improves the processing accuracy of the to-be-processed data.

[0080] Specifically, if the value of the obtained to-be-processed data satisfies the value range condition in the field declaration of the to-be-processed data, it is determined that the value of the to-be-processed data conforms to the preset rule, and the compliance check of the to-be-processed data for the value passes. Conversely, if the value of the obtained to-be-processed data does not satisfy the value range condition in the field declaration of the to-be-processed data, it is determined that the value of the to-be-processed data does not conform to the preset rule, and the compliance check of the to-be-processed data for the value fails. For example, the data format requirement of the value of birth_date is xxxx-xx-xx, and the value of birth_date in the to-be-processed data is “2021-10-01”. The value does not conform to the preset rule, and the data check for the value fails.

[0081] The compliance check of the to-be-processed data can be determined according to the corresponding additional attribute, or the to-be-processed data can be checked by using a check code block. The check code block rule can be defined by the field declaration. In this way, when the general determination method cannot be implemented, the to-be-processed data can be checked by using the check code block. Of course, when the check code block exists, the to-be-processed data can also be determined generally.

[0082] In implementation, the setting rule of the code block can include: 1) the code conforms to the python3 syntax (except return value); 2) the last return value is required (imagine that all submitted codes are in a function), and the return value will be used as the result of the calculation value; 3) the code can consider the exception condition or not, and the bottom strategy is considered when the exception condition is considered; 4) if the default value is not set, the value will not be updated if an error occurs, and the value will be updated in other cases; 5) the code is prohibited to contain the strings of "exec", "eval", "import", "__import__", "global"; 6) the package that needs to be imported needs to be in the whitelist and is transmitted by another parameter (need_pkg); 7) the tab indentation is replaced by 4 spaces. The following is an example of a code block for checking the code:

[0083]

[0084]

[0085] In S102, in response to a data acquisition instruction containing a preset information group from a target object, enhanced data belonging to the preset information group can be determined according to stored to-be-processed data, so as to realize fusion management of the to-be-processed data.

[0086] In the embodiment of the application, the enhanced data can be determined according to the to-be-processed data by using the data acquisition instruction. The data acquisition instruction can be generated according to the instruction of a data user. The data user is a user who has a demand for using the to-be-processed data, and can be a user different from a data uploading party, or can be the data uploading party, that is, the same user can be determined as the data uploading party according to the uploading operation of the to-be-processed data, or can be determined as the data user according to the acquisition operation of the to-be-processed data. In the embodiment of the application, the target object is the data user.

[0087] The data user can also register and log in through the data processing system provided by the embodiment of the application, and the data processing system can authenticate the identity of the data user. After the data user is authenticated, a preset information group from the data user can be obtained, the preset information group can be defined in the form of logical structure schema information, the data processing system can include an enhanced data management module, and the data user can define the preset information group through the enhanced data management module to call data. The preset information group can be predefined, which can be determined artificially or obtained by combining the keys of data according to rules. Each key in the preset information group has different necessary attributes as the value of the key according to the type. For example, the preset information group can be a basic information group, and the basic information group can include multiple keys such as name, gender, date of birth, height, weight, and body fat rate. The preset information group can also be a blood routine information group, and the blood routine information group can include multiple keys such as WBC, RBC, and PLT. If the preset information group is a basic information group, the enhanced data is the value corresponding to the multiple keys of the name, gender, date of birth, height, weight, and body fat rate of the patient.

[0088] When the to-be-processed data has a data weight corresponding to the data platform to which the to-be-processed data belongs, the enhanced data can be determined according to the data weight of the to-be-processed data. Specifically, the keys of the data in the to-be-processed data can be determined first, and then the original data matching the preset information group is determined. Then, according to the data weight of the original data, the multiple to-be-merged data corresponding to the first key in the original data are merged to obtain the merged data corresponding to the first key. After replacing the multiple to-be-merged data in the original data with the merged data, new original data is obtained, and the enhanced data is determined according to the new original data. The original data matching the preset information group can be obtained through the enhanced data management module.

[0089] Specifically, the data with the highest data weight in the multiple to-be-merged data is taken as the merged data corresponding to the first key, so that the data with the highest accuracy can be taken as the merged data corresponding to the first key, or the multiple to-be-merged data are weighted and averaged according to the data weight of the multiple to-be-merged data to obtain the merged data corresponding to the first key, so that the merged data corresponding to the first key can be obtained by comprehensively considering each data.

[0090] Before determining the enhanced data according to the data weight of the data to be processed, the data type of the enhanced data to which the data acquisition request is directed can also be determined, and the data type of the enhanced data can include non-computational and computational. Where the data type of the enhanced data is non-computational, the key in the preset information group exists in the original data, and the original data can be directly used as the corresponding enhanced data without calculation, for example, the enhanced data includes non-computational data such as height, date of birth, etc. Where the data type of the enhanced data is computational, the key in the preset information group does not exist in the original data, and the corresponding enhanced data needs to be calculated from the data in the original data, for example, the enhanced data includes computational data such as Body Mass Index (BMI), etc. BMI needs to be calculated from height and weight.

[0091] The data type of the enhanced data to which the data acquisition request is directed can be determined by the type identifier of the data acquisition request, and the non-computational corresponds to the first type identifier, and the computational corresponds to the second type identifier, for example, the first type identifier is 1, and the second type identifier is 2. When the data type of the enhanced data to which the data acquisition request is directed is computational, the data acquisition request also corresponds to a calculation code, an import package list, a default value information, a calculation frequency, a parameter list, etc. Wherein the parameter list can include at least one key of the enhanced data.

[0092] Specifically, if the data acquisition request is for non-computational data, taking the first key in the preset information group as an example, the original data includes multiple to-be-merged data corresponding to the first key. After replacing the multiple to-be-merged data with the merged data corresponding to the first key, the merged data corresponding to the first key can be used as the enhanced data corresponding to the first key. For example, the first key is name, gender, date of birth, weight, or height, etc. Of course, the merged data can be the data with the highest weight in the multiple to-be-merged data, or the weighted data obtained by weighting the multiple to-be-merged data. Similarly, if the data acquisition request is for non-computational data, taking the third key in the preset information group as an example, the original data includes a data corresponding to the third key, which can be used as the enhanced data corresponding to the third key. For example, the third key is name, gender, date of birth, weight, or height, etc.

[0093] Specifically, if the data acquisition request is for calculation type data, taking the second key in the preset information group as an example, the second key is related to the first key and does not exist in the original data, the original data includes a plurality of to-be-merged data corresponding to the first key, after replacing the plurality of to-be-merged data with the merged data corresponding to the first key, the enhanced data corresponding to the first key can be calculated according to the merged data corresponding to the first key. Of course, the merged data can be the data with the highest weight in the plurality of to-be-merged data, or the weighted data obtained by weighting the plurality of to-be-merged data. Calculating the enhanced data corresponding to the first key according to the merged data corresponding to the first key can specifically be determining the enhanced data corresponding to the first key according to the merged data corresponding to the first key and a calculation strategy, or determining the enhanced data corresponding to the first key according to the merged data corresponding to the first key, the data corresponding to the fourth key, and the calculation strategy, wherein the data corresponding to the fourth key can be the original data or the merged data, and the calculation strategy can be determined by a calculation code. For example, the second key is BMI, the first key is height, and the fourth key is weight.

[0094] Similarly, if the data acquisition request is for calculation type data, taking the fifth key in the preset information group as an example, the fifth key is related to the third key and does not exist in the original data, and the original data includes one data corresponding to the third key, then the enhanced data corresponding to the fifth key can be calculated according to the data. Calculating the enhanced data corresponding to the fifth key according to the data can specifically be determining the enhanced data corresponding to the fifth key according to the data and a calculation strategy, or determining the enhanced data corresponding to the fifth key according to the data, the data corresponding to the sixth key, and the calculation strategy, wherein the data corresponding to the sixth key can be the original data or the merged data, and the calculation strategy can be determined by a calculation code. For example, the fifth key is BMI, the third key is height, and the sixth key is weight.

[0095] In addition, after determining the enhanced data, if it is monitored that the original data is updated, the new enhanced data can be determined according to the updated original data, and the update of the original data is triggered by uploading new to-be-processed data related to the enhanced data by the data platform. Specifically, the new original data can be used as new to-be-merged data, the new merged data is determined according to the new to-be-merged data, and the new enhanced data is determined according to the new merged data. For example, the merged data is the to-be-merged data with the highest weight, and the data weight of the new original data is greater than the data weight of the merged data, then the new original data can be used as the new merged data, and the new merged data can be used as the new enhanced data or as the basis for calculating the new enhanced data.

[0096] It should be noted that a value range condition can be determined for the enhanced data, if the calculated enhanced data does not meet the value range condition, the data can be recorded as abnormal data, or the value of the data can be adjusted to a pre-set default value, and the default value is determined according to a pre-set bottom-up strategy.

[0097] As described above, the same original data can correspond to one enhanced data or multiple enhanced data, improving the overall planning of the to-be-processed data and the optimization space of data processing. In addition, after the enhanced data is generated, the enhanced data can be stored, and the enhanced data can be used as the original data for next time calculation of enhanced data to obtain new enhanced data. The step of calculating the enhanced data can be implemented by a calculation module in the data processing system. The calculation module can be an independent calculation component to facilitate the increase of calculation capacity, or a functional module with calculation function in the data processing system.

[0098] In S103, after the enhanced data is desensitized to obtain the desensitized data, the desensitized data can be provided to the target object. The desensitized data after desensitization is provided to the target object, reducing the leakage of sensitive data and improving the security of data in the data processing system, thereby realizing the security management of data and accurate data provision. The desensitization of the enhanced data can be performed according to the pre-defined desensitization rule. The desensitization rules corresponding to different target objects can be different. The desensitization can be realized by a desensitization module in the data processing system.

[0099] Providing the desensitized data to the target object can specifically be broadcasting the desensitized data and / or sending the desensitized data to the target object. Broadcasting the desensitized data can be realized by a broadcasting module in the data processing system. In actual operation, the desensitized data can be first sent to the target object, and then the change of the desensitized data is monitored through broadcasting. When the desensitized data changes, new desensitized data is provided to the target object through broadcasting, so that the target object can obtain accurate desensitized data.

[0100] The embodiment of the present application provides a data processing method. After obtaining a plurality of to-be-processed data from a plurality of data platforms, the to-be-processed data is stored in the form of key-value pairs. The naming of the key in the key-value pair conforms to a preset rule. In this way, the unified storage of the data of the plurality of data platforms can be realized. The unified naming of the key makes the stored to-be-processed data easy to process, improves the processing accuracy of the to-be-processed data, and responds to a data acquisition request containing a preset information group from a target object. According to the stored plurality of to-be-processed data, the enhanced data belonging to the preset information group is determined. In this way, the fusion management of the data of the plurality of data platforms can be realized. After the enhanced data is desensitized to obtain desensitized data, the desensitized data is provided to the target object, thereby realizing the security management of data. Therefore, the embodiment of the present application can realize the accurate fusion of the data of the plurality of data platforms, and take into account the security, thereby realizing the unified management and control of the secure and accurate data of the plurality of service platforms.

[0101] Based on the data processing method provided in the embodiment of the present application, the embodiment of the present application further provides a data processing device. ReferenceFigure 2 As shown, a structural block diagram of a data processing apparatus provided by an embodiment of the present application is provided, and the apparatus can include:

[0102] A data storage unit 110, configured to store, in a form of a key-value pair, a plurality of to-be-processed data from a plurality of data platforms after obtaining the plurality of to-be-processed data; a key in the key-value pair is named in accordance with a preset rule;

[0103] An enhanced data acquisition unit 120, configured to determine, in response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group according to the plurality of to-be-processed data stored;

[0104] A desensitized data acquisition unit 130, configured to provide, after desensitizing the enhanced data to obtain desensitized data, the desensitized data to the target object.

[0105] Optionally, the to-be-processed data has a data weight corresponding to a data platform to which the to-be-processed data belongs, and the enhanced data acquisition unit includes:

[0106] An original data searching unit, configured to determine, from the to-be-processed data, original data whose key matches the preset information group;

[0107] A data merging unit, configured to merge, according to a data weight of the original data, a plurality of to-be-merged data corresponding to a first key in the original data to obtain merged data corresponding to the first key;

[0108] A data replacing unit, configured to replace the plurality of to-be-merged data in the original data with the merged data to obtain new original data;

[0109] An enhanced data determining unit, configured to determine the enhanced data according to the new original data.

[0110] Optionally, the preset information group includes the first key, and the data merging unit is specifically configured to:

[0111] merge, as the merged data corresponding to the first key, data having a highest data weight in the plurality of first data;

[0112] The enhanced data determining unit is specifically configured to:

[0113] determine the merged data as the enhanced data corresponding to the first key.

[0114] Optionally, the preset information group includes a second key related to the first key and not existing in the original data group, and the data merging unit is specifically configured to:

[0115] weighting average the plurality of to-be-merged data according to data weights of the plurality of to-be-merged data, to obtain merged data corresponding to the first key;

[0116] The enhanced data determination unit is specifically configured to:

[0117] The enhanced data corresponding to the first key is calculated according to the merged data.

[0118] Optionally, the desensitization data acquisition unit is specifically configured to:

[0119] The desensitization data is broadcasted, and / or the desensitization data is sent to the target object.

[0120] Optionally, the apparatus further comprises:

[0121] The verification unit is configured to, before the to-be-processed data is stored in the form of key-value pairs, verify the to-be-processed data, and determine that a value of the to-be-processed data satisfies a value range condition in a field declaration of the to-be-processed data.

[0122] The embodiments of the present application provide a data processing apparatus. After a plurality of to-be-processed data from a plurality of data platforms is acquired, the to-be-processed data is stored in the form of key-value pairs, and the naming of the key in the key-value pair conforms to a preset rule. In this way, unified storage of data of the plurality of data platforms can be implemented, and the naming of the unified key makes the stored to-be-processed data easy to process, thereby improving the processing accuracy of the to-be-processed data. In response to a data acquisition request containing a preset information group from a target object, enhanced data belonging to the preset information group is determined according to the stored plurality of to-be-processed data. In this way, fusion management of data of the plurality of data platforms can be implemented. After the enhanced data is desensitized to obtain desensitization data, the desensitization data is provided to the target object, thereby realizing security management of data. Therefore, the embodiments of the present application can realize accurate fusion of data of multiple data platforms, and take into account security, thereby realizing unified management and control of secure and accurate data of multiple service platforms.

[0123] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments.

[0124] The above merely describes preferred embodiments of the present application, and the present application is not limited to the above. Any person skilled in the art, without departing from the technical scope of the present application, can make many possible changes and modifications to the technical solutions of the present application, or modify equivalent embodiments with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application, without departing from the technical scope of the present application, still falls within the protection scope of the technical solutions of the present application.

Claims

1. A data processing method, characterized by, The method comprises the following steps: After obtaining a plurality of to-be-processed data from a plurality of data platforms, the to-be-processed data are stored in the form of key-value pairs; The naming of the key in the key-value pair conforms to a preset rule; the preset rule is used to uniformly name the keys with the same meaning; if the name of the key is not defined and the to-be-processed data do not contain field declaration, it is determined that the name of the key does not conform to the preset rule; if the name of the key is defined or the to-be-processed data contain field declaration, it is determined that the name of the key conforms to the preset rule; In response to a data acquisition request containing a preset information group from a target object, the data type of enhanced data to which the data acquisition request is directed is determined; the data type includes non-computing type and computing type; The enhanced data belonging to the preset information group is determined according to the data type and the stored plurality of to-be-processed data; After the enhanced data is desensitized to obtain desensitized data, the desensitized data is provided to the target object; The to-be-processed data have data weights corresponding to the data platforms to which the to-be-processed data belong; the determination of the enhanced data belonging to the preset information group according to the data type and the stored plurality of to-be-processed data comprises the following steps: Original data in which the key of the data matches the preset information group is determined from the to-be-processed data; A plurality of to-be-merged data corresponding to a first key in the original data are merged according to the data weights of the original data to obtain merged data corresponding to the first key; The plurality of to-be-merged data in the original data are replaced by the merged data to obtain new original data; The enhanced data is determined according to the new original data; If the data acquisition request is directed to non-computing type data, the preset information group includes the first key, a plurality of to-be-merged data corresponding to the same key in the original data are merged according to the data weights of the original data to obtain merged data corresponding to the same key, which comprises the following steps: The data with the highest data weight in the plurality of to-be-merged data corresponding to the first key is taken as the merged data corresponding to the first key; The determination of the enhanced data according to the new original data comprises the following steps: The merged data is determined as the enhanced data corresponding to the first key; If the data acquisition request is directed to computing type data, the preset information group includes a second key related to the first key and not existing in the original data, a plurality of to-be-merged data corresponding to the same key in the original data are merged according to the data weights of the original data to obtain merged data corresponding to the same key, which comprises the following steps: The plurality of to-be-merged data are weightedly averaged according to the data weights of the plurality of to-be-merged data to obtain the merged data corresponding to the first key; The determination of the enhanced data according to the new original data comprises the following steps: The enhanced data corresponding to the first key is calculated according to the merged data.

2. The method of claim 1, wherein, The provision of the desensitized data to the target object comprises the following steps: The desensitized data is broadcasted and / or the desensitized data is sent to the target object.

3. The method of claim 1, wherein, Before the storing the to-be-processed data in the form of key-value pairs, the method further comprises: checking the to-be-processed data to determine whether the value of the to-be-processed data satisfies the value range condition in the field declaration of the to-be-processed data.

4. A data processing apparatus, characterized by, Comprise: a data storage unit, configured to store the to-be-processed data in the form of key-value pairs after obtaining a plurality of to-be-processed data from a plurality of data platforms; the naming of the key in the key-value pair conforms to a preset rule; the preset rule is used to uniformly name the keys with the same meaning; if the name of the key is not defined and the to-be-processed data does not contain a field declaration, it is determined that the name of the key does not conform to the preset rule; if the name of the key has been defined or the to-be-processed data contains a field declaration, it is determined that the name of the key conforms to the preset rule; a data type determination unit configured to determine the data type of enhanced data targeted by a data acquisition request from a target object in response to the data acquisition request containing a preset information group; the data type comprises a non-computing class and a computing class; an enhanced data acquisition unit configured to determine the enhanced data belonging to the preset information group according to the data type and the stored plurality of to-be-processed data; a desensitization data acquisition unit configured to provide desensitization data to the target object after desensitizing the enhanced data to obtain the desensitization data; The to-be-processed data has a data weight corresponding to the data platform to which it belongs, and the enhanced data acquisition unit comprises: an original data searching unit configured to determine original data with a key matching the preset information group from the to-be-processed data; a data merging unit configured to merge a plurality of to-be-merged data corresponding to a first key in the original data according to the data weight of the original data to obtain merged data corresponding to the first key; a data replacement unit configured to replace the plurality of to-be-merged data in the original data with the merged data to obtain new original data; an enhanced data determination unit configured to determine the enhanced data according to the new original data; if the data acquisition request is for non-computing class data, the preset information group includes the first key, and the data merging unit is specifically configured to: take the data with the highest data weight in the plurality of to-be-merged data corresponding to the first key as the merged data corresponding to the first key; the enhanced data determination unit is specifically configured to: determine the merged data as the enhanced data corresponding to the first key; if the data acquisition request is for computing class data, the preset information group includes a second key related to the first key and not existing in the original data, and the data merging unit is specifically configured to: weight average the plurality of to-be-merged data according to the data weight of the plurality of to-be-merged data to obtain the merged data corresponding to the first key; the enhanced data determination unit is specifically configured to: calculate the enhanced data corresponding to the first key according to the merged data.

Citation Information

Patent Citations

  • Data desensitization processing method and device, equipment and storage medium

    CN113792344A

  • System and methods for implementing a key-value data store

    US20220043585A1