Data management method and device, computer equipment and storage medium
By collecting insurance data from multiple data sources and using data quality management tools, sensitive identification models and data sharing strategies to process insurance data, the standardization and reliability issues of traditional insurance industry data management systems are resolved, efficient and secure data management and utilization are achieved, and system stability and data utilization efficiency are improved.
Patent Information
- Application Number
- CN202510591633.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-23
AI Technical Summary
The data management system in the traditional insurance industry lacks standardization and poor reliability, making it difficult to achieve comprehensive management and effective utilization of business data, resulting in low system stability and affecting risk assessment, customer service and marketing effectiveness.
By collecting business data related to business needs from multiple data sources, using data quality management tools for pre-processing, identifying sensitivity levels based on sensitive identification models, and storing the sensitivity level information in the data lake, data sharing is carried out in combination with data sharing strategies.
It achieves efficient, secure and reliable management of business data, improves the intelligence and standardization of data management, and enhances system stability and data utilization efficiency.
Smart Images

Figure CN120687512A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the field of financial technology, and in particular to data management methods, devices, computer equipment and storage media. Background Art
[0002] In the traditional insurance service model, data is a core element of insurance industry operations, and its management method directly affects the accuracy of insurance companies' risk assessments, the quality of customer service, and the effectiveness of marketing. However, the current data management systems in the insurance industry generally suffer from insufficient standardization and poor reliability, making it difficult to achieve comprehensive management and effective utilization of business data, which in turn leads to low system stability and limited business decision-making efficiency. Specifically, traditional data management methods often lack unified data standards, data quality monitoring mechanisms, and data security measures, making data prone to errors, omissions, or misuse during collection, storage, processing, and analysis, and unable to provide insurance companies with accurate, timely, and valuable information support.
[0003] For example, in the insurance claims process, traditional data management methods can lead to inaccurate and incomplete claim data entry, which in turn affects the fairness and efficiency of claim decision-making. If a customer submits a claim with key medical expense data that is lower due to an entry error, the insurance company may make an unreasonable claim decision based on the erroneous data, harming the customer's interests. This may also cause customers to question the insurance company's professionalism and damage its reputation.
[0004] Therefore, the limitations of traditional data management methods not only restrict insurance companies' ability to improve risk assessment, customer service, and marketing, but also hinder the overall development of the insurance industry. Therefore, there is an urgent need for an efficient, reliable, and intelligent data governance solution to achieve comprehensive management and effective utilization of business data, improve system stability, and promote the digital transformation and intelligent upgrade of the insurance industry. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to propose a data management method, apparatus, computer equipment and storage medium to solve the technical problems that existing data management systems generally have insufficient standardization and poor reliability, and it is difficult to achieve comprehensive management and effective utilization of business data.
[0006] In a first aspect, a data management method is provided, comprising:
[0007] Collect business data related to business needs from multiple preset data sources;
[0008] Performing data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data;
[0009] Performing sensitivity level identification on the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data;
[0010] Based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a preset data lake;
[0011] Get the preset data sharing policy;
[0012] Data sharing processing is performed on the target business data in the data lake based on the data sharing strategy.
[0013] In a second aspect, a data management device is provided, comprising:
[0014] The collection module is used to collect business data related to business needs from multiple preset data sources;
[0015] A preprocessing module, configured to perform data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data;
[0016] An identification module is used to identify the sensitivity level of the target business data based on a preset sensitivity identification model to obtain the sensitivity level information of the target business data;
[0017] A first storage module is configured to store the target business data in a preset data lake based on a target storage method corresponding to the sensitivity level information;
[0018] A first acquisition module is used to acquire a preset data sharing strategy;
[0019] A sharing module is used to perform data sharing processing on the target business data in the data lake based on the data sharing strategy.
[0020] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned data management method when executing the computer program.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned data management method are implemented.
[0022] In the solution implemented by the above-mentioned data management method, device, computer equipment and storage medium, business data related to business needs are first collected from preset multiple data sources; then, data preprocessing is performed on the business data based on a preset data quality management tool to obtain corresponding target business data; then, sensitivity level identification is performed on the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data; subsequently, based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a preset data lake; further, a preset data sharing strategy is obtained; finally, data sharing processing is performed on the target business data in the data lake based on the data sharing strategy. This application collects business data related to business needs from multiple data sources, and then pre-processes the business data based on the use of data quality management tools to obtain target business data. Thereafter, the sensitivity level of the target business data is identified based on the use of a sensitive identification model to obtain the sensitivity level information of the target business data. Based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a data lake. Finally, based on the use of a data sharing strategy, the target business data in the data lake is subjected to data sharing processing. In this way, this application automatically and intelligently realizes the comprehensive management and effective utilization of business data based on the combined use of data quality management tools, sensitive identification models, data lakes and data sharing strategies, provides a more efficient, secure and reliable solution for the system's data management, effectively improves the intelligence, reliability and standardization of data management, and thus helps to improve the stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0025] Figure 2 is a flow chart of an embodiment of a data management method according to the present application;
[0026] Figure 3 is a structural diagram of an embodiment of a data management device according to the present application;
[0027] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0031] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0032] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0033] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0034] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0035] It should be noted that the data management method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the data management device is generally set in the server / terminal device.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0037] Continue to refer Figure 2 , shows a flow chart of an embodiment of the data management method according to the present application. According to different needs, the order of the steps in the flow chart can be changed, and some steps can be omitted. The data management method provided in the embodiment of the present application can be applied to any scenario that requires data management, and the data management method can be applied to products in these scenarios, for example, data management in the financial field. The data management method includes the following steps:
[0038] Step S201 : collecting business data related to business needs from a plurality of preset data sources.
[0039] In this embodiment, the data management method is executed on the electronic device (eg Figure 1The server / terminal device shown in the figure) can obtain business data related to business needs through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G / 5G connections, Wi-Fi connections, Bluetooth connections, Wi-MAX connections, Zigbee connections, UWB (Ultra Wide Band) connections, and other wireless connection methods currently known or developed in the future. The execution entity of this application is specifically a data management system, which can be simply referred to as the system. The above-mentioned data sources may include at least the core business system database, the document file system uploaded by customers, the external market data API, etc. The above-mentioned business data may include structured data and unstructured data that match business needs. Specifically, by conducting a comprehensive inventory of data from different business systems and automatically scanning the data table structure, file content, etc., the type information of the data is identified and recorded. For example, data containing fields such as customer name and contact information is marked as customer data; data containing fields such as policy number, insurance amount, and insurance period is marked as policy data. Exemplary. This application can be applied to data management scenarios of insurance companies. Insurance companies access data from online insurance systems, offline counter business systems, and external credit reporting agencies. Through data collection tools, they can identify different types of data, such as basic customer information (name, ID number, etc.), insurance information (insured products, premiums, etc.), and claims information (claim reasons, claim amounts, etc.), to obtain business data related to business needs.
[0040] The aforementioned data collection tools can be configured with filtering rules and mapping relationships for data sources to clarify the purpose and scope of data collection. For example, when collecting customer data, only necessary fields such as name, contact information, and insurance information can be collected, filtering out information not relevant to the business. This ensures that only necessary information closely related to business needs is collected. For example, during the online insurance application process, an insurance company can configure rules at the data access layer to collect only basic customer information and insurance intentions, avoiding the collection of unnecessary customer privacy information.
[0041] Step S202: pre-process the business data based on a preset data quality management tool to obtain corresponding target business data.
[0042] In this embodiment, the above-mentioned specific implementation process of preprocessing the business data based on the preset data quality management tool to obtain the corresponding target business data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0043] Step S203: performing sensitivity level identification on the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data.
[0044] In this embodiment, the target business data can be input into the sensitivity identification model described above. The sensitivity identification model will perform real-time sensitivity assessment processing on the target business data and output sensitivity level information of the target business data. The sensitivity level information may include a high sensitivity level, a medium sensitivity level, or a low sensitivity level. The model construction process of the sensitivity identification model described above will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0045] Step S204: Based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a preset data lake.
[0046] In this embodiment, the specific implementation process of storing the target business data in a preset data lake based on the target storage method corresponding to the sensitivity level information will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0047] Step S205: Obtain a preset data sharing strategy.
[0048] In this embodiment, the policy content of the above-mentioned data sharing strategy includes: formulating a data sharing policy, and the corresponding specific implementation: clarifying the principles, scope and process of data sharing in the system. For example, specify which data can be shared between which departments, and which approval procedures are required before sharing. Establish a data sharing platform (or metadata management tool), and the corresponding specific implementation: build a safe and reliable data sharing platform based on the data lake. The data sharing platform has access control, data encryption and audit functions. The data sharing platform is used to maintain the definition, format and purpose of data, and promote the sharing and understanding of data among different departments. Use data lineage tools to maintain metadata and record the source, format, purpose and responsible person of the data. Use metadata management tools to implement data classification and label management to promote data sharing and understanding.
[0049] Step S206 performs data sharing processing on the target business data in the data lake based on the data sharing strategy.
[0050] In this embodiment, data sharing processing of the target business data in the above-mentioned data lake can be completed according to the data sharing policy and data sharing platform corresponding to the above-mentioned data sharing strategy, so as to achieve sharing and understanding of the target business data.
[0051] This application first collects business data related to business needs from preset multiple data sources; then pre-processes the business data based on a preset data quality management tool to obtain corresponding target business data; then identifies the sensitivity level of the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data; subsequently, stores the target business data in a preset data lake based on a target storage method corresponding to the sensitivity level information; further obtains a preset data sharing strategy; and finally, performs data sharing processing on the target business data in the data lake based on the data sharing strategy. This application collects business data related to business needs from multiple data sources, and then pre-processes the business data based on the use of data quality management tools to obtain target business data. Thereafter, the sensitivity level of the target business data is identified based on the use of a sensitive identification model to obtain the sensitivity level information of the target business data. Based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a data lake. Finally, based on the use of a data sharing strategy, the target business data in the data lake is subjected to data sharing processing. In this way, this application automatically and intelligently realizes the comprehensive management and effective utilization of business data based on the combined use of data quality management tools, sensitive identification models, data lakes and data sharing strategies, provides a more efficient, secure and reliable solution for the system's data management, effectively improves the intelligence, reliability and standardization of data management, and thus helps to improve the stability of the system.
[0052] In some optional implementations, step S202 includes the following steps:
[0053] The business data is cleansed based on the data quality management tool to obtain cleansed first business data.
[0054] In this embodiment, the data quality management tool is a pre-built automated tool with data quality management functions. The data cleaning process refers to removing missing values, duplicate values and noise data in the business data, thereby obtaining the first business data.
[0055] Perform data format verification on the first business data to filter out second business data that passes the data format verification from the first business data.
[0056] In this embodiment, the data format verification refers to verifying the format of the first business data using a regular expression to detect the accuracy and completeness of the data. Then, based on the obtained data format verification result, the second business data that passes the data format verification is screened out from the first business data.
[0057] A mandatory item check is performed on the second business data to filter out third business data that passes the mandatory item check from the second business data.
[0058] In this embodiment, mandatory items about business data are pre-set according to actual needs, and then the mandatory item matching detection is performed on the above-mentioned second business data based on the mandatory items. Then, according to the obtained matching detection results, the third business data that passes the mandatory item detection is filtered out from the second business data.
[0059] Perform data standardization processing on the third business data to obtain processed fourth business data.
[0060] In this embodiment, the format of the third service data may be replaced by using a predefined standard format, thereby obtaining standardized fourth service data.
[0061] The fourth business data is used as the target business data.
[0062] This application performs data cleaning on the business data based on the data quality management tool to obtain the cleaned first business data; then performs data format verification on the first business data to filter out the second business data that passes the data format verification from the first business data; then performs mandatory item detection on the second business data to filter out the third business data that passes the mandatory item detection from the second business data; subsequently performs data standardization on the third business data to obtain the processed fourth business data; finally, the fourth business data is used as the target business data. This application performs data cleaning, data format verification, mandatory item detection and data standardization on the business data based on the use of a data quality management tool, thereby automatically and accurately completing the pre-processing of the business data, ensuring the accuracy and reliability of the generated target business data, and achieving quality control of the business data, which is conducive to providing high-quality data for subsequent data storage and data use.
[0063] In some optional implementations of this embodiment, before step S203, the electronic device may further perform the following steps:
[0064] Obtain pre-collected initial business data.
[0065] In this embodiment, the initial business data may be all business data related to the business needs collected in a pre-collected historical time period. The initial business data may include both structured and unstructured data. Furthermore, there is no specific limitation on the time period of the historical time period and it may be determined based on actual business needs. For example, the past year may be used.
[0066] The initial business data is cleaned and sensitivity-level labeled to obtain corresponding processed data.
[0067] In this embodiment, the above-mentioned cleaning process may include: processing missing values, duplicate values and noise data. In addition, the unstructured data (such as text) is pre-processed by word segmentation, stop word removal, etc. In addition, the above-mentioned sensitivity level labeling process includes: first, defining the sensitivity level label according to business needs and compliance standards (such as GDPR, CCPA). For example: Low sensitivity level: does not contain PII or sensitive business data. Medium sensitivity level: contains part of PII (such as user name, email address). High sensitivity level: contains complete PII, financial or medical information. Then, a semi-automatic labeling method is used, combined with a rule engine (such as regular expression) to pre-label part of the data, and then the sensitivity level labeling process is performed by manual review.
[0068] The processed data is subjected to sample construction processing based on a preset feature screening strategy to obtain corresponding sample data.
[0069] In this embodiment, the strategy content of the above-mentioned feature screening strategy includes: 1) Feature extraction. For structured data: Data format: field type (such as string, number, date), length, character distribution. Content mode: Regular expression matching results (such as whether it contains ID card number, mobile phone number). Context information: Data source (such as system log, user table), storage location (such as database table name). For unstructured data: Text features: Keyword extraction (such as using TF-IDF or word embedding). Named entity recognition (NER) extracts sensitive entities (such as names, addresses). Metadata features: Document length, number of paragraphs, special character ratio. 2) Feature encoding. One-hot encoding (One-Hot Encoding) or label encoding (Label Encoding) is performed on classification features (such as data source). Text features are vectorized (such as bag-of-words model, Word2Vec, BERT embedding). 3) Feature selection. Use statistical methods (such as chi-square test, mutual information) or model importance (such as random forest feature importance) to screen key features. Specifically, the processed data may be subjected to feature screening according to the feature screening steps corresponding to the policy content of the feature screening policy to obtain key feature data, and the key feature data may be used as the sample data.
[0070] Call the preset machine learning model.
[0071] In this embodiment, the selection of the above-mentioned machine learning model is not specifically limited and can be determined according to actual business needs. For example, any one of a decision tree, random forest, convolutional neural network, long short-term memory network, and transformer model can be used.
[0072] Performing model training and optimization processing on the machine learning model based on the sample data until a specified model that meets preset requirements is obtained;
[0073] In this embodiment, the sample data can be divided into a training set, a validation set, and a test set according to a preset ratio (e.g., 70% training, 15% validation, and 15% testing). The process of the above-mentioned model training and optimization processing includes: (1) Training process: using the training set to train the model and adjust hyperparameters (e.g., learning rate, tree depth, and number of network layers). And evaluating the model performance on the validation set to prevent overfitting. Among them, the evaluation indicators include: Accuracy: the proportion of samples that are correctly classified. Precision: the proportion of data predicted to be sensitive that is actually sensitive. Recall: the proportion of data that is correctly predicted that is actually sensitive. F1 score: the harmonic mean of precision and recall. (2) Model evaluation and optimization. Model evaluation includes: evaluating the model performance on the test set to ensure its generalization ability on unseen data. And analyzing the confusion matrix to identify at which sensitivity levels the model performs poorly. Model optimization includes: Feature engineering optimization: adding or adjusting features (e.g., adding new contextual information). Model Tuning: Adjust model hyperparameters (e.g., the number of trees in a random forest, the size of the hidden layer in an LSTM). Experiment with different model architectures (e.g., switching from a decision tree to a random forest). Ensemble Learning: Combine the predictions of multiple models (e.g., random forest and BERT) to improve accuracy. Ultimately, through model training and optimization, construct a specific model that meets evaluation requirements.
[0074] The designated model is used as the sensitive recognition model.
[0075] In this embodiment, the trained sensitivity recognition model can be deployed as an API service to assess the sensitivity of new data in real time. Furthermore, metrics such as accuracy and recall can be tracked in the production environment of the sensitivity recognition model. Changes in data distribution can be monitored to prevent model performance degradation (e.g., data drift). Furthermore, user feedback (e.g., manual review results) can be collected to regularly update training data and models.
[0076] This application obtains pre-collected initial business data; then cleans and labels the initial business data with a sensitivity level to obtain corresponding processed data; and performs sample construction processing on the processed data based on a preset feature screening strategy to obtain corresponding sample data; then calls a preset machine learning model; subsequently, performs model training and optimization processing on the machine learning model based on the sample data until a specified model that meets the preset requirements is obtained; finally, the specified model is used as the sensitive recognition model, thereby achieving efficient and accurate completion of the construction processing of the sensitive recognition model, improving the construction efficiency of the sensitive recognition model, and ensuring the recognition effect of the obtained sensitive recognition model.
[0077] In some optional implementations, step S204 includes the following steps:
[0078] Determine whether the sensitivity level information is a highly sensitive level.
[0079] In this embodiment, the content of the sensitivity level information may include a high sensitivity level, a medium sensitivity level, or a low sensitivity level. The content of the sensitivity level information may be determined by performing content recognition on the sensitivity level information.
[0080] If the sensitivity level information is highly sensitive, the target business data is stored in a pre-built encrypted storage area in the data lake.
[0081] In this embodiment, the encrypted storage area refers to a pre-built secure storage area within the data lake with strict access controls, such as an encrypted database. If the sensitive information is detected as highly sensitive, the target business data is stored in the secure storage area. For example, an insurance company might store highly sensitive data such as customer health information and bank card numbers in an encrypted database.
[0082] If the sensitivity level information is not a highly sensitive level, determine whether the sensitivity level information is a moderately sensitive level.
[0083] In this embodiment, the content of the sensitive level information can be determined by performing content identification on the sensitive level information.
[0084] If the sensitivity level information is a medium sensitivity level, the target business data is stored in a pre-built general storage area in the data lake.
[0085] In this embodiment, the general storage area refers to a general database pre-built in the data lake. If the sensitivity level information is detected to be medium sensitivity, the target business data is stored in the general database.
[0086] If the sensitivity level information is not a medium sensitivity level, the target business data is stored in a low-cost storage medium pre-built in the data lake.
[0087] In this embodiment, the low-cost storage medium refers to low-cost object storage. If the sensitivity level information is not detected as medium, it indicates that the sensitivity level information is low, and the target business data is stored in the low-cost storage medium. For example, an insurance company stores general business statistics in object storage to reduce storage costs.
[0088] This application determines whether the sensitive level information is highly sensitive; if the sensitive level information is highly sensitive, the target business data is stored in the encrypted storage area pre-constructed in the data lake; and if the sensitive level information is not highly sensitive, determines whether the sensitive level information is moderately sensitive; if the sensitive level information is moderately sensitive, the target business data is stored in the ordinary storage area pre-constructed in the data lake; and if the sensitive level information is not moderately sensitive, the target business data is stored in the low-cost storage medium pre-constructed in the data lake. This application intelligently selects the appropriate storage method in the data lake to store the target business data based on the sensitive level information of the target business data, thereby effectively improving the storage intelligence and storage standardization of the target business data.
[0089] In some optional implementations, after step S206, the electronic device may further perform the following steps:
[0090] Call the preset data lineage tool.
[0091] In this embodiment, the data lineage tool is a pre-built automated tool with the function of data lineage tracing.
[0092] During the backup process of the target business data, the data flow path of the target business data is recorded and processed based on the data lineage tool to obtain corresponding flow records.
[0093] In this embodiment, after storing the target business data in the data lake, the data lake's backup and recovery functions can be further utilized to regularly back up the target business data. Data lineage tracking tools can be used to record the backup and recovery paths (i.e., data flow paths) of the target business data to obtain corresponding flow records, thereby ensuring the traceability of the target business data. Furthermore, the effectiveness of the recovery process can be regularly tested to prevent data loss.
[0094] During the data usage process of the target business data, usage data of the target business data is recorded based on the data lineage tool.
[0095] In this embodiment, during the data usage process of the above-mentioned target business data, the data access and usage can be recorded based on the use of the above-mentioned data lineage tool to obtain corresponding usage data.
[0096] Specifically, each type of business data should be pre-designated with a clear owner. The data owner is responsible for the accuracy, completeness, and security of that data. For example, the owner of customer data could be the customer service department, while the owner of insurance policy data could be the underwriting department. The responsibilities of data owners should also be clearly defined: data owners need to develop data management strategies, oversee data usage, promptly address data quality issues, and approve data changes. Furthermore, a communication mechanism for data owners should be established: regular data owner meetings should be held to share data management experiences and issues, and to jointly resolve data governance challenges.
[0097] A corresponding audit log is generated based on the usage data.
[0098] In this embodiment, a matching audit log may be generated by performing statistical analysis on the above usage data.
[0099] The flow record and the audit log are stored and processed.
[0100] In this embodiment, there is no specific limitation on the storage method of the above-mentioned circulation records and audit logs, which can be determined according to actual business needs. For example, any storage method such as a local database, a disk, a cloud server, or a blockchain storage can be used. Among them, by tracking data based on the above-mentioned circulation records, the traceability and transparency of business data can be ensured, as well as the compliance and transparency of business data. In addition, by analyzing the audit logs, abnormal behaviors can be discovered and handled in a timely manner. For example, through the data lineage tracking tool, the insurance company discovered that an employee had accessed customer health information multiple times on weekends. After investigation, it was found that the employee had violated the rules and was investigated and handled in a timely manner.
[0101] This application calls a preset data lineage tool; then, during the backup process of the target business data, records the data flow path of the target business data based on the data lineage tool to obtain corresponding flow records; and during the data use process of the target business data, records the usage data of the target business data based on the data lineage tool; then generates a corresponding audit log based on the usage data; and subsequently stores and processes the flow record and the audit log. This application records the data flow path of the target business data based on the use of the data lineage tool to obtain corresponding flow records during the backup process of the target business data; and during the data use process of the target business data, records the usage data of the target business data based on the use of the data lineage tool; then generates a corresponding audit log based on the usage data, and then stores and processes the flow record and the audit log, so that data query and tracing can be performed later, effectively ensuring the traceability, compliance and transparency of business data.
[0102] In some optional implementations of this embodiment, after step S204, the electronic device may further perform the following steps:
[0103] Determine whether a user-triggered access request for the target business data is received.
[0104] In this embodiment, the access request is a data access request for target business data stored in the data lake triggered by a user in the system.
[0105] If so, obtain data usage rules corresponding to the target business data.
[0106] In this embodiment, data usage rules are predefined based on actual data usage requirements, including who can access and use which data, as well as the purpose and method of use. For example, different data access permissions are assigned to business personnel to ensure that only authorized personnel can access sensitive customer information and that it can only be used for insurance business processing. For example, an insurance company may grant underwriters access to customer health and insurance information, but prohibit them from accessing customer financial information. Business personnel can only use customer information for underwriting purposes and not for other purposes.
[0107] Based on the data usage rule, it is determined whether the user has data access permission to the target business data.
[0108] In this embodiment, the identity type of the user may be obtained, and permission matching may be performed on the identity type based on the usage rules to determine whether the user has data access rights to the target business data.
[0109] If so, obtain the target business data from the data lake.
[0110] In this embodiment, if it is detected that the user has data access permissions for the target business data, the target business data will be automatically retrieved from the data lake. If it is detected that the user does not have data access permissions for the target business data, the response to the access request will be limited and a corresponding insufficient permission reminder will be issued.
[0111] The target business data is displayed and processed.
[0112] In this embodiment, the target business data may be displayed in a page display manner to complete the response processing for the access request.
[0113] This application determines whether a user-triggered access request for the target business data is received; if so, obtains the data usage rules corresponding to the target business data; then, based on the data usage rules, determines whether the user has data access rights to the target business data; if so, obtains the target business data from the data lake; and subsequently displays and processes the target business data. After receiving a user-triggered access request for the target business data, this application will automatically and intelligently determine whether the user has data access rights to the target business data based on the use of data usage rules, and only when it is detected that the user has data access rights to the target business data, will it obtain the target business data from the data lake and display and process it, effectively implementing standardized access control for business data, protecting business data from unauthorized access, use or leakage, and improving the access security of business data.
[0114] In some optional implementations of this embodiment, after step S204, the electronic device may further perform the following steps:
[0115] Obtain the data storage duration of the target business data.
[0116] In this embodiment, the data storage duration of the target business data can be obtained by obtaining the first storage time of the target business data in the data lake and then calculating the difference between the current time and the first storage time.
[0117] Determine whether the data storage period is greater than a preset retention period.
[0118] In this embodiment, there is no specific limitation on the numerical value of the above retention period, which can be determined according to actual data destruction requirements, for example, it can be set to 4 years.
[0119] If yes, get the preset destruction policy.
[0120] In this embodiment, the destruction policy includes: securely destroying data that has reached its retention period to prevent the disclosure of customer privacy information. Specifically, data cleaning and overwriting are performed to ensure that data cannot be recovered and prevent data leakage.
[0121] Perform corresponding data destruction processing on the target business data based on the destruction policy.
[0122] In this embodiment, the target business data may be destroyed according to the destruction policy. For example, for completed claims data that has reached its retention period, the system will automatically and securely destroy the data, thereby ensuring that customer privacy information is not leaked.
[0123] This application obtains the data storage duration of the target business data; then determines whether the data storage duration is greater than the preset retention period; if so, obtains a preset destruction policy; and subsequently performs corresponding data destruction processing on the target business data based on the destruction policy. This application obtains the data storage duration of the target business data, and when it detects that the data storage duration is greater than the preset retention period, it automatically and intelligently performs corresponding data destruction processing on the target business data based on the use of the destruction policy, thereby achieving the safe destruction of data that has reached the retention period, effectively ensuring that customer privacy information is not leaked, and improving the intelligence and standardization of data destruction processing.
[0124] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.
[0125] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0126] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0127] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned target business data, the above-mentioned target business data can also be stored in a node of a blockchain.
[0128] The blockchain referred to in this application refers to a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (to prevent counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0129] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0130] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0131] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0132] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0133] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of a data management device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0134] like Figure 3 As shown, the data management device 300 of this embodiment includes: a collection module 301, a pre-processing module 302, an identification module 303, a first storage module 304, a first acquisition module 305 and a sharing module 306.
[0135] The collection module 301 is used to collect business data related to business needs from a variety of preset data sources;
[0136] A preprocessing module 302 is configured to perform data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data;
[0137] An identification module 303 is configured to identify the sensitivity level of the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data;
[0138] A first storage module 304 is configured to store the target business data in a preset data lake based on a target storage method corresponding to the sensitivity level information;
[0139] A first acquisition module 305 is used to acquire a preset data sharing strategy;
[0140] The sharing module 306 is configured to perform data sharing processing on the target business data in the data lake based on the data sharing policy.
[0141] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0142] In some optional implementations of this embodiment, the preprocessing module 302 includes:
[0143] A first processing submodule is configured to perform data cleaning on the business data based on the data quality management tool to obtain cleansed first business data;
[0144] a second processing submodule, configured to perform data format verification on the first business data, so as to filter out second business data that passes the data format verification from the first business data;
[0145] a third processing submodule, configured to perform a mandatory item check on the second business data, so as to filter out third business data that passes the mandatory item check from the second business data;
[0146] a fourth processing submodule, configured to perform data standardization processing on the third service data to obtain processed fourth service data;
[0147] The determination submodule is configured to use the fourth service data as the target service data.
[0148] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0149] In some optional implementations of this embodiment, the data management device further includes:
[0150] The second acquisition module is used to acquire the pre-collected initial business data;
[0151] A first processing module is used to clean and label the initial business data with a sensitivity level to obtain corresponding processed data;
[0152] A construction module, configured to perform sample construction processing on the processed data based on a preset feature screening strategy to obtain corresponding sample data;
[0153] The first calling module is used to call the preset machine learning model;
[0154] A second processing module is used to perform model training and optimization processing on the machine learning model based on the sample data until a specified model that meets preset requirements is obtained;
[0155] A determination module is used to use the specified model as the sensitive recognition model.
[0156] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0157] In some optional implementations of this embodiment, the first storage module 304 includes:
[0158] A first judgment submodule is used to judge whether the sensitivity level information is a highly sensitive level;
[0159] A first storage submodule is configured to store the target business data in a pre-constructed encrypted storage area in the data lake if the sensitivity level information is highly sensitive;
[0160] A second judgment submodule is configured to judge whether the sensitivity level information is of a moderate sensitivity level if the sensitivity level information is not of a highly sensitive level;
[0161] A second storage submodule is configured to store the target business data in a pre-built common storage area in the data lake if the sensitivity level information is a moderate sensitivity level;
[0162] The third storage submodule is configured to store the target business data in a low-cost storage medium pre-built in the data lake if the sensitivity level information is not a medium sensitivity level.
[0163] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0164] In some optional implementations of this embodiment, the data management device further includes:
[0165] The second calling module is used to call the preset data lineage tool;
[0166] A first recording module is configured to record the data flow path of the target business data based on the data lineage tool during the backup process of the target business data to obtain a corresponding flow record;
[0167] A second recording module is configured to record usage data of the target business data based on the data lineage tool during the data usage of the target business data;
[0168] A generation module, configured to generate a corresponding audit log based on the usage data;
[0169] The second storage module is used to store and process the flow record and the audit log.
[0170] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0171] In some optional implementations of this embodiment, the data management device further includes:
[0172] A determination module, configured to determine whether a user-triggered access request for the target business data has been received;
[0173] a third acquisition module, configured to, if yes, acquire data usage rules corresponding to the target business data;
[0174] a first judgment module, configured to judge, based on the data usage rule, whether the user has data access authority to the target business data;
[0175] A fourth acquisition module, configured to, if yes, acquire the target business data from the data lake;
[0176] The display module is used to display the target business data.
[0177] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0178] In some optional implementations of this embodiment, the data management device further includes:
[0179] A fifth acquisition module is used to obtain the data storage time of the target business data;
[0180] A second judgment module is used to judge whether the data storage time is greater than a preset retention period;
[0181] A sixth acquisition module, configured to, if yes, acquire a preset destruction strategy;
[0182] The third processing module is used to perform corresponding data destruction processing on the target business data based on the destruction policy.
[0183] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the data management method in the aforementioned embodiment, and are not described in detail here.
[0184] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0185] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 having components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0186] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0187] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the data management method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0188] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the data management method.
[0189] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0190] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0191] In an embodiment of the present application, the present application collects business data related to business needs from multiple data sources, and then pre-processes the business data based on the use of data quality management tools to obtain target business data. Thereafter, the target business data is identified for sensitivity level based on the use of a sensitive identification model to obtain sensitivity level information of the target business data, and the target business data is stored in a data lake based on a target storage method corresponding to the sensitivity level information. Finally, the target business data in the data lake is processed for data sharing based on the use of a data sharing strategy. In this way, the present application automatically and intelligently realizes the comprehensive management and effective utilization of business data based on the combined use of data quality management tools, sensitive identification models, data lakes and data sharing strategies, provides a more efficient, secure and reliable solution for the system's data management, effectively improves the intelligence, reliability and standardization of data management, and thus helps to improve the stability of the system.
[0192] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the data management method as described above.
[0193] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0194] In an embodiment of the present application, the present application collects business data related to business needs from multiple data sources, and then pre-processes the business data based on the use of data quality management tools to obtain target business data. Thereafter, the target business data is identified for sensitivity level based on the use of a sensitive identification model to obtain sensitivity level information of the target business data, and the target business data is stored in a data lake based on a target storage method corresponding to the sensitivity level information. Finally, the target business data in the data lake is processed for data sharing based on the use of a data sharing strategy. In this way, the present application automatically and intelligently realizes the comprehensive management and effective utilization of business data based on the combined use of data quality management tools, sensitive identification models, data lakes and data sharing strategies, provides a more efficient, secure and reliable solution for the system's data management, effectively improves the intelligence, reliability and standardization of data management, and thus helps to improve the stability of the system.
[0195] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0196] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A data management method, characterized in that: The steps include: Collect business data related to business needs from multiple preset data sources; Performing data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data; Performing sensitivity level identification on the target business data based on a preset sensitivity identification model to obtain sensitivity level information of the target business data; Based on the target storage method corresponding to the sensitivity level information, the target business data is stored in a preset data lake; Get the preset data sharing policy; Data sharing processing is performed on the target business data in the data lake based on the data sharing strategy.
2. The data management method according to claim 1, wherein: The step of performing data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data specifically includes: Performing data cleansing on the business data based on the data quality management tool to obtain cleansed first business data; Performing data format verification on the first business data to filter out second business data that passes the data format verification from the first business data; Performing a mandatory item check on the second business data to filter out third business data that passes the mandatory item check from the second business data; performing data standardization processing on the third service data to obtain processed fourth service data; The fourth business data is used as the target business data.
3. The data management method according to claim 1, wherein: Before the step of identifying the sensitivity level of the target business data based on the preset sensitivity identification model to obtain the sensitivity level information of the target business data, the method further includes: Obtaining pre-collected initial business data; Cleaning and sensitivity-level labeling of the initial business data to obtain corresponding processed data; Performing sample construction processing on the processed data based on a preset feature screening strategy to obtain corresponding sample data; Call the preset machine learning model; Performing model training and optimization processing on the machine learning model based on the sample data until a specified model that meets preset requirements is obtained; The designated model is used as the sensitive recognition model.
4. The data management method according to claim 1, wherein: The step of storing the target business data in a preset data lake based on the target storage method corresponding to the sensitivity level information specifically includes: Determining whether the sensitivity level information is highly sensitive; If the sensitivity level information is highly sensitive, the target business data is stored in a pre-built encrypted storage area in the data lake; If the sensitivity level information is not a highly sensitive level, determining whether the sensitivity level information is a moderately sensitive level; If the sensitivity level information is of medium sensitivity, the target business data is stored in a pre-built general storage area in the data lake; If the sensitivity level information is not a medium sensitivity level, the target business data is stored in a low-cost storage medium pre-built in the data lake.
5. The data management method according to claim 1, wherein: After the step of performing data sharing processing on the target business data in the data lake based on the data sharing policy, the method further includes: Call the preset data lineage tool; During the backup process of the target business data, recording the data flow path of the target business data based on the data lineage tool to obtain corresponding flow records; and During the data usage of the target business data, recording usage data of the target business data based on the data lineage tool; generating a corresponding audit log based on the usage data; The flow record and the audit log are stored and processed.
6. The data management method according to claim 1, wherein: After the step of storing the target business data in a preset data lake based on the target storage method corresponding to the sensitivity level information, the method further includes: Determining whether a user-triggered access request for the target business data is received; If so, obtaining data usage rules corresponding to the target business data; Based on the data usage rules, determining whether the user has data access rights to the target business data; If yes, obtain the target business data from the data lake; The target business data is displayed and processed.
7. The data management method according to claim 1, wherein: After the step of storing the target business data in a preset data lake based on the target storage method corresponding to the sensitivity level information, the method further includes: Obtaining the data storage duration of the target business data; Determining whether the data storage period is greater than a preset retention period; If so, obtain the preset destruction strategy; Perform corresponding data destruction processing on the target business data based on the destruction policy.
8. A data management device, characterized in that: include: The collection module is used to collect business data related to business needs from multiple preset data sources; A preprocessing module, configured to perform data preprocessing on the business data based on a preset data quality management tool to obtain corresponding target business data; An identification module is used to identify the sensitivity level of the target business data based on a preset sensitivity identification model to obtain the sensitivity level information of the target business data; A first storage module is configured to store the target business data in a preset data lake based on a target storage method corresponding to the sensitivity level information; A first acquisition module is used to acquire a preset data sharing strategy; A sharing module is used to perform data sharing processing on the target business data in the data lake based on the data sharing strategy.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the data management method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data management method according to any one of claims 1 to 7.