Insurance industry customer information protection method and system

By employing natural language processing and machine learning algorithms in the insurance industry to automatically identify and label customer information, formulate diverse de-identification strategies, and generate and update these strategies in real time, the problem of traditional static de-identification methods being unable to adapt to changes in business needs has been solved, achieving precise protection and enhanced security of customer information.

CN121786077APending Publication Date: 2026-04-03CHINA LIFE INSURANCE CO LTD HEBEI BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional static data masking methods cannot flexibly adjust masking strategies according to real-time business needs, leading to either over-masking affecting business efficiency or under-masking causing information leakage. They are also difficult to effectively handle complex and interconnected data in the insurance industry and ensure data consistency and integrity.

Method used

Natural language processing and machine learning algorithms are used to automatically identify and label customer information, formulate diverse de-identification strategies, dynamically generate and update de-identification strategies by monitoring changes in business needs in real time, and apply the optimal strategy in real time during data transmission. De-identification strategies are generated by combining rule engines and machine learning algorithms, and de-identification processing is performed using replacement, encryption, truncation and masking algorithms.

Benefits of technology

It enables precise and dynamic protection of customer information in the insurance industry, reduces the risk of data leakage, adapts to complex business scenarios, ensures a balance between information security and business efficiency, and optimizes de-identification strategies through feedback mechanisms to improve the overall level of information security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786077A_ABST
    Figure CN121786077A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data security and privacy protection, and particularly provides an insurance industry customer information protection method, which comprises the following steps: S1, automatically identifying customer information, positioning sensitive data in the customer information, and marking the sensitive data; s2, making diversified desensitization strategies, and storing the desensitization strategies in a strategy rule base; s3, monitoring the change of service requirements, formulating a re-evaluation and generation mechanism of a desensitization strategy, and generating and updating the desensitization strategy in real time; s4, when the customer information is transmitted to the business system from the database, performing real-time desensitization processing on the customer information; and S5, collecting the use feedback of the service system on the desensitization data, and feeding back the feedback information to the step S3. According to the method, client information is strictly protected in each link through dynamic and accurate desensitization processing, so that the security of the client information in the insurance industry is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data security and privacy protection technology, and more specifically, it relates to a method and system for protecting customer information in the insurance industry. Background Technology

[0002] In the insurance industry, customer information is a core asset with extremely high sensitivity and value, including but not limited to personal identification information, financial status, and health information. With the accelerated digital transformation of the insurance industry, data faces numerous security challenges during storage, transmission, and use, such as data breaches and unauthorized access.

[0003] Traditional static data masking methods have significant limitations and are ill-suited to the complex and ever-changing business scenarios in the insurance industry. Static data masking is typically performed once while data is offline, making it impossible to flexibly adjust masking strategies according to real-time business needs. This can lead to over-masking, impacting business efficiency, or under-masking, resulting in information leaks. Furthermore, the insurance industry deals with massive amounts of highly interconnected data, making it difficult for static masking methods to effectively handle complex and related data and ensure data consistency and integrity. Summary of the Invention

[0004] Based on the above-mentioned technical problems, this application provides a method and system for protecting customer information in the insurance industry, in order to solve the technical problems existing in the prior art where traditional static de-identification methods cannot flexibly adjust de-identification strategies according to real-time business needs, which may lead to excessive de-identification affecting business efficiency, or insufficient de-identification causing information leakage.

[0005] To achieve the above objectives, the technical solution adopted in this application is: to provide a method for protecting customer information in the insurance industry, comprising the following steps: S1. Using natural language processing and machine learning algorithms, automatically identify customer information in the insurance business database, locate sensitive data in the customer information, and mark it. S2. Based on the sensitive data marked in step S1, formulate diverse desensitization strategies according to different business scenarios, user roles and customer information usage purposes, and store the desensitization strategies in the strategy rule base. S3. Based on the strategy rule base built in step S2, continuously monitor changes in business requirements. When a new business scenario, user role adjustment, or data usage purpose change is detected, immediately trigger the re-evaluation and generation mechanism of the de-identification strategy, and generate and update the de-identification strategy in real time. S4. When the business system initiates a data access request, and the request explicitly includes customer information fields that need to be de-identified, the latest de-identification strategy will be automatically applied for real-time processing when the customer information is transmitted from the database to the business system, based on the de-identification strategy generated in step S3. S5. Collect feedback from the business system on the use of de-identified data, and send the feedback information back to step S3 so that the de-identification strategy can be optimized and adjusted through step S3.

[0006] Furthermore, step S1 includes: S11. Clean, remove noise, and convert the format of customer information in the insurance business database to obtain customer data; S12. Use a convolutional neural network model to automatically identify the customer data in step S11; S13. Mark the sensitive data identified in step S12.

[0007] Furthermore, step S2 includes: S21. Analyze the data anonymization requirements in various business scenarios of the insurance industry and define anonymization rules for different scenarios; S22. Store the defined de-identification rules in the policy rule base to form the initial de-identification policy set; Furthermore, step S3 includes: S31. Real-time collection of contextual information during the processing of customer information requests in the business system, including identity identifier, permission level, operation time, geographical location, device information, and data sensitivity level; S32. Based on the potential impact of data breaches on insurance companies and customers, data is classified into different sensitivity levels such as public data, internal data, and confidential data, while clearly defining the roles of different visitors. S33. Utilizing a rule engine and machine learning algorithms, the collected context information is matched with the policy library. Through priority ranking and conflict resolution mechanisms, the optimal de-identification strategy for the current access scenario is quickly generated. The de-identification strategy generation algorithm is expressed as follows: Where S is the generated desensitization strategy, C is the set of collected context information, P is the set of strategy rules in the strategy library, and f is the strategy matching and decision function; S34. Based on the customer sensitive data marked in step S1, the data is parsed and identified at the field level using natural language processing and regular expression technology, and the sensitive fields are desensitized according to the generated desensitization strategy. S35. By monitoring changes in business needs, data security incidents, and regulatory information, adjust the de-identification strategy in a timely manner to ensure the effectiveness and adaptability of the strategy rule base.

[0008] Furthermore, step S4 includes: S41. The desensitization strategy includes replacement algorithm, encryption algorithm, truncation algorithm and masking algorithm; S42. Select at least one of the above desensitization algorithms to desensitize the marked sensitive data.

[0009] Furthermore, step S5 includes: S51. Set up a de-identification effect verification module in the business system. This de-identification effect verification module is responsible for collecting and analyzing the usage feedback of the de-identification data. S52. Set a feedback threshold. When the desensitization effect reaches the preset threshold, the feedback mechanism will be automatically triggered. S53. The information that triggers the feedback mechanism is transmitted to step S3, the desensitization strategy is optimized and adjusted accordingly, and updated to the strategy rule base.

[0010] This invention provides a customer information protection system for the insurance industry, comprising: a data identification and marking module, a dynamic desensitization strategy formulation module, a real-time desensitization execution module, and a desensitization effect verification and feedback module; The data identification and tagging module is used to automatically identify customer information in the insurance business database, locate and tag sensitive customer data; The dynamic desensitization strategy formulation module formulates diverse desensitization strategies based on customer data transmitted by the data identification and tagging module, and outputs the predetermined desensitization strategies to the real-time desensitization execution module. The real-time de-identification execution module dynamically selects a suitable de-identification algorithm based on the needs of the business system, and outputs the de-identified customer data to the business system.

[0011] The de-identification effect verification and feedback module receives feedback from business systems on the use of de-identified data, optimizes and adjusts the de-identification strategy based on the feedback information, and updates it to the strategy rule base.

[0012] Furthermore, the data recognition and labeling module includes a submodule integrating a natural language processing submodule and a machine learning algorithm. The Natural Language Processing submodule uses the NLTK library for text processing and can identify entity information in text. The machine learning algorithm submodule is implemented based on the Scikit-learn library and improves the accuracy of data recognition through a convolutional neural network model.

[0013] Furthermore, the real-time de-identification execution module includes: Replacement Algorithm Submodule: Replaces specific parts of sensitive data with preset placeholders or random characters; Encryption algorithm submodule: Uses encryption keys to encrypt sensitive data; The truncation algorithm submodule truncates sensitive data to a specified length, retaining some information while hiding the sensitive parts; The masking algorithm submodule performs masking processing on specific parts of sensitive data, such as using specific symbols to cover sensitive information. Furthermore, the desensitization effect verification and feedback module includes a verification module, which receives feedback information from the business system and sets a feedback threshold. When the feedback information from the business system exceeds the threshold, the feedback content is sent to the dynamic desensitization strategy formulation module.

[0014] Compared with existing technologies, the beneficial effects of the insurance industry customer information protection method and system provided in this application are: 1. Through dynamic and precise data anonymization, customer information is rigorously protected at every stage. Whether it's static storage in the database, transmission to business systems over the network, or usage within those systems, the risk of customer information leakage is effectively reduced. Even if data is unfortunately illegally obtained, because critical and sensitive information has been processed through dynamic and multiple anonymization algorithms, attackers will find it difficult to obtain valuable information, thus greatly enhancing the security of customer information in the insurance industry. 2. By deeply applying dynamic data masking technology to complex business scenarios in the insurance industry, it breaks through the limitations of traditional static data masking. It can perform real-time, flexible, and accurate data masking based on the specific needs of different business scenarios, ensuring effective protection of customer information under various complex circumstances. 3. Utilize a policy rule base to achieve dynamic policy formulation. By collecting contextual information from data access scenarios in real time, the rule engine and machine learning algorithms match the collected contextual information with the policy base to quickly generate the optimal de-identification policy for the current access scenario. Simultaneously, establish a real-time adjustment mechanism for de-identification policies. By monitoring changes in business requirements, data security incidents, and regulatory information, the de-identification policies are adjusted promptly to ensure they meet new requirements, maintain a balance between information security and business operations, and improve the overall level of information security protection.

[0015] 4. By setting up a desensitization effect verification module in the business system, when the desensitization effect reaches the preset threshold, a feedback mechanism is automatically triggered, and the feedback information is sent to the dynamic desensitization strategy formulation module so that the desensitization strategy can be optimized and adjusted in a targeted manner and updated to the strategy rule base, ensuring that the desensitization strategy is always in the optimal state and continuously improving the customer information protection effect. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a customer information protection method for the insurance industry according to the present invention. Detailed Implementation

[0018] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.

[0019] It should be noted that when a component is referred to as being "fixed to" or "set on" another component, it can be directly on or indirectly on that other component. When a component is referred to as being "connected to" another component, it can be directly connected to or indirectly connected to that other component.

[0020] It should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" or "several" means two or more, unless otherwise explicitly specified.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0023] Please refer to the following: Figure 1 As shown below, an insurance industry customer information protection method and system provided by embodiments of this application will be described. This invention provides an insurance industry customer information protection method to better protect the security of customer information in insurance business. Through a precise and dynamic process, it ensures the privacy of customer information during storage, transmission, and use, and includes the following steps: Step S1: Using natural language processing and machine learning algorithms, automatically identify customer information in the insurance business database, locate sensitive data in the customer information, and mark it. The sensitive data in the customer information includes customer name, customer age, ID number, bank card number, etc.

[0024] In this embodiment, step S1 includes: S11. Clean, denoise, and convert the customer information in the insurance business database to obtain customer data. Data cleaning involves removing duplicate records and correcting erroneous data. Denoising is achieved by setting a data range threshold to filter data that exceeds a reasonable range. Format conversion involves uniformly converting the data into JSON format for subsequent processing.

[0025] S12. A Convolutional Neural Network (CNN) model is used to automatically identify the customer data in step S11. Preferably, the CNN model uses the LeNet-5 architecture because it has a good feature extraction capability for structured data and can accurately identify sensitive data features in customer information.

[0026] S13. Mark the identified sensitive data by adding a specific prefix, such as "SENSITIVE_", before the identified sensitive data field to facilitate quick location and processing later.

[0027] Step S2: Based on different business scenarios, user roles, and the purpose of using customer information, formulate diverse de-identification strategies and store the de-identification strategies in the strategy rule base.

[0028] In this embodiment, step S2 includes: S21. Analyze the data anonymization requirements in various business scenarios of the insurance industry. For example, in the claims review scenario, it is necessary to ensure that claims personnel can obtain sufficient information for review, while protecting sensitive customer information from excessive disclosure. Or in the market research scenario, researchers only need to understand the general characteristics of customers, such as age range and geographical distribution, without needing to obtain specific sensitive information. Based on these requirements, define anonymization rules for different scenarios.

[0029] S22. Store the defined de-identification rules in the policy rule base to form an initial set of de-identification policies. Preferably, a relational database is used for storage to facilitate efficient querying and management of de-identification rules.

[0030] Step S3: Based on the policy rule base built in Step S2, continuously monitor changes in business requirements. When a new business scenario, user role adjustment, or data usage purpose is detected, immediately trigger the re-evaluation and generation mechanism of the de-identification policy, generating and updating the de-identification policy in real time. In specific implementation, JavaScript code can be embedded in the front end of the business system to collect user operation time, geographical location (via IP resolution or GPS positioning), and device information (such as browser type and operating system version). The back end records user identity and permission level through logs.

[0031] In this embodiment, step S3 includes: S31. Real-time collection of contextual information during the processing of customer information requests in the business system, including identity identifiers (such as user ID, department, position), permission levels (ordinary users, administrators, etc.), operation time, geographical location, device information, and data sensitivity level. This information can reflect the specific situation of the current access and provide a basis for generating de-identification strategies.

[0032] S32. Based on the potential impact of a data breach on insurance companies and customers, data is categorized into different sensitivity levels, such as public data, internal data, and confidential data. Different visitor roles are also clearly defined, such as ordinary employees, administrators, risk control personnel, and external partners. The scope of data that can be accessed and used, as well as the anonymization requirements, differ for data of different sensitivity levels and for different visitor roles.

[0033] S33. Utilizing a rule engine and machine learning algorithms, the collected context information is matched against a policy library. Specifically, the rule engine performs initial matching based on preset rules, while the machine learning algorithm optimizes the matching results through learning and analysis of historical data. Through priority ranking and conflict resolution mechanisms, the optimal de-identification strategy for the current access scenario is quickly generated. The de-identification strategy generation algorithm is expressed as follows: Where S represents the generated de-identification strategy, C represents the set of collected context information, P represents the set of strategy rules in the strategy library, and f represents the strategy matching and decision function. For example, when the visitor is an external partner and the visit occurs on a non-working day, a strategy is generated based on the algorithm to fully mask and de-identify fields such as the customer's ID number and mobile phone number.

[0034] S34. Based on the customer sensitive data marked in step S1, the data is parsed and identified at the field level using Natural Language Processing (NLP) and Regular Expression (RE) technologies. NLP technology can understand the semantic information of text, while REFES can accurately match data in specific formats. Sensitive fields are then de-identified according to the generated de-identification strategy, such as replacement, encryption, truncation, and masking.

[0035] S35. By monitoring changes in business needs, data security incidents, and regulatory information, timely adjustments to the de-identification strategy are made. For example, when the insurance business launches new products or services, it may be necessary to adjust the de-identification strategy in the relevant business scenarios. When a data security incident occurs, it is necessary to analyze the cause of the incident and strengthen the corresponding de-identification measures. When regulations and policies change, it is necessary to ensure that the de-identification strategy complies with the new regulatory requirements, so as to ensure the effectiveness and adaptability of the strategy rule base and ensure that customer information protection is always in the best state.

[0036] Step S4: When the business system initiates a data access request, and the request explicitly includes customer information fields that need to be de-identified, the latest de-identification strategy will be automatically applied for real-time processing when the customer information is transmitted from the database to the business system, based on the de-identification strategy generated in Step S3. The trigger condition for de-identification processing is that the business system initiates a data access request and the request includes customer information fields that need to be de-identified.

[0037] In this embodiment, step S4 includes: S41. Multiple desensitization strategies include replacement algorithms, encryption algorithms, truncation algorithms, and masking algorithms. Among them, the replacement algorithm uses random strings to replace parts of sensitive data, the encryption algorithm uses the AES algorithm which has high security and encryption efficiency, the truncation algorithm sets different truncation lengths according to data types, such as truncating the ID number into the first 4 digits and the last 4 digits, and the masking algorithm uses special symbols to cover the middle part of sensitive data.

[0038] S42. Select at least one of the above-mentioned desensitization algorithms to desensitize the marked sensitive data, wherein the selection is based on the combination of desensitization algorithms applicable to different business scenarios, user roles and data usage purposes as specified in the desensitization strategy.

[0039] Step S5 collects feedback from business systems regarding the use of anonymized data and relays this feedback to step S3 for optimization and adjustment of the anonymization strategy. Specifically, the optimization process includes analyzing indicators such as data accuracy and business process smoothness in the feedback information. Based on the degree to which these indicators deviate from preset thresholds, targeted strategy optimization plans are formulated. For example, if the system collects feedback from business systems that anonymized data is impacting business efficiency, it first analyzes the specific influencing factors, such as incomplete data display hindering the review process. Then, it adjusts the corresponding anonymization algorithm parameters or increases the length of displayed fields. Finally, the optimized anonymization strategy is updated to the strategy rule base.

[0040] In this embodiment, step S5 includes: S51. Set up a desensitization effect verification module in the business system. This module is responsible for collecting and analyzing user feedback on desensitized data. The collection method is to set up a feedback button on the business system operation interface, where users can actively submit their user experience. At the same time, the system automatically records error information, operation time and other indicators during the data usage process as feedback data. At this time, step S3 collects feedback from the business system, adjusts the desensitization strategy, and updates it to the strategy rule base.

[0041] S52. Set a feedback threshold. When the desensitization effect reaches the preset threshold, the feedback mechanism will be automatically triggered. The feedback threshold is set to data accuracy of no less than 95% (determined based on historical data analysis and business needs) and the number of user operation interruptions caused by the smoothness of the business process shall not exceed twice per hour (evaluated by user feedback and system logs).

[0042] S53. The feedback information is transmitted to step S3 so that the desensitization strategy can be optimized and adjusted in a targeted manner and updated in the strategy rule base. The specific transmission method can be a message queue to ensure the timely transmission and processing of feedback information.

[0043] This invention provides a customer information protection system for the insurance industry, comprising: a data identification and marking module, a dynamic desensitization strategy formulation module, a real-time desensitization execution module, and a desensitization effect verification and feedback module; The data identification and tagging module is used to automatically identify customer information in the insurance business database and locate and tag sensitive customer data.

[0044] In this embodiment, the data recognition and labeling module includes a sub-module that integrates a natural language processing sub-module and a machine learning algorithm. The natural language processing sub-module uses the NLTK library for text processing, which can identify entity information in the text, such as names of people, places, and organizations. At the same time, it can analyze the semantic structure of the text and accurately determine whether the text contains sensitive information. The machine learning algorithm submodule is implemented based on the Scikit-learn library and improves the accuracy of data recognition through a convolutional neural network model. This convolutional neural network model, trained on a large amount of labeled data, can automatically learn the feature patterns of customer data, thereby accurately classifying and identifying new customer data.

[0045] Specifically, the data recognition and labeling module uses the NLTK library for word segmentation and part-of-speech tagging, and combines it with a convolutional neural network model trained by Scikit-learn to perform high-precision recognition of sensitive entities (such as names and ID numbers) in customer information.

[0046] The dynamic data masking strategy formulation module develops diverse data masking strategies based on customer data transmitted by the data identification and tagging module, and outputs these strategies to the real-time data masking execution module. Specifically, it first analyzes the data masking requirements across various business scenarios within the insurance industry, defining detailed masking rules based on different visitor roles and data sensitivity levels. These rules are then stored in a strategy rule base, forming an initial set of masking strategies. Upon receiving customer data from the data identification and tagging module, the module uses a rule engine and machine learning algorithms to generate the optimal masking strategy based on real-time collected context information, and outputs the strategy to the real-time data masking execution module.

[0047] The real-time de-identification execution module dynamically selects a suitable de-identification algorithm based on the needs of the business system, and outputs the de-identified customer data to the business system.

[0048] In this embodiment, the real-time desensitization execution module is used to formulate the desensitization strategy output by the module according to the dynamic desensitization strategy, and select appropriate sub-modules to perform real-time desensitization processing on customer data.

[0049] Specifically, the real-time desensitization execution module includes: a replacement algorithm submodule, an encryption algorithm submodule, a truncation algorithm submodule, and a masking algorithm submodule.

[0050] Replacement Algorithm Submodule: Replaces specific parts of sensitive data with preset placeholders or random characters; Encryption algorithm submodule: Uses encryption keys to encrypt sensitive data, ensuring data security during transmission and storage; The truncation algorithm submodule truncates sensitive data to a specified length, retaining some information while hiding the sensitive parts; Masking algorithm submodule: performs masking processing on specific parts of sensitive data, such as using specific symbols to cover sensitive information.

[0051] The de-identification effect verification and feedback module receives feedback from business systems on the use of de-identified data, optimizes and adjusts the de-identification strategy based on the feedback information, and updates it to the strategy rule base.

[0052] In this embodiment, the desensitization effect verification and feedback module includes a verification module. This module receives feedback information from the business system, sets a feedback threshold, and sends feedback exceeding the threshold to the dynamic desensitization strategy formulation module. For example, if the business system reports that the desensitized data causes certain business functions to malfunction, the verification module feeds this information back to the dynamic desensitization strategy formulation module so that the desensitization strategy can be adjusted and optimized.

[0053] It is understood that the parts in the above embodiments can be freely combined or deleted to form different combined embodiments. The specific contents of each combined embodiment will not be repeated here. After this description, it can be considered that the present invention specification has recorded each combined embodiment and can support different combined embodiments.

[0054] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for protecting customer information in the insurance industry, characterized in that, Includes the following steps: S1. Using natural language processing and machine learning algorithms, automatically identify customer information in the insurance business database, locate sensitive data in the customer information, and mark it. S2. Based on the sensitive data marked in step S1, formulate diverse desensitization strategies according to different business scenarios, user roles and customer information usage purposes, and store the desensitization strategies in the strategy rule base. S3. Based on the strategy rule base built in step S2, continuously monitor changes in business requirements. When a new business scenario, user role adjustment, or data usage purpose change is detected, immediately trigger the re-evaluation and generation mechanism of the de-identification strategy, and generate and update the de-identification strategy in the strategy rule base in real time. S4. When the business system initiates a data access request, and the request explicitly includes customer information fields that need to be de-identified, the latest de-identification strategy is automatically applied for real-time processing when the customer information is transmitted from the database to the business system, based on the de-identification strategy generated in step S3. S5. Collect feedback from the business system on the use of de-identified data, and send the feedback information back to step S3 so that the de-identification strategy can be optimized and adjusted through step S3.

2. The method for protecting customer information in the insurance industry according to claim 1, characterized in that, Step S1 includes: S11. Clean, remove noise, and convert the format of customer information in the insurance business database to obtain customer data; S12. Use a convolutional neural network model to automatically identify the customer data in step S11; S13. Mark the sensitive data identified in step S12.

3. The method for protecting customer information in the insurance industry according to claim 1, characterized in that, Step S2 includes: S21. Analyze the data anonymization requirements in various business scenarios of the insurance industry and define anonymization rules for different scenarios; S22. Store the defined de-identification rules in the policy rule base to form the initial de-identification policy set.

4. The method for protecting customer information in the insurance industry according to claim 1, characterized in that, Step S3 includes: S31. Real-time collection of contextual information during the processing of customer information requests in the business system, including identity identifier, permission level, operation time, geographical location, device information, and data sensitivity level; S32. Based on the potential impact of data breaches on insurance companies and customers, data is classified into different sensitivity levels such as public data, internal data, and confidential data, while clearly defining the roles of different visitors. S33. Utilizing a rule engine and machine learning algorithms, the collected context information is matched with the policy library. Through priority ranking and conflict resolution mechanisms, the optimal de-identification strategy for the current access scenario is quickly generated. The de-identification strategy generation algorithm is expressed as follows: Where S is the generated desensitization strategy, C is the set of collected context information, P is the set of strategy rules in the strategy library, and f is the strategy matching and decision function; S34. Based on the customer sensitive data marked in step S1, the data is parsed and identified at the field level using natural language processing and regular expression technology, and the sensitive fields are desensitized according to the generated desensitization strategy. S35. By monitoring changes in business needs, data security incidents, and regulatory information, adjust the de-identification strategy in a timely manner to ensure the effectiveness and adaptability of the strategy rule base.

5. The method for protecting customer information in the insurance industry according to claim 1, characterized in that, Step S4 includes: S41. The desensitization strategy includes replacement algorithm, encryption algorithm, truncation algorithm and masking algorithm; S42. Select at least one of the above desensitization algorithms to desensitize the marked sensitive data.

6. The method for protecting customer information in the insurance industry according to claim 1, characterized in that, Step S5 includes: S51. Set up a de-identification effect verification module in the business system. This de-identification effect verification module is responsible for collecting and analyzing the usage feedback of the de-identification data. S52. Set a feedback threshold. When the desensitization effect reaches the preset threshold, the feedback mechanism will be automatically triggered. S53. The information that triggers the feedback mechanism is transmitted to step S3, the desensitization strategy is optimized and adjusted accordingly, and updated to the strategy rule base.

7. A customer information protection system for the insurance industry, characterized in that, include: The module includes data identification and labeling, dynamic desensitization strategy formulation, real-time desensitization execution, and desensitization effect verification and feedback. The data identification and tagging module is used to automatically identify customer information in the insurance business database, locate and tag sensitive customer data; The dynamic desensitization strategy formulation module formulates diverse desensitization strategies based on customer data transmitted by the data identification and tagging module, and outputs the predetermined desensitization strategies to the real-time desensitization execution module. The real-time de-identification execution module dynamically selects a suitable de-identification algorithm based on the needs of the business system, and outputs the de-identified customer data to the business system. The de-identification effect verification and feedback module receives feedback from business systems on the use of de-identified data, optimizes and adjusts the de-identification strategy based on the feedback information, and updates it to the strategy rule base.

8. The insurance industry customer information protection system according to claim 7, characterized in that, The data recognition and labeling module includes a submodule that integrates natural language processing and machine learning algorithms. The Natural Language Processing submodule uses the NLTK library for text processing and can identify entity information in text. The machine learning algorithm submodule is implemented based on the Scikit-learn library and improves the accuracy of data recognition through a convolutional neural network model.

9. The insurance industry customer information protection system according to claim 7, characterized in that, The real-time de-identification execution module includes: Replacement Algorithm Submodule: Replaces specific parts of sensitive data with preset placeholders or random characters; Encryption algorithm submodule: Uses encryption keys to encrypt sensitive data; The truncation algorithm submodule truncates sensitive data to a specified length, retaining some information while hiding the sensitive parts; Masking algorithm submodule: performs masking processing on specific parts of sensitive data, such as using specific symbols to cover sensitive information.

10. The insurance industry customer information protection system according to claim 7, characterized in that, The desensitization effect verification and feedback module includes a verification module, which receives feedback information from the business system and sets a feedback threshold. When the feedback information from the business system exceeds the threshold, the feedback content is sent to the dynamic desensitization strategy formulation module.