Intelligent underwriting method and device, equipment and medium
By obtaining the policy identification field information, constructing the original table width and performing encoding embedding processing, and combining it with the policy verification of the gradient boosting model, the problems of low efficiency and poor accuracy in the underwriting process are solved, and efficient and accurate intelligent underwriting is achieved.
Patent Information
- Application Number
- CN202510827862.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
The existing underwriting process in the financial insurance industry is characterized by low efficiency, poor accuracy, and heavy reliance on manual intervention. This is especially true when dealing with non-standard policies and complex cases, making it difficult to adapt to online and large-scale business needs. Existing models also perform poorly when processing multi-dimensional, highly sparse data and unstructured fields.
By obtaining the identification field information in the insurance policy, extracting the feature fields and constructing the original table width, encoding and embedding processing are performed, and the insurance policy verification model of the gradient boosting model is used for analysis to improve data structuring and model generalization capabilities.
It improves underwriting efficiency and accuracy, reduces manual review costs, enhances the ability to prevent complex fraud, optimizes review efficiency and reduces the need for manual intervention.
Smart Images

Figure CN120707310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an intelligent underwriting method, device, equipment and medium. Background Art
[0002] In the financial and insurance industries, underwriting is a crucial component of risk management. Particularly in the life insurance sector, manual underwriting processes are often complex, inefficient, and rely heavily on experience, making them difficult to adapt to the growing trend of online, scaled-up business. While some insurance companies have introduced automated rule-based judgment systems, non-standard policies or complex cases still require significant manual intervention, which is time-consuming and subject to inconsistent judgment criteria, leading to risk loopholes and customer complaints.
[0003] To address this issue, the development of FinTech has driven the emergence of intelligent underwriting methods based on machine learning. These methods intelligentize the underwriting process by constructing feature dimensions, training models, and automatically distinguishing between standard and non-standard items. However, existing models generally suffer from crude feature selection, low semantic information utilization, and difficulty generalizing. They are unable to effectively handle multi-dimensional, highly sparse data and unstructured fields. Furthermore, the lack of unified data processing and encoding standards hinders model deployment and interpretability. Therefore, it is necessary to propose an intelligent underwriting method that can improve underwriting efficiency and accuracy, reduce manual review costs, and ensure the compliance and robust operation of financial services. Summary of the Invention
[0004] The present invention provides an intelligent underwriting method, device, equipment and medium to solve the technical problems in related technologies such as low underwriting efficiency and accuracy and high manual review costs.
[0005] In a first aspect, a smart underwriting method is provided, the method comprising:
[0006] Obtain identification field information from the target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number;
[0007] Extracting the characteristic field corresponding to the identification field information, and constructing a corresponding original table width based on the characteristic field information;
[0008] Performing encoding and embedding processing on the original table width in sequence to obtain a structured input vector;
[0009] The structured input vector is analyzed by a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal, and the policy verification model is obtained based on historical policy data and gradient boosting model training.
[0010] In a second aspect, an intelligent underwriting device is provided, comprising:
[0011] An acquisition module is used to acquire identification field information in a target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number;
[0012] An extraction module, configured to extract a characteristic field corresponding to the identification field information, and construct a corresponding original table width based on the characteristic field information;
[0013] A processing module, configured to perform encoding processing and embedding processing on the original table width in sequence to obtain a structured input vector;
[0014] The underwriting module is used to analyze the structured input vector through a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal. The policy verification model is trained based on historical policy data and a gradient boosting model.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned intelligent underwriting method when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned intelligent underwriting method are implemented.
[0017] In the solution implemented by the above-mentioned intelligent underwriting method, apparatus, computer device, and storage medium, the method includes: first, obtaining identification field information from the target policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number. Furthermore, the feature fields corresponding to the identification field information can be extracted, and the corresponding original table width can be constructed based on the feature field information. Thus, the original table width can be sequentially encoded and embedded to obtain a structured input vector. Finally, the structured input vector can be analyzed using a policy verification model to obtain an analysis result; wherein the analysis result is used to characterize whether the target policy is abnormal. The policy verification model is trained based on historical policy data and a gradient boosting model. In the present invention, by extracting the identification field information from the target policy, further obtaining the corresponding feature fields, constructing the original table width, and then performing encoding and embedding, the structured semantic features carried by the identification field can be effectively preserved, helping to accurately reflect the relationship between the policyholder, insured, agent, and their affiliated institution. During the analysis phase of the policy verification model, the introduction of a gradient boosting model trained on historical policy data can fully tap into historical behavior patterns and risk characteristics, improving the model's sensitivity and accuracy in identifying abnormal policies. Compared to traditional verification methods based on rules or static field comparisons, this approach has stronger generalization and stability when processing large-scale, heterogeneous policy data, helping to prevent complex fraud, optimize audit efficiency, and reduce the need for manual intervention. In addition, inputting the model in the form of structured input vectors significantly improves modeling efficiency and adaptability to downstream tasks, providing technical support for risk prevention and control throughout the policy lifecycle management. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 This is a schematic diagram of an application environment of the intelligent underwriting method according to one embodiment of the present invention;
[0020] Figure 2 This is a flow chart of an intelligent underwriting method according to an embodiment of the present invention;
[0021] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S20;
[0022] Figure 4This is a schematic diagram of the structure of an intelligent underwriting device in one embodiment of the present invention;
[0023] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The intelligent underwriting method provided by the embodiment of the present invention can be applied in the following situations: Figure 1In an application environment, the client communicates with the server through a network. The server can obtain the identification field information in the target insurance policy through the client; wherein, the identification field information includes the policyholder number, agent number, insured number, agent department number and institution number; extract the feature field corresponding to the identification field information, and construct the corresponding original table width based on the feature field information; perform encoding and embedding processing on the original table width in turn to obtain a structured input vector; analyze the structured input vector through the insurance policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target insurance policy is abnormal. The insurance policy verification model is obtained based on historical insurance policy data and gradient boosting model training, and finally the analysis result is fed back to the client. In the present invention, by extracting the identification field information in the target insurance policy and further obtaining the corresponding feature field, constructing the original table width and then performing encoding and embedding processing, the structured semantic features carried by the identification field can be effectively retained, which helps to accurately reflect the relationship between the policyholder, the insured, the agent and their affiliated institutions. During the analysis phase of the policy verification model, the introduction of a gradient boosting model trained on historical policy data can fully tap into historical behavior patterns and risk characteristics, and improve the model's sensitivity and accuracy in identifying abnormal policies. Compared to traditional verification methods based on rules or static field comparisons, this approach has stronger generalization capabilities and stability when processing large-scale, heterogeneous policy data, helping to prevent complex fraud, optimize audit efficiency, and reduce the need for manual intervention. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0027] See also Figure 2 As shown, Figure 2 A flowchart of an intelligent underwriting method provided by an embodiment of the present invention is provided. The method includes the following steps:
[0028] S10: Obtain identification field information in the target insurance policy.
[0029] The identification field information includes the policyholder number, agent number, insured number, agent department number and institution number.
[0030] For example, in the financial sector, insurance policies are the core carrier of insurance product transactions, and their related identification fields are crucial for customer identification and transaction tracing. By obtaining the policyholder and insured number from the target policy, the risk responsible party and beneficiary can be clearly identified, facilitating accurate matching of insurance contract content. Simultaneously, extracting the agent number, department number, and institution number allows for tracing the source of business and clarifying the sales chain, playing a key role in subsequent compliance reviews and risk assessments.
[0031] By obtaining identification field information, it can be used as the basic unit of financial business data, providing structured support for customer risk profiling, agent behavior analysis and inter-institutional business linkage, effectively improving the transparency of insurance operations and data governance capabilities.
[0032] S20: Analyze the key element information through a text analysis model to obtain structured information corresponding to the target text.
[0033] The structured information includes title information, text information and text summary information.
[0034] For example, key element information exists in unstructured data such as insurance policy terms, customer application materials, claims instructions, and regulatory reports. This can be automatically analyzed using a text analysis model, enabling accurate identification and extraction of title information, body information, and text summaries from the target text, thereby achieving efficient structural transformation of the target text. This processing approach not only improves the automation level of data organization but also lays a solid foundation for downstream data analysis and model input.
[0035] Furthermore, the extraction of structured information enables quantifiable processing of complex target text, providing strong support for scenarios such as intelligent risk control, compliance review, and customer service. Title information facilitates quick identification of the subject matter, while the main text retains key business details. The text summary facilitates rapid understanding and judgment during high-level decision-making and auditing.
[0036] The above-mentioned analysis method based on natural language processing technology has significantly enhanced the capabilities of financial technology in data deconstruction, risk warning and knowledge mining, and laid a technical foundation for the intelligent management of large-scale unstructured financial information.
[0037] Among them, such as Figure 3 As shown, in step S20, that is, extracting the characteristic field corresponding to the identification field information and constructing the corresponding original table width based on the characteristic field information, the following steps are included:
[0038] S21: Extracting the characteristic field corresponding to the identification field information.
[0039] Among them, the characteristic fields include gender, age, number of insurance purchases, insured amount, product type, payment method, historical underwriting results and region.
[0040] S22: performing a standardization operation on the feature field to obtain a unified feature field.
[0041] The standardization operation includes unified field naming and unified format operations.
[0042] S23: Arrange the unified feature fields according to a preset field order to obtain arranged unified field features.
[0043] S24: Combining the arranged unified field features into a structured table to obtain the original table width.
[0044] It should be understood that although the identification field information can clearly identify the unique identity of the customer or agent, it alone cannot support in-depth analysis of dimensions such as behavioral patterns and risk levels. Therefore, it is particularly important to further extract feature fields associated with these identification field information. Through step S21, feature fields such as gender, age, number of insurance policies, insurance amount, product type, payment method, historical underwriting results and location are obtained, and a comprehensive risk profile can be constructed from demographic attributes, transaction behavior characteristics and geographical dimensions. These feature fields not only reflect the static information of the insured or the insured in the insurance behavior, but also include dynamic change factors, laying a solid data foundation for subsequent model building and data analysis. Especially in financial scenarios, such multidimensional features help to accurately portray customer preferences, identify potential moral risks, and predict potential fraudulent behavior.
[0045] To ensure the consistency and availability of feature fields, steps S22 to S24 can be performed around the standardization, sorting, and tabulation of feature fields to construct an original table width with a clear structure and uniform format. During the standardization process, unified field naming and formatting can effectively avoid field confusion caused by different formats of different systems or data sources, improving data quality and system compatibility. Subsequently, the fields are arranged and combined into a structured table using a preset order, forming an original table width with a fixed column structure. This not only facilitates subsequent coding, modeling, and storage, but also makes the model more stable and accurate when processing large quantities of insurance data.
[0046] The above steps improve the automation and standardization of target policy processing and are key links in improving data asset availability and risk management capabilities in FinTech scenarios.
[0047] S30: performing encoding processing and embedding processing on the original table width in sequence to obtain a structured input vector.
[0048] For example, in the processing of financial and insurance data, although the original table width already has a structured format, the feature fields contained therein may be categorical, numerical, or mixed data. Directly inputting the data into the model can easily lead to low computational efficiency or poor learning effects. Therefore, step S30 performs encoding and embedding processing on the original table width in sequence, first performing label encoding, unique hot encoding, or embedding vector conversion on the categorical field, normalizing or standardizing the numerical field, and then uniformly converting all features into structured input vectors of fixed dimensions. This processing method not only retains the semantic information of the original features, but also improves the model's ability to learn high-dimensional features, which helps to improve the accuracy and efficiency of tasks such as complex risk modeling, fraud identification, and customer behavior prediction in the financial field.
[0049] In some embodiments, the combining of the arranged unified field features into a structured table to obtain the original table width includes: performing a preprocessing operation on the original table width to obtain valid feature data; wherein the preprocessing operation includes integrity checking, filling missing fields, and filtering logical conflict fields; the encoding and embedding processing of the original table width in sequence to obtain a structured input vector includes: performing the encoding and embedding processing on the valid feature data in sequence to obtain the structured input vector.
[0050] For example, to ensure the quality and reliability of modeling data, after the arranged unified field features are combined into a structured table, that is, the original table width, further preprocessing operations are required to eliminate potential data problems. This preprocessing operation includes three key steps: integrity check, missing field filling, and logical conflict field filtering. Among them, integrity check is used to check whether each feature field has null values or illegal data formats to ensure the integrity and rationality of the data content; missing field filling uses the mean, median or model prediction to fill the data gaps to avoid affecting the model performance due to incomplete information; logical conflict field filtering checks the logical consistency between fields, for example, the age field should not be less than the age calculated by the year of birth, thereby improving the consistency and credibility of the data. These operations are particularly important in financial scenarios, because data quality directly determines the accuracy of risk judgment and the stability of the system.
[0051] Furthermore, after preprocessing, encoding and embedding can be performed sequentially on valid feature data to form structured input vectors for model learning. During this process, categorical fields such as product type and payment method can be mapped to high-dimensional numerical representations using one-hot encoding or embedding vectors, preserving their semantic information while reducing sparsity. Numerical fields such as insured amount and number of insurance policies are normalized to ensure that all features enter the model at the same numerical scale, which helps stabilize and accelerate the convergence of the model learning process.
[0052] In some embodiments, the encoding and embedding processes are performed on the original table width in sequence to obtain a structured input vector, including: encoding the classification feature field in the original table width to obtain a numerical feature field; normalizing the numerical feature field to standardize the numerical feature field to a preset interval range to obtain a standardized feature field; merging the standardized feature field with the embedding vector, and splicing them into the structured input vector in a preset order; wherein the embedding vector is a semantic vector of a preset dimension and is generated based on a pre-trained language model.
[0053] For example, in order to improve the model's ability to identify different types of features in insurance data, various fields in the original table width can be processed step by step according to the feature type. First, the categorical feature fields in the original table width (such as product type, payment method, location, etc.) are encoded. For example, label encoding, unique hot encoding or embedding mapping can be used to convert them into numerical feature fields to meet the model calculation requirements. Subsequently, all numerical feature fields are uniformly normalized and mapped to a preset interval (such as [0,1] or [-1,1]), thereby eliminating the interference of different dimensions on model training and ensuring that each feature has a balanced influence during the learning process. This standardization process is particularly critical in financial modeling, which helps to improve the model convergence speed, reduce the risk of overfitting, and enhance the comparability between features.
[0054] Furthermore, after encoding the categorical fields and normalizing the numerical fields, the standardized feature fields can be merged with the embedding vector and concatenated according to the preset field order to ultimately construct a structured input vector. This processing method, which combines semantic embedding with standardized numerical features, significantly enhances the model's ability to understand complex financial data structures, not only improving prediction accuracy but also providing stronger data support and a foundation for intelligent decision-making in financial technology scenarios such as insurance fraud identification and customer risk assessment.
[0055] In some embodiments, the encoding processing of the classification feature fields in the original table width to obtain numerical feature fields includes: identifying the classification feature fields in the original table width, wherein the classification feature fields include occupational category, product type, insurance channel, and payment method; converting the classification feature fields with a frequency higher than a preset threshold into corresponding numerical feature fields using a target encoding method; and converting the classification feature fields with a frequency lower than the preset threshold into corresponding numerical feature fields using a label mapping encoding method.
[0056] For example, in order to more efficiently process the categorical feature fields in the original table width, key categorical fields such as occupational category, product type, insurance channel, and payment method can be first identified and differentially coded according to their frequency of occurrence in the data set. For high-frequency categorical fields with a frequency higher than a preset threshold, target coding is used to convert them into numerical features, that is, the numerical representation is performed by the statistical average of the target variable corresponding to the category (such as risk label or loss ratio), thereby retaining its relevance to the model output; for low-frequency categorical fields with a frequency lower than the threshold, label mapping coding is used, that is, a unique numerical number is assigned to each category to avoid rare categories generating noise in the target coding.
[0057] This hierarchical coding strategy for classification fields improves the model's ability to learn high-frequency patterns in financial scenarios, effectively suppresses low-frequency interference, and enhances the stability and generalization ability of the overall feature expression.
[0058] S40: Analyze the structured input vector through the insurance policy verification model to obtain an analysis result.
[0059] The analysis result is used to characterize whether the target policy is abnormal, and the policy verification model is trained based on historical policy data and a gradient boosting model.
[0060] For example, in step S40, the policy verification model analyzes the constructed structured input vector, outputting an analysis result used to determine whether the target policy contains anomalies. This policy verification model is trained on a large amount of historical policy data and uses a gradient boosting model as its core algorithm. It possesses strong nonlinear modeling and feature interactive learning capabilities, enabling it to accurately capture potential risk characteristics and anomaly patterns in policies.
[0061] By comparing and learning the current policy characteristics with normal and abnormal cases in historical data, the policy verification model can effectively identify risks such as suspicious insurance behavior, duplicate insurance, and false agency, providing intelligent risk control support for insurance business.
[0062] In some embodiments, the method further includes: obtaining a training data set and the gradient boosting model; wherein the training data set includes the historical policy data; labeling the training data set to obtain a labeling result, wherein the labeling result includes an analysis result corresponding to the historical policy data; training the gradient boosting model through the training data set and the labeling result to obtain the policy verification model.
[0063] On the basis of the above embodiment, after obtaining the policy verification model, it also includes: iteratively training the policy verification model based on the training data set and the annotation results to extract data features, and calculate the loss function; using a preset method to iteratively train the loss function for the purpose of reducing the loss function value until the loss function value is less than the expected threshold; based on the loss function after iterative training, obtaining the iterative policy verification model.
[0064] Specifically, a training dataset containing historical policy data may be collected for training. For example, the training dataset may be obtained through manual collection, web crawlers, or public datasets, and this application does not limit this.
[0065] Furthermore, each historical policy data set can be labeled to obtain a corresponding labeling result. The labeling result is then used as the label for the input data set. Each labeled training data set is then input into the gradient boosting model for supervised learning. Training ends when a training end condition is met, such as when the number of training times reaches a threshold or when the model output accuracy reaches a threshold, resulting in a trained policy verification model.
[0066] In the embodiment of the present application, the training data set and the annotation results can be input into the gradient boosting model for supervised learning, and then the policy verification model can be trained to obtain the annotation results.
[0067] The above embodiments can enhance data quality and diversity during the policy verification model training process, and improve the model's generalization ability and practical application effect.
[0068] It is understandable that in order to train a policy verification model with higher accuracy, the policy verification model can be repeatedly trained to continuously reduce the loss function until the loss function meets the expected threshold requirements, and then a more accurate labeling result can be obtained based on the iterated policy verification model.
[0069] It should be noted that this application does not limit the above-mentioned preset method and expected threshold. For example, the preset method can be a gradient descent algorithm, a batch gradient descent algorithm, a stochastic gradient descent algorithm, etc. This application takes the gradient descent algorithm as an example for illustration.
[0070] The purpose of the gradient descent algorithm is to find the minimum value of the loss function through iteration, or to converge to the minimum value. Geometrically speaking, the gradient descent algorithm is that where the function changes most rapidly, the gradient decreases most rapidly in the direction opposite to the vector, making it easier to find the minimum value of the function. Based on this, in an embodiment of the present application, the gradient descent algorithm can be used to iteratively train the policy verification model so that the loss function is continuously reduced, thereby reducing the error of the calculation results.
[0071] In an embodiment of the present application, the policy verification model is repeatedly iteratively trained using a gradient descent algorithm so that the loss function is continuously reduced to obtain an iterated policy verification model, and then a more accurate labeling result can be obtained based on the iterated policy verification model.
[0072] It can be seen that in the above scheme, by extracting the identification field information in the target insurance policy and further obtaining the corresponding feature fields, encoding and embedding the original table width, the structured semantic features carried by the identification field can be effectively retained, which helps to accurately reflect the relationship between the policyholder, the insured, the agent and their affiliated institutions. In the analysis stage of the insurance policy verification model, the introduction of the gradient boosting model trained based on historical insurance policy data can fully explore historical behavior patterns and risk characteristics, and improve the sensitivity and accuracy of the model in identifying abnormal insurance policies. Compared with traditional verification methods based on rules or static field comparisons, this method has stronger generalization capabilities and stability when processing large-scale, heterogeneous insurance policy data, which helps to prevent complex fraud, optimize audit efficiency, and reduce the need for manual intervention.
[0073] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0074] In one embodiment, an intelligent underwriting device is provided, which corresponds one-to-one with the intelligent underwriting method in the above embodiment. Figure 4 As shown, the intelligent underwriting device includes an acquisition module 101, an extraction module 102, a processing module 103, and an underwriting module 104. The functional modules are described in detail as follows:
[0075] The acquisition module 101 is used to acquire identification field information in the target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number and institution number;
[0076] An extraction module 102 is configured to extract a characteristic field corresponding to the identification field information and construct a corresponding original table width based on the characteristic field information;
[0077] The processing module 103 is configured to perform encoding processing and embedding processing on the original table width in sequence to obtain a structured input vector;
[0078] The underwriting module 104 is used to analyze the structured input vector through a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal. The policy verification model is obtained based on historical policy data and gradient boosting model training.
[0079] The extraction module 102 is used to extract the characteristic fields corresponding to the identification field information, wherein the characteristic fields include gender, age, number of insurance policies, insurance amount, product type, payment method, historical underwriting results and location; perform standardization operations on the characteristic fields to obtain unified characteristic fields; wherein the standardization operations include unified field naming and unified format operations; arrange the unified characteristic fields according to a preset field order to obtain the arranged unified field characteristics; combine the arranged unified field characteristics into a structured table to obtain the original table width.
[0080] The extraction module 102 is configured to perform a preprocessing operation on the original table width to obtain valid feature data; wherein the preprocessing operation includes integrity checking, filling missing fields, and filtering logical conflict fields.
[0081] The processing module 103 is configured to perform the encoding process and the embedding process on the effective feature data in sequence to obtain the structured input vector.
[0082] The processing module 103 is used to encode the classification feature fields in the original table width to obtain numerical feature fields; perform normalization on the numerical feature fields to standardize the numerical feature fields to a preset interval range to obtain standardized feature fields; merge the standardized feature fields with the embedding vectors, and splice them into the structured input vectors in a preset order; wherein the embedding vectors are semantic vectors of preset dimensions and are generated based on a pre-trained language model.
[0083] Processing module 103 is used to identify classification feature fields in the original table width, wherein the classification feature fields include occupation category, product type, insurance channel, and payment method; convert the classification feature fields whose frequency is higher than a preset threshold into corresponding numerical feature fields using a target coding method; and convert the classification feature fields whose frequency is lower than the preset threshold into corresponding numerical feature fields using a label mapping coding method.
[0084] In one embodiment, the acquisition module 101 is further used to: obtain a training data set and the gradient boosting model; wherein the training data set includes the historical policy data; annotate the training data set to obtain an annotation result, wherein the annotation result includes an analysis result corresponding to the historical policy data; train the gradient boosting model through the training data set and the annotation result to obtain the policy verification model.
[0085] In one embodiment, the acquisition module 101 is also used to: iteratively train the policy verification model based on the training data set and the annotation results to extract data features, and calculate the loss function; iteratively train the loss function using a preset method for the purpose of reducing the loss function value until the loss function value is less than the expected threshold; and obtain the iterative policy verification model based on the loss function after iterative training.
[0086] The present invention provides an intelligent underwriting device that extracts the identification field information in the target insurance policy, further obtains the corresponding feature fields, constructs the original table width, and then performs encoding and embedding processing. It can effectively retain the structured semantic features carried by the identification field, and helps to accurately reflect the relationship between the policyholder, the insured, the agent and their affiliated institutions. In the analysis stage of the insurance policy verification model, the introduction of the gradient boosting model trained based on historical insurance policy data can fully explore historical behavior patterns and risk characteristics, and improve the sensitivity and accuracy of the model in identifying abnormal insurance policies. Compared with traditional verification methods based on rules or static field comparisons, this method has stronger generalization capabilities and stability when processing large-scale, heterogeneous insurance policy data, which helps to prevent complex fraud, optimize audit efficiency, and reduce the need for manual intervention.
[0087] The specific definition of the intelligent underwriting device can be found in the definition of the intelligent underwriting method above and will not be repeated here. Each module in the above-mentioned intelligent underwriting device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0088] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of an intelligent underwriting method.
[0089] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an intelligent underwriting method.
[0090] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0091] Obtain identification field information from the target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number;
[0092] Extracting the characteristic field corresponding to the identification field information, and constructing a corresponding original table width based on the characteristic field information;
[0093] Performing encoding and embedding processing on the original table width in sequence to obtain a structured input vector;
[0094] The structured input vector is analyzed by a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal, and the policy verification model is obtained based on historical policy data and gradient boosting model training.
[0095] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0096] Obtain identification field information from the target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number;
[0097] Extracting the characteristic field corresponding to the identification field information, and constructing a corresponding original table width based on the characteristic field information;
[0098] Performing encoding and embedding processing on the original table width in sequence to obtain a structured input vector;
[0099] The structured input vector is analyzed by a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal, and the policy verification model is obtained based on historical policy data and gradient boosting model training.
[0100] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0101] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0102] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0103] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An intelligent underwriting method, characterized in that: The method comprises: Obtain identification field information from the target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number; Extracting the characteristic field corresponding to the identification field information, and constructing a corresponding original table width based on the characteristic field information; Performing encoding and embedding processing on the original table width in sequence to obtain a structured input vector; The structured input vector is analyzed by a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal, and the policy verification model is obtained based on historical policy data and gradient boosting model training.
2. The method according to claim 1, characterized in that The extracting the characteristic field corresponding to the identification field information and constructing a corresponding original table width based on the characteristic value includes: Extracting characteristic fields corresponding to the identification field information, wherein the characteristic fields include gender, age, number of insurance purchases, insured amount, product type, payment method, historical underwriting results, and region; Performing a standardization operation on the characteristic fields to obtain unified characteristic fields; wherein the standardization operation includes a unified field naming and a unified format operation; Arranging the unified feature fields according to a preset field order to obtain arranged unified field features; The arranged unified field features are combined into a structured table to obtain the original table width.
3. The method according to claim 2, characterized in that After combining the arranged unified field features into a structured table to obtain the original table width, the following steps are performed: The original table width is preprocessed to obtain valid feature data; wherein the preprocessing operation includes integrity checking, filling missing fields, and filtering logical conflict fields; The encoding and embedding processing is performed on the original table width in sequence to obtain a structured input vector, including: The encoding process and the embedding process are sequentially performed on the effective feature data to obtain the structured input vector.
4. The method according to claim 1, wherein The encoding and embedding processing is performed on the original table width in sequence to obtain a structured input vector, including: Encoding the classification feature fields in the original table width to obtain numerical feature fields; Normalizing the numerical feature field to standardize the numerical feature field to a preset range to obtain a standardized feature field; The standardized feature field is merged with the embedding vector and spliced into the structured input vector in a preset order; wherein the embedding vector is a semantic vector of a preset dimension and is generated based on a pre-trained language model.
5. The method according to claim 4, characterized in that The encoding process of the classification feature field in the original table width to obtain a numerical feature field includes: Identifying classification feature fields in the original table width, wherein the classification feature fields include occupation category, product type, insurance channel, and payment method; Converting categorical feature fields with frequencies higher than a preset threshold into corresponding numerical feature fields using a target encoding method; and The classification feature fields whose frequencies are lower than the preset threshold are converted into corresponding numerical feature fields using a label mapping encoding method.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining a training data set and the gradient boosting model; wherein the training data set includes the historical policy data; Annotating the training data set to obtain an annotation result, wherein the annotation result includes an analysis result corresponding to the historical policy data; The gradient boosting model is trained using the training data set and the annotation results to obtain the insurance policy verification model.
7. The method according to claim 6, characterized in that After obtaining the insurance policy verification model, the method further includes: Iteratively train the policy verification model based on the training data set and the annotation results to extract data features and calculate the loss function; Iteratively training the loss function using a preset method for the purpose of reducing the loss function value until the loss function value is less than an expected threshold; Based on the loss function after iterative training, an iterative policy verification model is obtained.
8. An intelligent underwriting device, characterized in that: include: An acquisition module is used to acquire identification field information in a target insurance policy; wherein the identification field information includes the policyholder number, agent number, insured number, agent department number, and institution number; An extraction module, configured to extract a characteristic field corresponding to the identification field information, and construct a corresponding original table width based on the characteristic field information; A processing module, configured to perform encoding processing and embedding processing on the original table width in sequence to obtain a structured input vector; The underwriting module is used to analyze the structured input vector through a policy verification model to obtain an analysis result; wherein, the analysis result is used to characterize whether the target policy is abnormal. The policy verification model is trained based on historical policy data and a gradient boosting model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the intelligent underwriting method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the intelligent underwriting method according to any one of claims 1 to 7 are implemented.