A Sample Sampling Method and Device for Insurance Texts

The sample sampling method of text vectorization and semi-supervised learning through the Sent-Bert model solves the problem of large workload in text sample annotation in the financial insurance industry, achieves higher sample diversity and model robustness, and improves classification accuracy and training efficiency.

CN114741504BActive Publication Date: 2025-07-22ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210219956.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-07-22
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

In the financial insurance industry, text sample annotation work is large and difficult to achieve sample diversity and model robustness, especially in terms of difficult sample mining.

Method used

Text vectorization is used by Sent-Bert model, combined with the farthest point sampling and distribution-based resampling method, sample sampling is performed through semi-supervised learning to ensure the consistency of sampling sample space and sample balance between classes.

Benefits of technology

While reducing the annotation workload, it achieves higher sample diversity and model robustness, improving the classification accuracy and training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741504B_ABST
    Figure CN114741504B_ABST
Patent Text Reader

Abstract

The present invention discloses a sample sampling method and device for insurance texts. The method includes two parts: text vectorization based on semantics and semi-supervised sampling. The semi-supervised sampling is further divided into steps such as farthest point sampling and annotation, resampling based on distribution and annotation of resampled samples, and verification of model classification accuracy. The method of the present invention performs sample sampling based on semantic vectorization combined with semi-supervised learning method. Under the condition of extremely few labeled samples, it can achieve model accuracy and robustness comparable to those of fully labeled samples, while significantly reducing the computational and time costs of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of financial insurance text recognition, and particularly relates to a sample sampling method and device for insurance texts. Background Art

[0002] With the development of informatization in the financial insurance industry, relevant business data has grown rapidly. The shortage of manpower and the need for refined management have led to the application of more and more deep learning models, but the corresponding data annotation task volume has also increased rapidly. How to annotate fewer samples to achieve better sample diversity and model robustness has become an important direction in the current model research of the financial insurance industry, which is called the problem of hard sample mining. Hard sample mining is also an important research content in deep learning. Related research is divided into two directions: one is to increase the learning rate of hard samples by weighting. Related research includes Focal loss. The advantage is that it can improve the model convergence speed, but the disadvantage is that the annotation workload is not reduced; the other is to sample all samples in an unsupervised or semi-supervised manner to find confusing hard samples. This method can not only reduce the number of annotated samples but also improve the model convergence speed, which is more effective in actual engineering applications.

[0003] Text sample sampling usually includes two important steps: vectorization and uniform sampling. The vectorization process ensures that the similarity remains unchanged before and after the text is converted into a vector. Uniform sampling ensures that the sample space coverage range and spatial structure remain unchanged before and after sampling. Text vectorization methods include keyword-based vectorization such as TF-IDF, BM25, etc., and semantic-based vectorization such as Topic-embedding, Sent-Bert. Uniform sampling methods include farthest point sampling, etc. Chinese Patent CN 112364130A discloses a text sampling method that uses character encoding for text vectorization and uses the edit distance to calculate the text distance. However, this method cannot well represent the semantic similarity between texts. Chinese Patent CN 112329427A discloses a method for obtaining short message samples, which uses a multiple de-duplication method for short message sampling, uses short message templates combined with features such as the short message source time for similarity quantification, and uses the classification uncertainty index as the last screening criterion for annotated samples. This method is effective for short message texts but also does not consider the semantic similarity of samples. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a sample sampling method and device for insurance texts.

[0005] To achieve the above technical objectives, the technical solution of the present invention is as follows:

[0006] The first aspect of the embodiment of the present invention provides a sample sampling method for insurance texts, including the following steps:

[0007] (1) Construct a pre-trained model for text similarity determination, and vectorize the text through this model to obtain a total vector set;

[0008] (2) Perform farthest point initial sampling and annotation on the total vector set to achieve uniform distribution of the sampling in space, and obtain a selected point set;

[0009] (3) Set the number of samples, and resample the initial sample set based on the inter-class distribution model to update the selected point set;

[0010] (4) Set the sampling radius ratio and accuracy threshold, and use the updated selected point set to perform model training and accuracy verification until the accuracy is met, and complete the sample sampling.

[0011] Further, the pre-trained model for text similarity determination is Sent-Bert; Sent-Bert is a text similarity measurement model, which uses the pre-trained Bert as the underlying model, and adds a pair of pooling-based embedding layers to this underlying model to form a siamese network with shared underlying parameters.

[0012] Further, the fine-tuning training is specifically as follows: Sent Bert is fine-tuned and trained through a Chinese database including LCQMC, STS-B, and ATEC with manually annotated similarities.

[0013] Further, input a pair of insurance texts into the pre-trained model for text similarity determination, and the output is two vectors; the first vector is the result of text vectorization, and together they form the total vector set; the second vector is empty.

[0014] Further, the step (2) specifically includes the following sub-steps:

[0015] (2.1) Set the number of samples in the initial sampling set according to the similarity of the samples and few-shot learning;

[0016] (2.2) Select an initial point, select the point farthest from the data center. For text data, calculate the similarity between vectors using cosine similarity, sort all similarities, and use the maximum similarity as the vector farthest from other text vectors to establish a selected point set;

[0017] (2.3) Calculate the distance between other points and the selected point set, select the farthest point, and update the selected point set;

[0018] (2.4) Repeat the above steps (2.1) to (2.3) until the number of samples in the selected point set reaches the number of samples set in the initial sampling set;

[0019] (2.5) For the sampling samples obtained in step (2.4), perform manual annotation according to text classification.

[0020] Further, step (3) is specifically as follows: Assume that each class of samples conforms to a Gaussian distribution, calculate the center points and within-class densities of different classes of samples; calculate the class boundaries and the boundary points between different class centers, and represent them as the weighted means of the two class center points; calculate the sampling quantity according to the boundary point density, set the sampling radius using the unbiased estimation of the Gaussian standard deviation of the large-density class, and perform resampling around the boundary points to update the selected point set.

[0021] Further, assume that each class of samples conforms to a Gaussian distribution, calculate the center points C = [c0,...] of different classes of samples, the within-class densities D = [d0...], and the centers of different classes are the means of the samples within the class. The calculation formula is as follows:

[0022]

[0023]

[0024] where t i represents the label value of the i-th category, and l k is the label value of the k-th sample in the selected point set; a k represents the vector of the k-th sample, and c i is the center point of the i-th class;

[0025] Calculate the class boundaries. The boundary points between different class centers are represented as the weighted means of the two class center points. The calculation formula is as follows:

[0026]

[0027] In the above formula, b ij represents the boundary point between the i-th and j-th classes, c i represents the center point of the i-th class, is the normalization weight; norm represents array normalization;

[0028] Sample around the boundary points; the boundary points involve two classes. Calculate the densities of the two classes at the boundary points, that is:

[0029] d′ i = count(b ij - s k <r i ), s k ∈S - A

[0030] In the formula, d’ represents the density, s k is the k-th sample outside the selected point set A, S is the set of all samples, and r iis the sampling radius for the i-th class, which is an unbiased estimate of the Gaussian standard deviation of the samples within the sampling radius class, i.e.:

[0031]

[0032] where n is the total number of samples in the selected point set of the i-th class, a ij is the j-th sample of the i-th class, and c i is the center point of the i-th class; the resampling point is the number of samples within a certain radius from the boundary point. To avoid duplication, only the class with the higher density among the two classes corresponding to the boundary point is sampled, which is defined as:

[0033]

[0034] where s k is the k-th sample outside the selected point set A, and b ij represents the boundary point between the i-th and j-th classes. When the density of the i-th class is higher, the resampling points are selected according to r i and the samples s k that meet the above conditions are added to the selected point set to update the selected point set.

[0035] Furthermore, the specific steps of step (4) are as follows: Set the sampling radius ratio and the accuracy threshold, sample several pieces of data as the test set, and use the updated selected point set as the training set; Use the training set to train the classifier, and then use the classifier to classify and predict the test set; Calculate the accuracy rate. If the preset accuracy threshold is reached, the sample sampling is completed; If the prediction accuracy threshold is not met, adjust the sampling radius ratio and repeat step (3) until the preset accuracy threshold is reached to complete the sample sampling.

[0036] The second aspect of the embodiments of the present invention provides a sample sampling device for insurance texts, including one or more processors for implementing the above-mentioned sample sampling method for insurance texts.

[0037] The third aspect of the embodiments of the present invention provides a computer-readable storage medium, on which a program is stored. The program, when executed by a processor, is used to implement the above-mentioned sample sampling method for insurance texts.

[0038] The beneficial effects of the present invention: Using a text semantic model for text vectorization can calculate the similarity between texts more accurately. Secondly, a semi-supervised progressive sampling method based on farthest point and distribution density resampling is proposed to ensure the consistency of the sampling sample space and the overall sample space and the balance of the samples between classes. Thus, under the condition of the same accuracy, the effect of significantly reducing the annotation workload is achieved. The method of the present invention can achieve higher sample diversity and model robustness with fewer labeled samples, and is applicable to insurance text classification and information extraction tasks with a large number of samples and heavy annotation tasks. Brief Description of the Drawings

[0039] Figure 1 is a framework diagram of the method of the present invention;

[0040] Figure 2 is a process flow of the method of the present invention;

[0041] Figure 3 is a schematic diagram of boundary points and sampling radius;

[0042] Figure 4 is a schematic diagram of the device of the present invention. Detailed Description of the Preferred Embodiments

[0043] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0044] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0045] The following will describe in detail the sample sampling method and device for insurance texts of the present invention with reference to the drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0046] The present invention proposes a semi-supervised sample sampling method for insurance texts. The so-called semi-supervised sampling herein includes an unsupervised part and a supervised part. The unsupervised part includes farthest point sampling and annotation, resampling based on distribution, and annotation of resampled samples; the supervised part includes text quantization and verification of the classification accuracy of the model. First, a general pre-trained semantic model is used to vectorize the text. Then, farthest point sampling is performed based on the text vectors. The number of samples needs to be set for sampling, and the output result is the initial sampling set. Then, resampling is performed based on the prior sample distribution. The resampling radius ratio needs to be set for resampling, and the output is the resampled set. Finally, training is performed on the sampling set and accuracy verification is performed on the test set. If the accuracy exceeds the set threshold, it is determined that the sampling is completed. The overall framework diagram and process are as Figure 1 、 2 shown. Specifically, it includes the following sub-steps:

[0047] (1) Vectorize the text using a pre-trained model for text similarity determination, specifically as follows:

[0048] In the embodiments of the present invention, Sent-Bert (i.e., Sentence-BERT) is used as the pre-trained model for text similarity determination to vectorize the text. The Sent-Bert is a text similarity metric model with the pre-trained Bert as the underlying model. The pre-training is to fine-tune the Sent-Bert using Chinese databases such as LCQMC, STS-B, and ATEC with manually annotated similarities; on this basis, a pair of pooling-based embedding layers are added to form a siamese network sharing underlying parameters, namely Sent Bert. The total sample set S is a Chinese database including LCQMC, STS-B, and ATEC. The total sample set S is input into the pre-trained model for text similarity determination. The input of the model is a pair of texts, and the output is two vectors. When converting a text into a vector, the second text needs to be filled with a null value, and then the first output vector is taken as the output result. After vectorization, the output result is recorded as the total vector set V = [v0, v1,...]. The Sent-Bert model combines the generality of Bert in various fields and its own semantic similarity representation ability, and can obtain a better text similarity measurement effect than the keyword-based model.

[0049] (2) Perform farthest point initial sampling and annotation on the total vector set V obtained in step (1) based on uniform space sampling, specifically including the following sub-steps:

[0050] After text vectorization, sample the samples. First, use the unsupervised sampling method - the farthest point algorithm - for sampling to ensure that the collected samples are evenly distributed in the sample space. The farthest point sampling is divided into the following steps:

[0051] (2.1) According to the target task, set the sampling quantity ń of the initial sampling set;

[0052] The sampling quantity can be estimated according to the similarity of samples and few-shot learning. The commonly used sampling quantities in current few-shot learning are 1, 5, 10, 20, 50, etc. When the sample similarity is high, 5-10 items are sampled for each type of sample to calculate the total sampling set. When the sample similarity is low, the quantity of the total sampling samples needs to be appropriately increased. Most of the lengths of car insurance description texts are between 30 and 60 characters, and the description objects are all related to car insurance. In the embodiments of the present invention, 0.8 is taken as the cosine similarity determination criterion, and a value higher than this is considered to have a high similarity. In addition, in few-shot learning, the class labels are known, but the class labels of the sampled samples are unknown. In order to enable the samples to cover different classes, considering the above factors comprehensively, in the embodiments of the present invention, the number of classes * 10 * 2 is used as the initial sampling quantity. The setting of the initial sample quantity will affect the subsequent resampling times. Setting an appropriate sample quantity can reduce the corresponding time cost, but has little impact on the size of the final sample set and the model accuracy.

[0053] (2.2) Select the initial point. Usually, the point farthest from the data center is selected. For text data, the cosine similarity is used to calculate the similarity between pairwise vectors, all the similarities are sorted, and the maximum similarity is used as the vector a0 = v that is farthest from other text vectors. i Establish the selected point set A = [a0];

[0054] (2.3) Calculate the distances between other points and the selected point set A, select the farthest point as v1, a1 = v1, and update the selected point set A = [a0, a1];

[0055] (2.4) Repeat the above steps (2.1) to (2.3) until the sample quantity of the selected point set reaches the sample quantity of the set initial sampling set, and the farthest point sampling is completed.

[0056] (2.5) Manually annotate the sampled samples according to text classification. In the embodiments of the present invention, taking text multi-class classification as an example, the annotation results corresponding to the sampled samples are L = [l0, l1...], where l i represents the text category, such as accident types like 'collision' and'rear-end collision' in car insurance.

[0057] (3) Resample the initial sample set based on the inter-class distribution model, which specifically includes the following steps:

[0058] Although farthest point sampling ensures that sample points are scattered relatively evenly in the sample space, it cannot guarantee the balance of samples from different classes. Since the data distributions of different classes are different, some classes of data are relatively compact while some are relatively sparse. At the same time, there are also differences in the distances between classes. Some classes are closer to each other and are more likely to be confused, while other classes are farther apart and are easier to distinguish. Therefore, for data with closer inter-class distances, more data needs to be sampled additionally, and vice versa, less data is sampled. Therefore, it is necessary to design a sampling method according to the distribution density of sample classes and the position of the inter-class boundary.

[0059] In the embodiments of the present invention, it is assumed that each class of samples conforms to a Gaussian distribution. The center points C = [c0,...] of different classes of samples and the within-class density D = [d0...] are calculated. Since farthest point sampling is uniform sampling by distance, the within-class density is approximately inversely proportional to the number of initial sampling samples, and the centers of different classes are the means of the within-class samples. The calculation formulas are as follows:

[0060]

[0061]

[0062] where, t i represents the label value of the i-th class, and l k is the label value of the k-th sample in the selected point set. a k represents the vector of the k-th sample, and c i is the center point of the i-th class.

[0063] Next, the class boundaries are calculated. The boundary points between the centers of different classes can be expressed as the weighted means of the center points of the two classes. The calculation formula is as follows:

[0064]

[0065] In the above formula, b ij represents the boundary point between the i-th and j-th classes, c i represents the center point of the i-th class, is the normalized weight. norm represents array normalization. The sum of the normalized array is 1, and the normalization operator is softmax. The greater the class density, the greater the normalized value, and the obtained boundary point is closer to the end with fewer within-class samples, as Figure 3 shown.

[0066] Sampling around the boundary points will start below. The boundary points and the sampling radius are the center and radius of the circle as Figure 3 shown. The boundary points involve two classes. First, calculate the densities of the two classes at the boundary points, that is:

[0067] d i ′ = count(b ij -sk <r i ), s k ∈S - A

[0068] In the formula, d’ represents density, s k is the k-th sample outside the selected point set A, S is the set of all samples, r i is the sampling radius of the i-th class, which is the unbiased estimate of the Gaussian standard deviation of the samples within the sampling radius class, that is:

[0069]

[0070] In the formula, n is the total number of samples in the selected point set of the i-th class, a ij is the j-th sample of the i-th class, c i is the center point of the i-th class. The resampled point is the number of samples within a certain radius from the boundary point. Here, to avoid repetition, only the class with the larger density among the two classes corresponding to the boundary point is sampled, defined as:

[0071]

[0072] In the formula, s k is the k-th sample outside the selected point set A, b ij represents the boundary point between the i-th and j-th classes. When the density of the i-th class is larger, the resampled points are selected according to r i , and the sample s k that meets the above conditions is added to the selected point set.

[0073] (4) Perform model training and testing on the selected point set updated in step (3). Use the sampled set as the training set. Among the total samples S minus the selected sample set A, sample several pieces of data as the test set B. Use the selected point set as the training set to train the classifier, and then use the classifier to perform classification prediction on the test set B. For text data classification, the current general classifier is Bert, which has good stability and accuracy and is based on a relatively good pre-trained model. Then, conduct manual recheck on the prediction results and calculate the prediction accuracy rate. Finally, according to the data situation, the business expert sets the accuracy threshold to evaluate whether the selected samples meet the requirements. The text classification accuracy threshold is usually set to 85%, and values higher than this are considered usable. If the accuracy does not reach the threshold after training with the selected samples, repeat the resampling and accuracy verification steps (3) - (4) until the accuracy meets the requirements. Among them, the resampling radius can be scaled proportionally with the increase in the number of sampling times. When the number of newly added samples in resampling is less than half of the number of newly added samples in the previous round, the sampling radius can be manually increased.

[0074] Example:

[0075] Taking the vehicle insurance claim text as an example, in this embodiment, the traffic accident description needs to be classified to predict the accident type, including collision, scratch, flat tire, rollover, falling object injury, etc. The claim text is as shown in Table 1 below. Therefore, a multi-class classification model needs to be trained. Before training, the most important task is to label the samples and divide them into a training set and a test set. The following table intercepts part of the actual data, and it can be seen that the vehicle insurance claim text has a strong correlation. For example, 'two vehicles collided', 'two vehicles crashed', and 'hit a third party' during a turn are of the same accident type. If all samples are labeled without discrimination, the labor cost is high, and due to factors such as unbalanced data samples, the final model will converge slowly and the classification effect will be poor. Therefore, according to the method proposed in the present invention, semi-supervised inter-class balanced sampling is performed on the samples to improve the model convergence speed and accuracy.

[0076] Table 1: Initial Sampling Set

[0077]

[0078] (1) Vectorization

[0079] In the embodiment, the Sent Bert model is finely tuned and trained using Chinese databases such as LCQMC, STS-B, and ATEC with artificially labeled similarity. The training set includes corpora such as finance, current politics, sports, and entertainment. Therefore, it can be used for semantic vectorization of vehicle insurance claim text. The input of the Sent Bert model is a pair of texts, and the output is the corresponding two vectors. In order to output the vector of a single text, a null text needs to be supplemented during input, and the vector corresponding to the actual claim text is selected as the vectorization result during output.

[0080] (2) Farthest Point Sampling and Labeling

[0081] After the text is vectorized, the farthest point method is used for sampling. First, to ensure the invariance of the sampling point set, the point farthest from the sample center is selected as the first sample, and other samples are screened in turn according to the subsequent steps. The size of the initial sampling set is set to about 10% of the total sample number, and the result is shown in the first column of the following table. The following table intercepts the accidents that occurred during a 'turn'. Comparing with Table 1, the similarity of the text has been greatly reduced, that is, the diversity of the same number of samples has increased. Finally, the text sampling set is labeled, and the labeled categories are shown in the second column of Table 2 below, including 12 types such as scratch, slip and fall, collision, etc.

[0082] Table 2: Resampled Set

[0083]

[0084] (3) Resampling Based on Density Distribution

[0085] Vehicle insurance claim texts exhibit strong inter-class imbalance. From the intercepted samples in the above table, there are more samples in categories such as 'collision' and fewer samples in categories such as 'natural combustion'. Resampling can make the inter-class distribution of samples more balanced. First, calculate the vector center point of each class of samples. For high-dimensional feature vectors, the class center point coordinate value is the mean of each dimension of the in-class sample points. Then, for each class center point, calculate the inter-class distance. In the embodiments of the present invention, the K-nearest neighbor method is used to find adjacent categories. Let K be 5, and there are a total of 12 classes, including duplicate adjacent categories, for a total of 60 pairs of adjacent classes. Next, calculate the boundary points and sampling radii of adjacent classes. Taking the boundary points as the center, find the samples within the sampling radius and perform farthest point sampling on these samples to form a resampling set. Finally, label the resampling set.

[0086] After resampling, the proportion of the number of samples between unbalanced classes is reduced from 50:1 to 10:1, achieving the purpose of supplementing more samples of small classes. Furthermore, through inter-class balance, the discrimination of the training model is improved.

[0087] (4) Model accuracy verification

[0088] An initial training set is obtained through resampling. In this step, it is necessary to determine whether it is necessary to supplement more labeled samples. First, extract 100 samples from the remaining samples after sampling as the test set. Then, use the labeled samples as the training set to train a text classification model. Next, use the model to predict the test set and record the prediction results. Finally, manually recheck the prediction results. Set the accuracy threshold to 85%, that is, only 15 samples in the test set are allowed to be misclassified. If the recheck accuracy exceeds the threshold, the sampling ends; otherwise, repeat step (3). The setting of the accuracy threshold depends on the specific application scenario. For scenarios with simple text semantics and short text lengths, a higher threshold value is set; for longer texts with complex semantics, a lower threshold value is set. The recommended range of the accuracy threshold is 75-90%. If the text classification accuracy is lower than 75%, the sample labels are usually considered to be less accurate and need to be relabeled.

[0089] Corresponding to the embodiments of the sample sampling method for insurance texts described above, the present invention also provides embodiments of a sample sampling device for insurance texts.

[0090] See Figure 4 , a sample sampling device for insurance texts provided by an embodiment of the present invention includes one or more processors for implementing the sample sampling method for insurance texts in the above embodiments.

[0091] The embodiments of the sample sampling device for insurance texts of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented through software, or through hardware, or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the sample sampling device for insurance texts of the present invention is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, usually according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0092] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.

[0093] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0094] The embodiments of the present invention also provide a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the sample sampling method for insurance texts in the above embodiments.

[0095] The computer-readable storage medium may be an internal storage unit of any data processing-capable device described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any data processing-capable device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any data processing-capable device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any data processing-capable device, and may also be used to temporarily store the data that has been output or is to be output.

[0096] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A sample sampling method for insurance texts, characterized in that, It includes the following steps: (1) Construct a pre-trained model for text similarity determination, and vectorize the text through this model to obtain a total vector set; (2) Perform farthest point initial sampling and annotation on the total vector set to achieve a uniform distribution of sampling in space and obtain a selected point set; (3) Set the number of samples, and resample the initial sample set based on the inter-class distribution model to update the selected point set; The specific steps of step (3) are as follows: Assume that each class of samples conforms to a Gaussian distribution, calculate the center points and intra-class densities of different classes of samples; Calculate the class boundaries and the boundary points between the centers of different classes, and represent them as the weighted means of the two class center points; Calculate the sampling quantity according to the boundary point density, set the sampling radius using the unbiased estimate of the Gaussian standard deviation of the large density class, and perform resampling around the boundary points to update the selected point set; (4) Set the sampling radius ratio and accuracy threshold, and use the updated selected point set to train the model and verify the accuracy until the accuracy is met to complete the sample sampling.

2. The sample sampling method for insurance texts according to claim 1, wherein The pre-trained model for text similarity determination is Sent-Bert; Sent-Bert is a text similarity measurement model, which uses the pre-trained Bert as the underlying model, and adds a pair of pooling-based embedding layers to this underlying model to form a twin network with shared underlying parameters.

3. The sample sampling method for insurance texts according to claim 2, wherein The specific pre-training is as follows: Fine-tune and train Sent Bert through a Chinese database including LCQMC, STS-B, and ATEC with manually annotated similarities.

4. The sample sampling method for insurance texts according to claim 1, characterized in that, Input a pair of insurance texts into the pre-trained model for text similarity determination, and the output is two vectors; The first vector is the result after text vectorization, and together they form the total vector set; The second vector is empty.

5. The sample sampling method for insurance texts according to claim 1, wherein The specific steps of step (2) include the following sub-steps: (2.1) Determine the number of samples in the initial sampling set according to the similarity of the samples and few-shot learning; (2.2) Select the initial point, select the point farthest from the data center. For text data, calculate the similarity between vectors using cosine similarity, sort all similarities, and use the maximum similarity as the vector farthest from other text vectors to establish a selected point set; (2.3) Calculate the distances between other points and the selected point set, select the farthest point, and update the selected point set; (2.4) Repeat the above steps (2.1) to (2.3) until the number of samples in the selected point set reaches the number of samples in the set initial sampling set; (2.5) Manually annotate the sampling samples obtained in step (2.4) according to text classification.

6. The sample sampling method for insurance texts according to claim 1, characterized in that Assume that each class of samples conforms to a Gaussian distribution, calculate the center points C = [c0,...] of different classes of samples, the intra-class densities D = [d0...], and the centers of different classes are the means of the samples within the class. The calculation formula is as follows: d i = 1 / ; ; Among them, t i represents the label value of the i-th category, and l k is the label value of the k-th sample in the selected point set; a k represents the vector of the k-th sample, and c i is the center point of the i-th category; Calculate the class boundaries, and the boundary points between the centers of different classes are represented as the weighted means of the two class center points. The calculation formula is as follows: ; In the above formula, b ij represents the boundary point between classes i and j, and c i represents the center point of class i, is the normalized weight; norm represents the normalization of the array; Sample around the boundary points; The boundary points involve two classes, calculate the densities of the two classes at the boundary points, that is: ; where d’ represents density, s k is the k-th sample outside the selected point set A, S is the set of all samples, r i is the sampling radius of the i-th class, which is the unbiased estimate of the Gaussian standard deviation of the samples within the sampling radius class, that is: ; where n is the total number of samples in the selected point set of the i-th class, and a ij is the j-th sample of the i-th class, and c i is the center point of the i-th class; the resampled points are the number of samples within a certain radius from the boundary points. To avoid duplication, only the class with a higher density among the two classes corresponding to the boundary points is sampled, which is defined as: ; where s k is the k-th sample outside the selected point set A, and b ij represents the boundary point between classes i and j. When the density of the i-th class is large, the resampled points are selected according to r i . The sample s k satisfying the above conditions is added to the selected point set to update the selected point set.

7. The sample sampling method for insurance texts according to claim 1, characterized in that, The specific steps of step (4) are as follows: Set the sampling radius ratio and the accuracy threshold, sample a number of data as the test set, and use the updated selected point set as the training set; Use the training set to train the classifier, and then use the classifier to classify and predict the test set; Calculate the accuracy rate. If the preset accuracy threshold is reached, the sample sampling is completed; If the predicted accuracy threshold is not met, adjust the sampling radius ratio and repeat step (3) until the preset accuracy threshold is reached to complete the sample sampling.

8. A sample sampling device for insurance texts, characterized in that, Comprising one or more processors for implementing the insurance text-oriented sample sampling method according to any one of claims 1-7.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program is used to implement the insurance text-oriented sample sampling method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Short message sample acquisition method and device

    CN112329427A

  • Sample sampling method and device and readable storage medium

    CN112364130A

  • Method for generating blue noise meshes on basis of farthest point optimization

    CN104036552A

  • Optimized downsampling SVM classification method based on potential function clustering and storage medium

    CN110059764A