A method and apparatus for identifying a risk insurance business

By identifying risky insurance business, the method obtains target risk identifiers, determines the semantic feature vector of each target address, and calculates similarity based on the relationship between addresses, thereby identifying risky addresses and associated risky insurance policies. This solves the problem in existing technologies that cannot identify multiple fake accounts of fraudulent agents and their related risky insurance policies, achieving a more efficient identification effect.

CN115760441BActive Publication Date: 2025-12-05SHENZHEN PAKRY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211380565.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-12-05
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

Existing technology cannot identify multiple aliases used by fraudulent agents and their associated risky insurance policies, resulting in low identification efficiency.

Method used

By acquiring the target risk identifier and the text information of the target address, the semantic feature vector of each target address is determined. Based on the association between each address and other addresses, the semantic feature vectors of N candidate addresses associated with each target address are determined, the similarity is calculated, the risk address is identified, and the risk policy and agent are associated.

Benefits of technology

It accurately identifies risky policies and risky agents associated with all risky addresses, resulting in higher identification efficiency and more comprehensive results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760441B_ABST
    Figure CN115760441B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of risk control management, and provides a method and device for identifying risk insurance business, which comprises the following steps: obtaining text information of each target address related to a target risk identifier; determining a semantic feature vector of each target address according to the text information of each target address; determining semantic feature vectors of N candidate addresses associated with each target address according to the semantic feature vector of each target address and the association relationship between each address and other addresses; obtaining N similarities of each target address according to the semantic feature vector of each target address and the semantic feature vectors of the N candidate addresses; determining a risk address of each target address according to the N similarities of each target address; and determining a risk policy related to each risk address as a risk policy and an agent associated with each risk address as a risk agent. Thus, the efficiency of identifying risk policies and risk agents can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of risk control management technology, and in particular relates to a method and apparatus for identifying risky insurance business. Background Technology

[0002] With the increasing demand for insurance services, insurance fraud has become more prevalent. One type of insurance fraud involves insurance agents deceiving customers into purchasing insurance policies, causing losses to insurance companies and seriously hindering the healthy development of the insurance industry.

[0003] Currently, insurance fraud is often identified manually. This process typically occurs when a customer applies for policy cancellation, at which point the agent handling the policy may discover fraudulent activity. The agent can then mark the policy as risky and the agent as fraudulent.

[0004] However, insurance agents may use multiple aliases registered at the same address. The above identification methods cannot identify multiple aliases used by fraudulent agents, nor can they identify risky insurance policies associated with these aliases, resulting in low identification efficiency. Summary of the Invention

[0005] This application provides a method and apparatus for identifying risky insurance business, which can solve the problem that the existing technology cannot identify multiple aliases of fraudulent agents, nor can it identify the risky insurance policies related to these aliases, resulting in low identification efficiency.

[0006] In a first aspect, embodiments of this application provide a method for identifying risk insurance business, including:

[0007] Obtain the target risk identifier, which may be the identifier of the target risk policy or the identifier of the target risk agent;

[0008] Query the text information for each target address associated with the target risk identifier, where the target address is either the address of the target risk policy or the address of the target risk agent;

[0009] Based on the text information of each target address, determine the semantic feature vector of each target address;

[0010] Based on the semantic feature vector of each target address and the association between each address and other addresses besides each address, determine the semantic feature vectors of N candidate addresses associated with each target address, where N is a positive integer;

[0011] Based on the semantic feature vector of each target address and the semantic feature vectors of the N candidate addresses associated with each target address, N similarities are obtained for each target address.

[0012] Based on N similarities for each target address, determine the risk address for each target address;

[0013] Based on the risk address of each target address, the insurance policies associated with each risk address are identified as risk policies, and the agents associated with each risk address are identified as risk agents.

[0014] In one possible implementation, the semantic feature vector of each target address is determined based on the text information of each target address, including:

[0015] The text information of each target address is input into the semantic matching dual-tower model to obtain the semantic feature vector of each target address.

[0016] In one possible implementation, a semantic matching dual-tower model is generated, including:

[0017] Obtain text information from multiple addresses and semantic feature vectors for each of the multiple addresses;

[0018] Based on the text information of multiple addresses and the semantic feature vector of each address, the original semantic matching dual-tower model is trained to obtain the semantic matching dual-tower model.

[0019] In one possible implementation, based on the semantic feature vector of each target address and the association relationships between each address and other addresses besides each address, the semantic feature vectors of N candidate addresses associated with each target address are determined, including:

[0020] Retrieve the associations between each address in the entire set of addresses and all other addresses except for each address;

[0021] Based on the semantic feature vector of each target address and the association between each address and other addresses besides each address, determine the semantic feature vector of all candidate addresses associated with each target address;

[0022] The semantic feature vectors of all candidate addresses associated with each target address are sorted according to the distance between the semantic feature vector of each target address and the semantic feature vectors of all candidate addresses associated with each target address.

[0023] Among the semantic feature vectors of all candidate addresses associated with each target address, the semantic feature vectors of the N candidate addresses with the smallest distance associated with each target address are determined as the semantic feature vectors of the N candidate addresses associated with each target address.

[0024] In one possible implementation, based on the semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address, N similarities are obtained for each target address, including:

[0025] The semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address are respectively input into the semantic similarity model to obtain N feature values ​​for each target address. The feature values ​​are used to characterize the similarity between the target address and the candidate address.

[0026] The N feature values ​​of each target address are input into the activation function to obtain N similarities for each target address. The activation function is used to transform the feature values ​​into similarities within a preset range.

[0027] In one possible implementation, the risk address of each target address is determined based on N similarities, including:

[0028] Among the N similarities of each target address, the candidate addresses with similarity greater than the first threshold are determined as the first similar address of each target address;

[0029] Obtain the tag information of the first similar address for each target address. The tag information is used to indicate all agents and all policies associated with the first similar address.

[0030] Based on the tag information of the first similar address of each target address, the first similar address that is the same as the agent's identifier is determined as the second similar address of each target address, and the first similar address that is the same as the policy's identifier is determined as the third similar address of each target address;

[0031] Determine the similarity of the second similar address for each target address, and determine the similarity of the third similar address for each target address;

[0032] Among the similarity scores of the second and third similar addresses for each target address, the second and third similar addresses with similarity scores greater than a second threshold are identified as risk addresses for each target address.

[0033] In one possible implementation, the address of the target risk policy includes at least one of the following: the policyholder's document address, the policyholder's permanent address, the policyholder's login latitude and longitude address, the policyholder's login street address, the insured's document address, or the insured's permanent address;

[0034] The address of the target risk agent includes at least one of the following: document address, login latitude and longitude address, login street address, permanent address, or delivery contact address.

[0035] This application embodiment uses a technical means to determine the risk address of each target address by obtaining the semantic feature vector of each target address and the association relationship between each address and other addresses besides each target address, and calculating the N similarities of each target address. This solves the technical problem that existing solutions cannot identify all risk policies and risk agents, and achieves the technical effect of identifying all risk policies and all risk agents related to all risk addresses with higher identification efficiency.

[0036] Secondly, embodiments of this application provide an apparatus for identifying risk insurance business, which is used to perform the method in the first aspect or any possible implementation thereof. Specifically, the apparatus may include modules for performing the method for identifying risk insurance business in the first aspect or any possible implementation thereof.

[0037] Thirdly, embodiments of this application provide a server including a memory and a processor. The memory stores instructions; the processor executes the instructions stored in the memory, causing the server to perform the method for identifying risk insurance business in the first aspect or any possible implementation thereof.

[0038] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the method for identifying risk insurance business in the first aspect or any possible implementation thereof.

[0039] Fifthly, a computer program product containing instructions is provided, which, when executed on a server, cause the server to perform the method for identifying risk insurance business in the first aspect or any possible implementation thereof.

[0040] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0043] Figure 2 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0044] Figure 3 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0045] Figure 4 This is a structural block diagram of the semantic matching dual-tower model generated in a method for identifying risk insurance business provided in this application embodiment;

[0046] Figure 5 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0047] Figure 6 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0048] Figure 7 This is a flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application;

[0049] Figure 8 This is an application block diagram of a method for identifying risk insurance business provided in an embodiment of this application;

[0050] Figure 9 This is a schematic block diagram of a device for identifying risk insurance business provided in an embodiment of this application. Detailed Implementation

[0051] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0052] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0053] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0054] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0055] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0057] Please see Figure 1 , Figure 1 A flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application is shown. Figure 1 As shown, the method for identifying risk insurance business in this application embodiment may include:

[0058] S101. Obtain the target risk identifier, which is the identifier of the target risk policy or the identifier of the target risk agent.

[0059] Users enter the target risk identifier on the server's display interface.

[0060] The target risk identifier is used to characterize insurance business that has been identified as having risk. Risk insurance business refers to insurance business involving insurance fraud.

[0061] The target risk identifier is a labeling information related to the identified risk insurance business. For example, the target risk identifier can be set as a number, symbol, etc.

[0062] The target risk identifier may include: the identifier of the target risk policy and the identifier of the target risk agent.

[0063] The identifier for the target risk policy can be its policy number. For example, the policy number for a target risk policy could be: Policy No. 1, Policy No. 2, etc.

[0064] The identifier for the target risk agent can be their agent account. For example, the target risk agent's agent accounts could include: Agent Account 1, Agent Account 2, etc.

[0065] Thus, in response to user input, the server can obtain one or more target risk identifiers.

[0066] It should be understood that users in the embodiments of this application may include: risk control personnel, customer service personnel, insurance agents, etc.

[0067] S102. Query the text information of each target address related to the target risk identifier, where the target address is the address of the target risk policy or the address of the target risk agent.

[0068] After obtaining the target risk identifier, the server can perform a search query in the database based on the target risk identifier.

[0069] Thus, the server can query the text information of each target address associated with the target risk identifier.

[0070] The target address is the address corresponding to the already identified risk insurance business.

[0071] The target address may include: the address of the target risk policy or the address of the target risk agent.

[0072] The address of the target risk policy may include at least one of the following: the policyholder's document address, the policyholder's permanent address, the policyholder's login latitude and longitude address, the policyholder's login street address, the insured's document address, or the insured's permanent address.

[0073] The address of the target risk agent may include at least one of the following: document address, login latitude and longitude address, login street address, permanent address, or delivery contact address.

[0074] Therefore, the number of target risk identifiers and target addresses is not specifically limited in the embodiments of this application.

[0075] Based on a target risk identifier obtained, the server can query one or more target addresses related to the target risk identifier.

[0076] Alternatively, the server can query one or more target addresses that are related to multiple target risk identifiers based on the obtained target risk identifiers.

[0077] S103. Determine the semantic feature vector of each target address based on the text information of each target address.

[0078] The server extracts features from the text information of each target address, transforming the text information of each target address into a semantic feature vector for each target address.

[0079] Thus, the server can determine the semantic feature vector of each target address.

[0080] The semantic feature vector is used to represent the textual semantics expressed by the textual information of the address. The semantic feature vector can be a 768-dimensional vector.

[0081] S104. Based on the semantic feature vector of each target address and the association between each address and other addresses besides each address, determine the semantic feature vectors of N candidate addresses associated with each target address, where N is a positive integer.

[0082] The server can pre-establish the association between each address and other addresses, and store the pre-established association between each address and other addresses in the database.

[0083] After obtaining the semantic feature vector of each target address, the server can retrieve the association between each address and other addresses in the database.

[0084] In summary, the server can determine the semantic feature vectors of N candidate addresses associated with each target address based on the semantic feature vector of each target address and the association between each address and other addresses.

[0085] Candidate addresses are those that, after initial screening, contain characters or words that are identical to those in the text information of the target address.

[0086] S105. Based on the semantic feature vector of each target address and the semantic feature vectors of the N candidate addresses associated with each target address, obtain the N similarities for each target address.

[0087] The server can obtain the semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address.

[0088] When a server receives multiple target addresses, it can calculate the similarity between any target address and any candidate address by performing a similarity calculation on the semantic feature vector of the target address and the semantic feature vector of any candidate address associated with the target address.

[0089] It should be understood that the similarity calculation process for the other candidate addresses among the N candidate addresses associated with any target address is the same as the similarity calculation process for any candidate address. The calculation process for each target address among multiple target addresses is the same as that for any target address, and will not be elaborated here.

[0090] In summary, by performing the above calculations on the semantic feature vector of each of the N candidate addresses associated with each target address, the server can obtain the similarity between each target address and each of the N candidate addresses, thus obtaining N similarity scores.

[0091] The similarity score is used to characterize the probability of similarity between the candidate address and the target address.

[0092] S106. Based on the N similarities of each target address, determine the risk address of each target address.

[0093] After obtaining N similarities for each target address, the server can recalculate, sort, and filter the N similarities for each target address.

[0094] Therefore, based on the filtering results, the server can determine the risky address for each target address.

[0095] Among them, the risk address can be an address among the candidate addresses that has exactly the same characters or words as the target address in the text information.

[0096] S107. Based on the risk address of each target address, identify the insurance policy related to each risk address as a risk policy, and identify the agent associated with each risk address as a risk agent.

[0097] The server can obtain the risky address for each target address.

[0098] Based on the risk address of each target address, the server can obtain the insurance policy associated with each risk address and the agent associated with each risk address.

[0099] Thus, the server can identify the policy associated with each risk address as a risk policy and the agent associated with each risk address as a risk agent.

[0100] Among these, a risk insurance policy is an insurance policy corresponding to risk insurance business. A risk agent is an agent who handles risk insurance business.

[0101] The method for identifying risky insurance business provided in this application involves obtaining a target risk identifier. Based on the obtained target risk identifier, text information for each target address related to the target risk identifier can be retrieved. Based on the retrieved text information for each target address, a semantic feature vector for each target address can be determined. Based on the determined semantic feature vectors of each target address and the association relationships between each address and other addresses besides itself, semantic feature vectors for N candidate addresses associated with each target address can be determined. Based on the determined semantic feature vectors of each target address and the semantic feature vectors of the N candidate addresses associated with each target address, N similarities for each target address can be obtained. Based on the obtained N similarities for each target address, the risk address for each target address can be determined. Based on the determined risk addresses for each target address, the insurance policies associated with each risk address can be identified as risky insurance policies, and the agents associated with each risk address can be identified as risky agents. Therefore, all risk addresses related to the target risk identifier can be accurately identified, and all risky insurance policies and all risky agents associated with all risk addresses can be obtained, resulting in higher identification efficiency and more comprehensive identification results.

[0102] Based on the above description, an example is given to illustrate one implementation method for the server to determine the semantic feature vector of each target address.

[0103] In some embodiments, the server can input the text information of each target address into the semantic matching dual-tower model to obtain the semantic feature vector of each target address.

[0104] Below, in conjunction with Figures 2-4 For example, we can illustrate the specific implementation process of the server determining the semantic matching dual-tower model.

[0105] Please see Figure 2 , Figure 2 A flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application is shown. Figure 2 As shown, the method for identifying risk insurance business in this application embodiment may include:

[0106] S201. Obtain text information from multiple addresses and semantic feature vectors for each of the multiple addresses.

[0107] The server can retrieve text information from multiple addresses from a pre-built database.

[0108] The server can obtain the semantic feature vector of each address from multiple addresses in the database. The semantic feature vector of each address is obtained by the server performing vector transformation on the text information of the multiple addresses.

[0109] S202. Based on the text information of multiple addresses and the semantic feature vector of each address, the original semantic matching dual-tower model is trained to obtain the semantic matching dual-tower model.

[0110] The server can input the text information of multiple addresses, as well as the semantic feature vector of each address, into the original semantic matching dual-tower model to train the original semantic matching dual-tower model.

[0111] Among them, the original semantic matching dual-tower model can be the Baidu RocketQA model, which is a deep semantic retrieval model based on the pre-trained language representation model (bidirectional encoder representation from transformers, BERT) to achieve end-to-end question answering retrieval.

[0112] Below, in conjunction with Figure 3 This describes the process of generating the semantic matching dual-tower model. The specific implementation process of S301-S305 includes the following:

[0113] S301. Divide the obtained text information of multiple addresses and the semantic feature vector of each address into training sample data and validation sample data.

[0114] The server can divide the text information of multiple addresses and the semantic feature vector of each address to obtain training sample data and validation sample data.

[0115] The training sample data includes: text information from multiple training addresses and semantic feature vectors for each of the multiple training addresses.

[0116] The validation sample data includes: text information of multiple validation addresses and semantic feature vectors for each of the multiple validation addresses. The multiple validation addresses are those addresses other than the multiple training addresses. The semantic feature vector of each of the multiple validation addresses is a semantic feature vector derived from the semantic feature vectors of each of the multiple validation addresses, excluding the semantic feature vectors of each of the multiple training addresses.

[0117] S302. Input the text information of multiple training addresses in the training sample data into the original semantic matching dual-tower model for training, and obtain multiple semantic feature vectors output by the original semantic matching dual-tower model.

[0118] S303. When the multiple semantic feature vectors output by the original semantic matching dual-tower model are the same as the semantic feature vectors of each training address in the multiple training addresses, the trained semantic matching dual-tower model is obtained.

[0119] The server inputs the text information of multiple training addresses from the training sample data into the original semantic matching dual-tower model. After the original semantic matching dual-tower model performs vector transformation on the text information of the training addresses, multiple semantic feature vectors output by the original semantic matching dual-tower model can be obtained.

[0120] Furthermore, the server compares the multiple semantic feature vectors output by the original semantic matching dual-tower model with the semantic feature vectors of each of the multiple training addresses. When the server determines that the multiple semantic feature vectors output by the original semantic matching dual-tower model are identical to the semantic feature vectors of each of the multiple training addresses, it obtains the trained semantic matching dual-tower model.

[0121] like Figure 4 As shown, by inputting the training sample data into the original semantic matching dual-tower model for training, a trained semantic matching dual-tower model can be obtained.

[0122] Because the amount of training sample data is limited, the trained semantic matching dual-tower model may not be applicable to text information at all addresses other than the text information at multiple training addresses. Therefore, validation sample data is needed to validate the trained semantic matching dual-tower model. The validation process can be found in S304-S305.

[0123] S304. Input the text information of multiple verification addresses in the verification sample data into the trained semantic matching dual-tower model to obtain multiple semantic feature vectors output by the trained semantic matching dual-tower model.

[0124] S305. When the semantic feature vectors output by the trained semantic matching dual-tower model are the same as the semantic feature vectors of each of the multiple verification addresses, the semantic matching dual-tower model is obtained.

[0125] The server inputs the text information of multiple verification addresses in the verification sample data into the trained semantic matching dual-tower model. After the trained semantic matching dual-tower model performs vector transformation on the text information of multiple verification addresses, multiple semantic feature vectors output by the trained semantic matching dual-tower model can be obtained.

[0126] Furthermore, the server compares the multiple semantic feature vectors output by the trained semantic matching dual-tower model with the semantic feature vectors of each of the multiple verification addresses. The server obtains the semantic matching dual-tower model when it determines that the multiple semantic feature vectors output by the trained semantic matching dual-tower model are identical to the semantic feature vectors of each of the multiple verification addresses.

[0127] In summary, after training the original semantic matching dual-tower model, the server obtains a trained semantic matching dual-tower model. After validating the trained semantic matching dual-tower model, the server obtains the final semantic matching dual-tower model. Therefore, in S103, the server can input the text information of each target address into the semantic matching dual-tower model to obtain the semantic feature vector of each target address after transformation by the semantic matching dual-tower model.

[0128] In some embodiments, the semantic matching dual-tower model includes a data tower sub-model and a query tower model.

[0129] The data tower sub-model is used to transform the text information of each address in all addresses into a semantic feature vector for each address.

[0130] The query pyramid model is used to transform the text information of each target address in all target addresses into a semantic feature vector for each target address.

[0131] During the training phase, the data pyramid model and the query pyramid model are trained using the same training sample data.

[0132] During the application phase, the data pyramid sub-model and the query pyramid sub-model are used separately. The data pyramid sub-model in the semantic matching dual-pyramid model pre-converts the text information of each address in all addresses into a corresponding semantic feature vector for each address. This improves the conversion efficiency of the semantic matching dual-pyramid model.

[0133] Based on the above description, after the server uses the semantic matching dual-tower model to transform the text information of each target address into the semantic feature vector of each target address, it can further determine the semantic feature vectors of N candidate addresses associated with each target address.

[0134] Below, in conjunction with Figure 5 This illustrates a feasible implementation method in S104 for the server to determine the semantic feature vectors of N candidate addresses associated with each target address.

[0135] Please see Figure 5 , Figure 5 A flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application is shown. Figure 5As shown, the method for identifying risk insurance business in this application embodiment may include:

[0136] S401. Obtain the association between each address in all addresses and all other addresses except each address.

[0137] In some embodiments, the server establishes the association between each semantic feature vector and other semantic feature vectors besides each semantic feature vector using the hierarchical navigable small world (HNSW) algorithm, thereby obtaining a feature vector query index.

[0138] The feature vector query index is a data structure. It provides pointers to semantic feature vectors stored in specified columns of the database, allowing the server to quickly retrieve semantic feature vectors from the database and obtain the text information corresponding to the addresses of those vectors.

[0139] Since each address has only one corresponding semantic feature vector, the association between each semantic feature vector and all other semantic feature vectors is the same as the association between each address and all other addresses. Therefore, the feature vector lookup index can be used to represent the association between each address and all other addresses.

[0140] In summary, by querying the index based on the feature vector, the server can obtain the association between each address and all other addresses except for each address.

[0141] S402. Based on the semantic feature vector of each target address and the association between each address and other addresses besides each address, determine the semantic feature vector of all candidate addresses associated with each target address.

[0142] In some embodiments, after obtaining the semantic feature vector of each target address, the server may retrieve the feature vector query index from the database.

[0143] The server can determine the semantic feature vectors of all candidate addresses associated with each target address by using the semantic feature vector of each target address and the feature vector query index.

[0144] S403. Sort the semantic feature vectors of all candidate addresses associated with each target address according to the distance between the semantic feature vector of each target address and the semantic feature vectors of all candidate addresses associated with each target address.

[0145] The server can obtain the distance between the semantic feature vector of each target address and the semantic feature vector of all candidate addresses by calculating the distance between the semantic feature vector of each target address and the semantic feature vector of each candidate address.

[0146] The server sorts the semantic feature vectors of all candidate addresses associated with each target address based on the distance between the semantic feature vector of each target address and the semantic feature vectors of all candidate addresses. For example, the server sorts the semantic feature vectors of all candidate addresses associated with each target address in ascending order of the distance between them and the semantic feature vector of each target address.

[0147] In some embodiments, the server can indirectly determine the distance between the semantic feature vector of each target address and the semantic feature vector of all candidate addresses by using the cosine similarity calculation formula.

[0148] The smaller the cosine similarity between the semantic feature vector of the target address and the semantic feature vector of the candidate address, the greater the distance between them.

[0149] The greater the cosine similarity between the semantic feature vector of the target address and the semantic feature vector of the candidate address, the smaller the distance between them.

[0150] In addition, the server can also directly calculate the distance between the semantic feature vector of the target address and the semantic feature vector of the candidate address using the Euclidean distance calculation formula.

[0151] S404. Among the semantic feature vectors of all candidate addresses associated with each target address, the semantic feature vectors of the N candidate addresses with the smallest distance associated with each target address are determined as the semantic feature vectors of the N candidate addresses associated with each target address.

[0152] Based on the ranking of the semantic feature vectors of all candidate addresses associated with each target address, the server can determine the semantic feature vectors of the top N candidate addresses with the smallest distances associated with each target address. Therefore, the server can select the semantic feature vectors of the top N candidate addresses with the smallest distances associated with each target address from among all the semantic feature vectors of the candidate addresses associated with each target address as the semantic feature vectors of the N candidate addresses associated with each target address.

[0153] For example, if the target address is address A, the N candidate addresses with the smallest distance associated with address A are address 1, address 2, ..., address n, and the distance between the semantic feature vector of address 1 and the semantic feature vector of address A is less than the distance between the semantic feature vector of address 2 and the semantic feature vector of address A, and so on.

[0154] In some embodiments, the server can use the topN algorithm to select the semantic feature vectors of the N candidate addresses with the smallest distance associated with each target address, where N is a positive integer.

[0155] Based on the above description, the server uses a feature vector query index and a top-N algorithm to filter out the semantic feature vectors of the N candidate addresses with the smallest distance associated with each target address from the database. This reduces the filtering time from N to logN, thereby reducing the computational load, shortening the filtering time, and improving the filtering efficiency.

[0156] After obtaining the semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address, the server can further determine the N similarities of each target address.

[0157] In some embodiments, the server may input the semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address into the semantic similarity model to obtain N feature values ​​for each target address.

[0158] Among them, the feature value is used to characterize the similarity between the target address and the candidate address.

[0159] Therefore, the server can input the N feature values ​​of each target address into the activation function to obtain N similarities for each target address.

[0160] The activation function is used to transform feature values ​​into similarity within a preset range.

[0161] The formula for the semantic similarity model is F(x1, x2), where F represents the semantic similarity model, x1 represents the semantic feature vector of each target address, x2 represents the semantic feature vector of any one of the N candidate addresses associated with each target address, and F(x1, x2) represents a feature value. The feature value is a real number, which includes positive real numbers and negative real numbers.

[0162] The activation function can be the sigmoid function. Typically, the output of the sigmoid function ranges from [0, 1].

[0163] Since the output of the semantic similarity model is a real number, and the preset range of similarity is generally set to [0, 1], the server needs to use the sigmoid function to convert the feature values ​​into similarity within the preset range.

[0164] It should be understood that the training process of the semantic similarity model is similar in principle to that of the semantic matching dual-tower model. The difference lies in the use of different training and validation sample data. The training sample data of the semantic similarity model consists of the semantic feature vector of each target address among multiple target addresses, the semantic feature vector of each candidate address among multiple candidate addresses associated with each target address, and multiple feature values.

[0165] In summary, the server can further calculate the similarity between the N candidate addresses associated with each target address from the initially selected N candidate addresses using a trained semantic similarity model, thus obtaining the N similarities for each target address in S105. Furthermore, the server can precisely sort the N similarities for each target address according to their similarity magnitude. Therefore, the server can further filter the risk addresses associated with each target address from the N candidate addresses based on their similarity magnitude.

[0166] Below, in conjunction with Figure 6 This illustrates a feasible implementation method in S106 for the server to determine the risk address of each target address.

[0167] Please see Figure 6 , Figure 6 A flowchart illustrating a method for identifying risk insurance business provided in an embodiment of this application is shown. Figure 6 As shown, the method for identifying risk insurance business in this application embodiment may include:

[0168] S501. Among the N similarities of each target address, the candidate addresses with similarity greater than the first threshold are determined as the first similar address of each target address.

[0169] The server can obtain N similarities between each target address and N candidate addresses, as well as a first threshold.

[0170] The first threshold represents the probability of similarity between the target address and the candidate address. This first threshold is a coarse similarity threshold preset by the server. For example, the first threshold can be set to 0.87.

[0171] The server can filter out candidate addresses whose similarity to N candidate addresses is greater than a first threshold from N similarity scores between each target address and N candidate addresses.

[0172] Therefore, the server can determine the candidate addresses with a similarity greater than the first threshold as the first similar address for each target address.

[0173] S502. Obtain the tag information of the first similar address for each target address. The tag information is used to indicate all agents and all policies associated with the first similar address.

[0174] After determining the first similar address for each target address, the server can retrieve the tag information of the first similar address for each target address from the database.

[0175] The tag information may include: the agent's address, the agent's identifier, the policy's address, and the policy's identifier. The tag information is used to represent the aggregated text and identifier information of any given address.

[0176] In addition, tag information can be stored in a key-value structure, where the key represents the text information of any address and the value represents the identification information of any address.

[0177] For example, the label for address A could be {"Agent's Identification Address": ["Agent's Account Number"], "Policyholder's Address": ["Policy Number"]}. Here, both the agent's identification address and the policyholder's address are text information for address A. The agent's account number and the policy number are both identifying information for address A.

[0178] In cases where a single address has multiple corresponding agent accounts and policy numbers, the tag information for any address includes multiple agent accounts and policy numbers. For example, the tag for address A could be {"Agent ID Address": ["Agent Account 1", "Agent Account 2"], "Policyholder Address": ["Policy Number 1", "Policy Number 2"]}.

[0179] S503. Based on the tag information of the first similar address of each target address, the first similar address that is the same as the agent's identifier is determined as the second similar address of each target address, and the first similar address that is the same as the policy's identifier is determined as the third similar address of each target address.

[0180] When the first, second, and third similar addresses all include multiple addresses, the server aggregates addresses with the same agent identifier within the first similar addresses based on the tag information of the first similar addresses obtained for each target address, thus obtaining the second similar address for each target address. For example, the first similar addresses include: Address B: "Agent's ID Address", Address C: "Agent's Login Latitude and Longitude Address", Address D: "Policyholder's Address", and Address E: "Insured's Address".

[0181] The tag information for the first similar address includes: Address B: {“Agent’s ID address”: [“Agent’s account 1”]}, Address C: {“Agent’s login latitude and longitude address”: [“Agent’s account 1”]}, Address D: {“Policyholder’s address”: [“Policy number 1”]}, Address E: {“Insured’s address”: [“Policy number 1”]}.

[0182] For example, the second similar address includes: {"Agent Account 1": "Agent ID Address": {"addr": "Address B", "Similarity": 0.89}, "Agent Login Latitude and Longitude Address": {"addr": "Address C", "Similarity": 0.9}}.

[0183] The second similar address is an address with a similarity greater than the first threshold and with the same agent's identifier.

[0184] Therefore, the server can identify the address with the same agent identifier as the first similar address of each target address as the second similar address of each target address.

[0185] Based on the tag information of the first similar address of each target address, the server aggregates addresses with the same policy identifier in the first similar address to obtain the third similar address of each target address.

[0186] For example, the third similar address includes: {"Policy No. 1": "Insured's Address": {"addr": "Address D", "Similarity": 0.89}, "Insured's Address": {"addr": "Address E", "Similarity": 0.88}}.

[0187] The third similar address is an address with a similarity greater than the first threshold and with the same policy identifier.

[0188] Therefore, the server can identify the address with the same policy identifier as the first similar address of each target address as the third similar address of each target address.

[0189] It should be understood that the aforementioned first similar address, second similar address, and third similar address may include one or more, and the embodiments of this application do not specifically limit the number of first similar addresses, second similar addresses, and third similar addresses.

[0190] S504. Determine the similarity of the second similar address for each target address, and determine the similarity of the third similar address for each target address.

[0191] The server calculates the similarity of the second similar addresses for each target address by weighting the similarity of each similar address in the second similar address set.

[0192] For example, when the weight of the similarity of each similar address in the second similar address is the same, the similarity of the second similar address is {"Agent Account 1": {"Similarity of the second similar address": 0.89+0.9, "Address Information": {"Agent ID Address": {"addr": "Address B", "Similarity": 0.89}, "Agent Login Latitude and Longitude Address": {"addr": "Address C", "Similarity": 0.9}}}}.

[0193] The server calculates the similarity of the third similar address to each target address by weighting the similarity of each similar address in the third similar address.

[0194] For example, when the similarity weights of each similar address in the third similar address are the same, the similarity of the third similar address is {"Policy No. 1": {"Similarity of the third similar address": 0.89 + 0.88, "Address Information": {"Insured's Address": {"addr": "Address D", "Similarity": 0.89}, "Insured's Address": {"addr": "Address E", "Similarity": 0.88}}}}.

[0195] The similarity of the second similar address for each target address obtained by the server is 1.79, and the similarity of the third similar address for each target address is 1.77.

[0196] In addition, the server can also use calculation formulas such as sorting and weighting in the fusion formula to calculate the similarity of the second similar address of each target address, as well as the similarity of the third similar address of each target address.

[0197] S505. Among the similarity of the second similar address and the similarity of the third similar address of each target address, the second similar address and the third similar address with a similarity greater than the second threshold are identified as risk addresses of each target address.

[0198] Since the tag information of the first similar address may include multiple agent accounts and multiple policy numbers, with multiple agent accounts belonging to the agent dimension and multiple policy numbers belonging to the policy dimension, the server can obtain multiple similarities for the second similar addresses by calculating the similarity of the second similar addresses under each agent account. Similarly, the server can obtain multiple similarities for the third similar addresses by calculating the similarity of the third similar addresses under each policy number.

[0199] It should be understood that the second similar addresses under each agent account may be partially the same, completely identical, or completely different. The third similar addresses under each policy number may be partially the same, completely identical, or completely different.

[0200] The server filters out second and third similar addresses whose similarity is greater than a second threshold from multiple similarity values ​​of second and third similar addresses.

[0201] The second threshold is a precise similarity threshold preset by the server. For example, the second threshold can be set to 1.75.

[0202] In summary, the server can identify the second and third similar addresses with a similarity greater than the second threshold as risk addresses for each target address. For example, if the similarity of the second similar address is 1.79 and the similarity of the third similar address is 1.77, the risk addresses for each target address include: address B, address C, address D, and address E.

[0203] In addition, the server can sort the similarity of the second similar address and the third similar address of each target address in descending order of similarity, and identify one or more addresses with the highest similarity as risk addresses for each target address.

[0204] This ensures that the risk addresses identified by the server are among the multiple addresses most similar to the target address, resulting in higher accuracy. Furthermore, users only need to analyze the risk policies and agents associated with the risk addresses identified by the server, improving analysis efficiency.

[0205] In actual insurance business, there are many identical addresses between the agent's address and the policy address. Therefore, the server can create an address-dimensional profile table to aggregate the text information of the same address into the tag information of a single address.

[0206] Below, in conjunction with Figure 7 This section describes the specific implementation process of establishing the address dimension profile table in the embodiments of this application. For example... Figure 7 As shown, the method for identifying risk insurance business in this application embodiment may include:

[0207] S601. Collect text information from all addresses.

[0208] Users enter business data on the server.

[0209] The business data includes agent registration information and policy information. Agent registration information includes the agent's name, ID card information, permanent address, and account information. Policy information includes the policyholder's name, address, insured's name, address, and policy number.

[0210] The server responds to user input, receives business data, and extracts text information for all addresses from the business data. Thus, the server can collect text information from all addresses.

[0211] Since databases store data in tables or files, servers can store business data in the database in the form of business tables or business logs.

[0212] S602. Based on the text information of all collected addresses, determine the tag information for each address among all addresses.

[0213] During the process of collecting text information from all addresses, the server can aggregate text information from the same address scattered in the business data, along with multiple agent accounts and / or multiple policy numbers, to obtain tag information for each address among all addresses.

[0214] S603. Based on the text information of all addresses and the label information of each address in all addresses, establish an address dimension profile table.

[0215] The server stores the text information of all addresses, along with the tag information for each address, in the database in the form of an address dimension profile table.

[0216] The address dimension profile table is an address database used to store text information about all addresses.

[0217] In summary, the server can establish an address-dimensional profile table. By establishing this table, the server can avoid redundant calculations for the same addresses, reducing computational costs and improving computational speed and analysis efficiency.

[0218] Based on the above description, and in combination Figure 8 The following example illustrates the specific implementation process of the method for identifying risk insurance business in this application.

[0219] Assume the server includes: a semantic matching dual-tower model, a semantic similarity model, a feature vector query index built using the HNSW algorithm, and an address-dimensional profile table.

[0220] The server sets the first threshold to 0.87, the second threshold to 1.75, and N = 500 candidate addresses.

[0221] The identifier of the target agent entered by the user is agent account A.

[0222] Please see Figure 8 , Figure 8 This illustration shows an application block diagram of a method for identifying risk insurance business provided in an embodiment of this application. For example... Figure 8 As shown, the method for identifying risk insurance business in this application embodiment may include:

[0223] Step 11: The user enters agent account A.

[0224] The specific implementation method of step 11 can be found in the description of S101, and will not be repeated here.

[0225] Step 12: The server queries the text information of each target address related to the agent's account or policy number in the established address dimension profile table.

[0226] For details on the implementation of step 12, please refer to the description in S102. For the process of establishing the address dimension profile table, please refer to the descriptions in S601-S603. These details will not be repeated here.

[0227] Step 13: The server inputs the text information of each target address into the query tower sub-model of the trained semantic matching dual-tower model to obtain the semantic feature vector of each target address.

[0228] Step 13 uses the query tower sub-model in the semantic matching dual-tower model obtained in S201-S202 and S301-S305 to perform vector transformation on the text information of each target address.

[0229] By separating the data tower sub-model and the query tower sub-model in the semantic matching dual-tower model, the server can improve the conversion efficiency of the text information for each target address during the user query phase.

[0230] Step 14: Based on the semantic feature vector of each target address, the established feature vector query index, and the top500 algorithm, the server determines the semantic feature vectors of 500 candidate addresses associated with each target address from the semantic feature vectors of all addresses pre-converted by the data tower sub-model in the semantic matching dual-tower model.

[0231] The specific implementation method of step 14 can be found in the description of S401-S404, and will not be repeated here.

[0232] Step 15: The server inputs the semantic feature vector of each target address and the semantic feature vector of each of the 500 candidate addresses into the semantic similarity model to calculate the 500 similarities for each target address.

[0233] The specific implementation method of step 15 can be found in the relevant description of similarity calculation above, and will not be repeated here.

[0234] Step 16: Among the 500 similarities for each target address, the server determines the candidate addresses with a similarity greater than 0.87 as the first similar address for each target address, resulting in 100 first similar addresses.

[0235] Step 17: The server retrieves the tag information of the first similar address from the established address dimension profile table.

[0236] Step 18: The server calculates the similarity under the agent dimension and the similarity under the policy dimension based on the tag information and similarity of the first similar address.

[0237] Step 19: Based on the similarity scores of the agent and the policy, the server identifies the first similar address with a similarity score greater than 1.75 as the risk address for each target address, resulting in 5 risk addresses.

[0238] The specific implementation methods of steps 16-19 can be found in the descriptions of S501-S505, and will not be repeated here.

[0239] Step 20: The server identifies the policies associated with each risk address as risk policies and the agents associated with each risk address as risk agents. The risk policies associated with the five risk addresses are Policy No. 1 and Policy No. 2, and the risk agents are Agent Account 1, Agent Account 2 and Agent Account 3.

[0240] The specific implementation of step 20 can be found in the description of S107, and will not be repeated here.

[0241] Therefore, the server can identify the policy number 1, policy number 2, agent account 1, agent account 2, and agent account 3 of the fraud gang based on agent account A, and quickly mark and block the fraud gang containing policy number 1, policy number 2, agent account 1, agent account 2, and agent account 3, thereby improving the efficiency of risk insurance business analysis.

[0242] In summary, the method for identifying risk insurance business in this application embodiment can identify multiple policy numbers of a risk insurance policy and multiple agent accounts of a risk agent through a single policy number of a risk insurance policy or a single agent account of a risk agent, thereby improving the efficiency of analyzing, identifying, and intercepting risk insurance business.

[0243] Please see Figure 9 , Figure 9 This diagram illustrates a structural schematic block diagram of an apparatus for identifying risk insurance business, as provided in an embodiment of this application. Figure 9 As shown, the apparatus 700 for identifying risk insurance business in this application may include:

[0244] The acquisition module 701 is used to acquire the target risk identifier, which is either the identifier of the target risk policy or the identifier of the target risk agent.

[0245] The query module 702 is used to query the text information of each target address related to the target risk identifier, where the target address is the address of the target risk policy or the address of the target risk agent.

[0246] The processing module 703 is used to: determine the semantic feature vector of each target address based on the text information of each target address; determine the semantic feature vectors of N candidate addresses associated with each target address based on the semantic feature vector of each target address and the association relationship between each address and other addresses besides each target address, where N is a positive integer; obtain N similarities for each target address based on the semantic feature vector of each target address and the semantic feature vectors of the N candidate addresses associated with each target address; determine the risk address of each target address based on the N similarities of each target address; and determine the insurance policy related to each risk address as a risk policy and the agent associated with each risk address as a risk agent based on the risk address of each target address.

[0247] In some embodiments, the processing module 703 is specifically used to input the text information of each target address into the semantic matching dual-tower model to obtain the semantic feature vector of each target address.

[0248] In some embodiments, the processing module 703 is specifically used to generate a semantic matching dual-tower model, including: obtaining text information of multiple addresses and semantic feature vectors of each of the multiple addresses; training the original semantic matching dual-tower model based on the text information of the multiple addresses and the semantic feature vectors of each of the multiple addresses to obtain the semantic matching dual-tower model.

[0249] In some embodiments, the processing module 703 is specifically configured to: obtain the association relationship between each address and all other addresses in all addresses; determine the semantic feature vectors of all candidate addresses associated with each target address based on the semantic feature vector of each target address and the association relationship between each address and all other addresses; sort the semantic feature vectors of all candidate addresses associated with each target address based on the semantic feature vector of each target address and the distance between the semantic feature vectors of all candidate addresses associated with each target address; and determine the semantic feature vectors of the N candidate addresses with the smallest distances among the semantic feature vectors of all candidate addresses associated with each target address as the semantic feature vectors of the N candidate addresses associated with each target address.

[0250] In some embodiments, the processing module 703 is specifically used to input the semantic feature vector of each target address and the semantic feature vectors of N candidate addresses associated with each target address into a semantic similarity model to obtain N feature values ​​for each target address. The feature values ​​are used to characterize the similarity between the target address and the candidate addresses. The N feature values ​​of each target address are input into an activation function to obtain N similarity activation functions for each target address. These activation functions are used to convert the feature values ​​into similarity within a preset range.

[0251] In some embodiments, the processing module 703 is specifically configured to: determine candidate addresses with similarity greater than a first threshold as first similar addresses for each target address from N similarities of each target address; obtain tag information of the first similar addresses of each target address, the tag information being used to indicate all agents and all policies associated with the first similar address; based on the tag information of the first similar addresses of each target address, determine first similar addresses with the same identifier as an agent as second similar addresses for each target address, and determine first similar addresses with the same identifier as a policy as third similar addresses for each target address; determine the similarity of the second similar addresses of each target address, and determine the similarity of the third similar addresses of each target address; and, based on the similarity of the second similar addresses of each target address and the similarity of the third similar addresses of each target address, determine the second similar addresses and the third similar addresses with similarity greater than a second threshold as risk addresses for each target address.

[0252] In some embodiments, the address of the target risk policy includes at least one of the following: the policyholder's document address, the policyholder's permanent address, the policyholder's registered latitude and longitude address, the policyholder's registered street address, the insured's document address, or the insured's permanent address; the address of the target risk agent includes at least one of the following: document address, registered latitude and longitude address, registered street address, permanent address, or delivery contact address.

[0253] This application also provides a server, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the server performs the steps described in the various method embodiments above.

[0254] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0255] This application also provides a computer program product containing instructions that, when run on a server, cause the server to perform the steps described in the various method embodiments above.

[0256] This application also provides a chip, including: an interface circuit and a logic circuit. The interface circuit is used to receive signals from other chips outside the chip and transmit them to the logic circuit, or to send signals from the logic circuit to other chips outside the chip. The logic circuit is used to implement the method for identifying risk insurance business in the foregoing embodiments.

[0257] This application also provides a chip system applied to a server including a memory and sensors; the chip system includes a processor; when the processor executes computer instructions stored in the memory, the server executes the method for identifying risk insurance business in the preceding embodiments.

[0258] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0259] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0260] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0261] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0262] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0263] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0264] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0265] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method of identifying a risk insurance business, characterized by, The method comprises the following steps: obtaining a target risk identifier, the target risk identifier being an identifier of a target risk policy or an identifier of a target risk agent; querying text information of each target address related to the target risk identifier, the target address being an address of the target risk policy or an address of the target risk agent; determining a semantic feature vector of each target address according to the text information of each target address; determining semantic feature vectors of N candidate addresses associated with each target address according to the semantic feature vector of each target address and an association relationship between each address and other addresses except for each address, N being a positive integer; obtaining N similarities of each target address according to the semantic feature vector of each target address and the semantic feature vectors of the N candidate addresses associated with each target address; determining a risk address of each target address according to the N similarities of each target address; determining a risk policy related to each risk address and a risk agent associated with each risk address according to the risk address of each target address; the step of determining the risk address of each target address according to the N similarities of each target address comprises the following steps: in the N similarities of each target address, determining a first similar address of each target address from candidate addresses with a similarity greater than a first threshold value; the first threshold value is a pre-set rough similarity threshold value; obtaining label information of the first similar address of each target address, the label information being used to indicate all agents and all policies associated with the first similar address; determining a second similar address of each target address from the first similar address with an agent identifier identical to that of an agent and determining a third similar address of each target address from the first similar address with a policy identifier identical to that of a policy according to the label information of the first similar address of each target address; determining a similarity of the second similar address of each target address and a similarity of the third similar address of each target address; in the similarity of the second similar address of each target address and the similarity of the third similar address of each target address, determining a risk address of each target address from the second similar address and the third similar address with a similarity greater than a second threshold value.

2. The method of claim 1, wherein, the step of determining the semantic feature vector of each target address according to the text information of each target address comprises the following step: inputting the text information of each target address into a semantic matching double-tower model to obtain the semantic feature vector of each target address.

3. The method of claim 2, wherein, generating the semantic matching double-tower model comprises the following steps: obtaining text information of a plurality of addresses and a semantic feature vector of each address in the plurality of addresses; training an original semantic matching double-tower model according to the text information of the plurality of addresses and the semantic feature vector of each address in the plurality of addresses to obtain the semantic matching double-tower model.

4. The method of claim 1, wherein, the step of determining the semantic feature vectors of the N candidate addresses associated with each target address according to the semantic feature vector of each target address and the association relationship between each address and other addresses except for each address comprises the following steps: obtaining an association between each of all addresses and other addresses except the each of all addresses; determining semantic feature vectors of all candidate addresses associated with each of the target addresses according to the semantic feature vector of each of the target addresses and the association between each of the target addresses and other addresses except the each of the target addresses; sorting the semantic feature vectors of all candidate addresses associated with each of the target addresses according to distances between the semantic feature vector of each of the target addresses and the semantic feature vectors of all candidate addresses associated with the each of the target addresses; determining semantic feature vectors of N candidate addresses associated with each of the target addresses from the semantic feature vectors of all candidate addresses associated with the each of the target addresses, wherein the semantic feature vectors of the N candidate addresses associated with the each of the target addresses are determined as the semantic feature vectors of the N candidate addresses associated with the each of the target addresses.

5. The method of claim 1, wherein, the N similarities of each of the target addresses are obtained according to the semantic feature vector of each of the target addresses and the semantic feature vectors of the N candidate addresses associated with the each of the target addresses, including: the N feature values of each of the target addresses are obtained by inputting the semantic feature vector of each of the target addresses and the semantic feature vectors of the N candidate addresses associated with the each of the target addresses into a semantic similarity model, wherein the feature values are used to represent similarities between the target addresses and the candidate addresses; the N similarities of each of the target addresses are obtained by inputting the N feature values of each of the target addresses into an activation function, wherein the activation function is used to convert the feature values into the similarities within a preset range.

6. The method according to any one of claims 1 to 5, wherein, the address of the target risk policy includes at least one of a certificate address of a policyholder, a permanent address of the policyholder, a login latitude and longitude address of the policyholder, a login street address of the policyholder, a certificate address of a person insured, or a permanent address of the person insured. the address of the target risk agent includes at least one of a certificate address, a login latitude and longitude address, a login street address, a permanent address, or a consignee contact address.

7. An apparatus for identifying a risk insurance business, characterized by The computer program is executed by the processor to implement the method of any one of claims 1-6.

8. A server comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-6.

9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Credit fraud analysis method, device and equipment and computer readable storage medium

    CN109584041A