A method, device and equipment for detecting illegal arbitrage, and a storage medium
By integrating operational behavior data of enterprise users and enterprise characteristic data, and using neural network models to identify illegal arbitrage behavior, the problem of inaccurate identification in traditional systems has been solved, achieving more accurate detection of illegal arbitrage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PING AN PROPERTY INSURANCE CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-09
Smart Images

Figure CN122175639A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of risk monitoring, and in particular to a method, apparatus, equipment, and storage medium for detecting illegal arbitrage. Background Technology
[0002] In the operation of B2B customer service platforms (covering fields such as insurance and finance), platforms often launch promotional activities, lotteries, and coupon distributions to attract enterprise users and improve user stickiness and activity. However, while these marketing activities bring business growth, they also breed illegal arbitrage by enterprise users or individuals using improper means—such as registering sub-accounts in bulk, sharing devices / IPs, and abusing activity rules to occupy large amounts of platform resources and illegally obtain prizes or discounts. This not only seriously damages the platform's economic interests but also undermines the fair competition environment and reduces the participation experience of enterprise users who genuinely need services.
[0003] To address these issues, existing industry solutions largely rely on traditional rule-based anti-fraud systems, which filter and restrict user behavior by setting fixed conditions in advance. However, in B2B scenarios, enterprise users' behavioral patterns and risk factors are more complex, leading to inaccurate identification of illegal arbitrage by enterprise users under traditional rules. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for detecting illegal arbitrage, in order to solve the problem of inaccurate identification of illegal arbitrage by enterprise users in the prior art.
[0005] Firstly, a method for detecting illegal arbitrage is provided, including: Acquire operational behavior data of target enterprise users who need to be detected for illegal arbitrage; Obtain enterprise characteristic data of the target enterprise user's company; By integrating operational behavior data and enterprise characteristic data, integrated characteristic data is obtained. Based on the fused feature data, the system identifies whether target enterprise users are engaging in illegal arbitrage activities, and obtains the identification results.
[0006] Secondly, a device for detecting illegal arbitrage is provided, comprising: The first acquisition module is used to acquire the operational behavior data of the target enterprise users that need to be detected for illegal arbitrage. The second acquisition module is used to acquire enterprise characteristic data of the enterprise to which the target enterprise user belongs; The fusion module is used to merge operational behavior data and enterprise characteristic data to obtain fused characteristic data. The identification module is used to identify whether target enterprise users have engaged in illegal arbitrage activities based on the fused feature data, and to obtain the identification results.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described illegal arbitrage detection method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described illegal arbitrage detection method.
[0009] The aforementioned methods, devices, equipment, and storage media for detecting illegal arbitrage can be applied to the financial and insurance sectors. They simultaneously acquire "target enterprise user operational behavior data" (such as batch registration, IP sharing, and other arbitrage-related operations) and "enterprise characteristic data" (such as enterprise size, certification status, and compliance records), constructing a detection foundation from two dimensions to avoid identification bias caused by incomplete data. By fusing data from both dimensions, the correlation between the two types of data can be directly demonstrated. Using the fused data as the basis for identification solves the problem of missed detections caused by the inability to identify associated risks in existing technologies. Therefore, this invention solves the problem of inaccurate identification of illegal arbitrage by enterprise users in existing technologies. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for the illegal arbitrage detection method in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for detecting illegal arbitrage in one embodiment of the present invention; Figure 3 This is a schematic diagram of a device for detecting illegal arbitrage in one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The illegal arbitrage detection method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain the operational behavior data of the target enterprise user to be detected for illegal arbitrage, as well as the enterprise characteristic data of the enterprise to which the target enterprise user belongs, through the client; the operational behavior data and enterprise characteristic data are fused to obtain fused characteristic data; based on the fused characteristic data, the server identifies whether the target enterprise user has engaged in illegal arbitrage behavior, and obtains the identification result. In this invention, both "operational behavior data of the target enterprise user" (such as arbitrage-related operations such as batch registration and sharing of IPs) and "enterprise characteristic data of the enterprise to which the target enterprise belongs" (such as enterprise size, certification status, compliance records, etc.) are obtained simultaneously, constructing the detection foundation from two dimensions and avoiding identification bias caused by incomplete data; the fusion of data from the two dimensions can directly reflect the correlation between the two types of data, and the fused data is used as the identification basis, solving the problem of missed detection caused by the inability to identify associated risks in existing technologies. Therefore, this invention solves the problem of inaccurate identification of illegal arbitrage by enterprise users in existing technologies. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a separate server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] The illegal arbitrage detection method provided in this invention can be used in the financial and insurance sectors. Please refer to [link / reference]. Figure 2 As shown, Figure 2 A flowchart illustrating the illegal arbitrage detection method provided in this embodiment of the invention includes the following steps: S21: Obtain the operational behavior data of the target enterprise users who need to be detected for illegal arbitrage.
[0015] The operational behavior data of target enterprise users refers to dynamic operation records collected through the platform's data collection channels that are directly related to the target enterprise users' illegal arbitrage activities within the platform. Target enterprise users include all relevant operational entities of the enterprise, such as administrator users associated with the enterprise's main account, ordinary employee users, and operational users corresponding to sub-accounts.
[0016] In one example, if a financial platform needs to detect whether a target enterprise user (Enterprise E, associated with 3 administrator users and 8 sub-account users) is engaging in illegal arbitrage activities, the following operational behavior data is obtained: Account operation data: In the past 15 days, Enterprise E's associated user registered 6 sub-accounts, all of which completed real-name authentication (but the bound mobile phone numbers belong to the same person). The main account frequently changed the administrator permissions of the sub-accounts (4 times in 3 days). Activity participation data: All sub-accounts participated in the platform's "Enterprise Financial Management Interest Rate Increase Activity" within 2 hours of its launch, accumulating 12 interest rate increase coupons, and the collection time was concentrated in the same time period (10:00-10:30). Login trajectory data: The login IPs of the 6 sub-accounts all belong to the same IP range (192.168.XX.XX-192.168.XX.XX), and the login device fingerprints show that only 3 devices (2 computers and 1 mobile phone) alternately logged into all sub-accounts. Profit acquisition data: All interest rate increase coupons collected by the sub-accounts were used for the same batch of financial management orders (the order receiving account is Enterprise E's corporate account), the order submission interval was no more than 5 minutes, and the financial management amount was close to the upper limit of the interest rate increase coupon usage threshold.
[0017] S22: Obtain enterprise characteristic data of the enterprise to which the target enterprise user belongs.
[0018] Enterprise characteristic data of the target enterprise user refers to the core data such as static attributes, qualification information, and historical business status of the enterprise to which the target enterprise user belongs, collected through legal collection channels such as platform registration information, compliance filing data, and business history records.
[0019] In one example, if an insurance B-end platform needs to detect whether a target enterprise user (Enterprise F, associated with 5 sub-accounts) is engaging in illegal arbitrage activities, the following enterprise characteristic data is obtained: Basic enterprise attribute data: Enterprise F is a micro-enterprise, established only 22 days ago, with a registered capital of 500,000 yuan, 3 registered employees, and a business scope of "enterprise management consulting," lacking insurance-related business qualifications; Platform qualification certification data: Only completed basic platform registration certification, without submitting advanced certification materials such as corporate bank account statements and proof of business premises, and failed the platform's "enterprise insurance qualification review"; Historical business data: No formal enterprise insurance purchase or policy renewal records in the past 30 days, only bulk purchase of "employee accident insurance" for associated individuals and application for the platform's "new enterprise insurance discount subsidy," with an enterprise account balance of only 3,000 yuan (not matching the total discount subsidy of 12,000 yuan applied for); Compliance record data: No historical compliance records related to insurance business, belonging to the platform's newly added low-credit-rating enterprises, and not included in the platform's enterprise trustworthy insurance list.
[0020] In one embodiment, acquiring operational behavior data and enterprise characteristic data includes: acquiring a preset window length and a preset sliding step length corresponding to each type of data in the operational behavior data and enterprise characteristic data; and acquiring the corresponding data according to the preset window length and preset sliding step length corresponding to each type of data.
[0021] This embodiment refers to first defining the "time statistical range (window length)" and "data update interval (sliding step size)" of different types of data before collecting two types of data. Then, effective data within a specific time range is accurately extracted according to the preset window length and preset sliding step size. This avoids indiscriminate collection of all historical data and ensures that the acquired data provides an accurate and efficient basis for subsequent fusion and recognition.
[0022] The preset window length refers to the "time span" for collecting a certain type of data, that is, only data within this time period (such as 7 days or 30 days) is extracted, and invalid data that is too long is filtered out; the preset sliding step size refers to the "update interval" for data collection, that is, how often data is collected again according to the window length (such as 1 day or 3 days), to ensure the real-time nature and dynamic updates of the data.
[0023] It should be noted that since the arbitrage activities are mostly short-term and concentrated, the window length is usually set to short-term (e.g., 7 days, 15 days) and the sliding step is set to high-frequency (e.g., 1 day) to ensure that recent operational data is captured. Since enterprise attributes (e.g., certification status, size) are not easily changed, the window length can be set to long-term (e.g., 90 days, 180 days) and the sliding step can be set to low-frequency (e.g., 7 days, 30 days) to avoid repeatedly collecting stable data.
[0024] In one example, an insurance B-end platform detects whether company G (linked to 6 sub-accounts) is engaging in illegal arbitrage activities such as "bulk insurance purchases to exploit subsidies." The data obtained according to these steps is as follows: Step 1: Obtain the preset window length and sliding step size for each type of data. Based on the characteristics of violations in the financial and insurance sectors, the platform presets the following parameters: Step 2: Obtain the corresponding data based on the preset parameters.
[0025] Collect operational behavior data, and extract the sub-account registration data (3 newly registered sub-accounts) and subsidy receipt data (each of the 3 sub-accounts received 2 insurance subsidy vouchers) of Enterprise G in the past 7 days according to the "7-day window length + 1-day sliding step"; and extract the login IP / device records of all sub-accounts in the past 15 days according to the "15-day window length + 1-day sliding step" (concentrated in 2 IP segments and sharing 4 devices).
[0026] Collect enterprise characteristic data, and extract the certification status (basic certification only, not passed the advanced qualification review) and business scope (no insurance-related qualifications) of enterprise G in the past 90 days by "90-day window length + 30-day sliding step". Extract the account balance (5,000 yuan) and historical insurance records of enterprise G in the past 30 days by "30-day window length + 7-day sliding step".
[0027] S23: Integrate operational behavior data and enterprise characteristic data to obtain integrated characteristic data.
[0028] This step involves organically integrating the operational behavior data of target enterprise users acquired in the early stages with the enterprise characteristic data of their respective enterprises, according to preset fusion rules (such as feature vector concatenation, weighted association of key information, entity graph construction, etc.), to form fused feature data that simultaneously contains "operational behavior characteristics + enterprise attribute characteristics". Fusion is not a simple data overlay, but rather a relational integration, focusing on "which operational behavior data combined with which enterprise characteristic data would constitute a risk of illegal arbitrage," such as the association between "uncertified enterprises" and "bulk registration of sub-accounts," or the association between "small-scale enterprises" and "receiving large discounts," embedding this relational information into the fused feature data.
[0029] In one embodiment, before fusing operational behavior data and enterprise characteristic data to obtain fused characteristic data, the process includes: calculating the weighted value of each type of data within a preset window length according to a preset time decay algorithm; and fusing the weighted data of each type to obtain fused characteristic data.
[0030] This embodiment refers to first assigning weighted values (higher weight for recent data and lower weight for older data) to different types of operational behavior data (such as sub-account registration, subsidy receipt, and IP login) and different types of enterprise characteristic data (such as authentication status and account balance) within their respective preset window lengths, using a time decay algorithm. Finally, all weighted data are integrated into fused feature data.
[0031] In one example, an insurance platform detects whether company H is engaging in illegal arbitrage. It calculates the weighted value of each data type within a preset window length using a pre-defined time decay algorithm, and then merges the weighted data types to obtain the fused feature data. The specific steps are as follows: The first step is to calculate the time decay weight of a single data point within a preset window for any type of data. The calculation formula is as follows: ,in, This represents the time decay weight of the j-th data item within a preset window for the i-th data type. This indicates the preset attenuation coefficient. This represents the time interval between the occurrence time of the j-th data item in the preset window of the i-th data type and the current time.
[0032] The second step is to perform a weighted summation on each data point in the preset window for this type of data. The calculation formula is as follows: ,in, This represents the j-th data item in the preset window representing the i-th type of data. This indicates the number of data entries of the i-th type within the preset window. Indicates the first The overall weighted value of the data.
[0033] The third step is to merge the weighted values of various types of data in the operational behavior data and various types of data in the enterprise feature data according to the preset fusion rules to obtain the fused feature data.
[0034] S24: Based on the fused feature data, identify whether the target enterprise user has engaged in illegal arbitrage activities, and obtain the identification results.
[0035] This step involves taking the integrated feature data generated in the early stage, which contains both enterprise characteristic data and operational behavior data, and substituting it into a preset risk identification logic (such as rule matching or model calculation). By analyzing the correlation risks between the enterprise characteristic data and operational behavior data reflected in the data, the final determination can be made as to whether the target enterprise user has engaged in illegal arbitrage activities.
[0036] The aforementioned method for detecting illegal arbitrage can be applied to the financial and insurance sectors. It simultaneously acquires "target enterprise user operational behavior data" (such as bulk registration, IP sharing, and other arbitrage-related operations) and "enterprise characteristic data" (such as enterprise size, certification status, and compliance records), constructing a detection foundation from two dimensions to avoid identification bias caused by incomplete data. By fusing data from both dimensions, the correlation between the two types of data can be directly demonstrated. Using the fused data as the basis for identification solves the problem of missed detections caused by the inability to identify associated risks in existing technologies. Therefore, this invention solves the problem of inaccurate identification of illegal arbitrage by enterprise users in existing technologies.
[0037] In one embodiment, fusing operational behavior data and enterprise feature data to obtain fused feature data includes: extracting core related entities from operational behavior data and enterprise feature data; using each core related entity as a node and the relationship between entities as the connection edge between nodes; generating feature vectors for each node based on operational behavior data and enterprise feature data; calculating the weights of the connection edges based on operational behavior data, enterprise feature data, and a preset weight calculation rule; and fusing the feature vectors of each node and the weights of the connection edges according to a preset fusion rule to obtain fused feature data.
[0038] This embodiment integrates operational behavior data and enterprise characteristic data by constructing an entity association graph to obtain fused characteristic data. The core is to transform the two disparate types of data into a graph structure of "node-edge-feature vector," achieving deep data association and integration. The specific logic is as follows: First, core related entities directly associated with illegal arbitrage are extracted from the two types of data. These entities are then used as graph nodes, and the business or behavioral relationships between entities are used as edges between nodes. Next, a feature vector containing its own attributes is generated for each node. Simultaneously, the weight of each edge (reflecting the degree of association) is calculated according to a preset weighting rule. Finally, the node feature vectors and edge weights are integrated according to the fusion rules to form fused characteristic data that combines "entity attributes" and "associations."
[0039] Extracting core entities is to focus on key participants in illegal arbitrage and avoid interference from irrelevant data. Core entities typically include companies, accounts, login devices, IP addresses, etc.
[0040] Node feature vectors integrate the attribute information of an entity itself. For example, the feature vector of an "account" node includes operational attributes such as registration time, number of coupons claimed, and login frequency; the feature vector of an "enterprise" node includes attributes such as authentication status, enterprise size, and account balance.
[0041] Edge weights reflect the degree of connection between two entities. For example, the edge weight of "enterprise-account" can be calculated based on the number of account registrations and the frequency of operations, while the edge weight of "account-device" can be calculated based on the number of logins. The higher the weight, the closer the connection and the greater the suspicion of arbitrage.
[0042] In one example, the target of the detection is enterprise user H on an insurance B-end platform. Their operational behavior data includes "3 sub-accounts, 2 login devices, 1 IP address, and receiving 5 insurance subsidy coupons." Their enterprise characteristic data includes "no advanced authentication, micro-enterprise, and account balance of 1235 yuan." Step 1: Extract core related entities. From the operational behavior data and enterprise characteristic data, extract six types of core related entities: target enterprise H, sub-accounts A / B / C, login devices X / Y, login IP 192.168.XX.XX, subsidy coupons T1-T5, and enterprise accounts.
[0043] Step 2: Construct an entity relationship graph (nodes + related edges). Nodes are: target company H, sub-accounts A / B / C, devices X / Y, IP 192.168.XX.XX, subsidy coupons T1-T5, and company account; related edges are: company H - sub-accounts A / B / C (company registration sub-accounts), sub-accounts A / B / C - devices X / Y (sub-account login devices), sub-accounts A / B / C - IP 192.168.XX.XX (sub-account login IP), sub-accounts A / B / C - subsidy coupons T1-T5 (sub-accounts receiving subsidy coupons), and subsidy coupons T1-T5 - company account (sub-accounts collecting subsidy coupon redemption funds).
[0044] Step 3: Generate node feature vectors. Enterprise H node feature vector: [No advanced certification (1), micro-enterprise (1), account balance of RMB 1235, no historical insurance record (1)]; Sub-account A node feature vector: [Registration time D-1 day, login frequency 4 times, number of coupons received 2, belonging to enterprise H (1)]; Sub-account B / C node feature vector: [Registration time D-1 day, login frequency 3 times, number of coupons received 1-2, belonging to enterprise H (1)]; Device X / Y node feature vector: [Number of login sub-accounts 3, login time concentrated on D-1 day, belonging IP 192.168.XX.XX (1)]; Subsidy coupon T1-T5 node feature vector: [Receiving entity is sub-account A / B / C, redemption time D-1 day, funds are collected to enterprise account (1)].
[0045] Step 4: Calculate the weights of the associated edges (according to the preset weighting rule: the higher the frequency of association, the greater the weight). The preset weighting calculation rule is: Weight = Frequency of association / Total frequency × 1.0 (maximum value is 1.0), and the specific calculation is as follows: The edge weight of enterprise H-sub-accounts A / B / C is: 3 / 3 = 1.0 (number of sub-accounts). Edge weights of sub-accounts A / B / C and devices X / Y: All three sub-accounts log in to two devices, indicating high frequency of association, and their weights are 0.9. Edge weights for sub-accounts A / B / C with IP 192.168.XX.XX: All sub-accounts use the same IP, and the weight is 1.0; The edge weight of sub-accounts A / B / C for subsidy coupons T1-T5 is: coupon quantity 5 / 5 = 1.0; Side weight of subsidy vouchers T1-T5-enterprise accounts: All subsidy vouchers are aggregated into this account with a weight of 1.0.
[0046] Step 5: Generate fused feature data according to the preset fusion rules. The preset fusion rules are: Fusion feature data = concatenation of all node feature vectors + weighted sum of all edge weights. Node feature vector concatenation: Integrate all feature vectors of enterprise H, sub-accounts, devices, IPs, subsidy coupons, and enterprise accounts to form a vector set with unified dimensions; Weighted sum of edge weights: 1.0 (enterprise-sub-account) + 0.9 (sub-account-device) + 1.0 (sub-account-IP) + 1.0 (sub-account-sub-voucher) + 1.0 (sub-voucher-enterprise account) = 4.9; The final generated fused feature data contains both the attribute feature vectors of each entity and the comprehensive value of 4.9 of the associated edge weights, clearly presenting the complete arbitrage chain of "enterprise H using 3 sub-accounts, sharing the same device and IP, to receive subsidy coupons in batches and collect funds".
[0047] In one embodiment, identifying whether a target enterprise user is engaging in illegal arbitrage based on fused feature data includes: inputting the fused feature data into a preset neural network model to obtain a risk coefficient, and identifying whether the target enterprise user is engaging in illegal arbitrage based on the risk coefficient.
[0048] The neural network model needs to be trained in advance. The training data consists of historically labeled "fusion feature data of non-compliant enterprises" and "fusion feature data of normal enterprises." Through supervised learning, the model learns the "feature patterns corresponding to arbitrage behavior," such as the risk features corresponding to high-weighted association links of "enterprise-sub-account-equipment-subsidy voucher." In some examples, because the fusion feature data contains node feature vectors and edge weight information, graph neural network models or their variants are usually chosen. These models can directly process the graph structure data of "node-edge" and accurately capture the association risks between entities, rather than traditional fully connected neural networks.
[0049] In one embodiment, based on fused feature data, identifying whether a target enterprise user engages in illegal arbitrage behavior and obtaining an identification result includes: presetting at least one risk threshold to classify different risk levels, comparing the risk coefficient with each risk threshold, and if the risk coefficient is higher than a preset high-risk threshold, the identification result is that the target enterprise user engages in illegal arbitrage behavior; if the risk coefficient is lower than a preset low-risk threshold, the identification result is that the target enterprise user does not engage in illegal arbitrage behavior; if the risk coefficient is between the preset low-risk threshold and the high-risk threshold, the identification result is that the target enterprise user has a potential risk of illegal arbitrage.
[0050] The core logic of this embodiment is to pre-set at least two risk thresholds (low risk threshold and high risk threshold) to construct a three-level risk classification standard of "low risk - potential risk - high risk"; compare the risk coefficients calculated by the neural network model with the preset thresholds one by one, and output three clear identification results according to the threshold range in which the risk coefficient is located: "no illegal arbitrage behavior", "potential illegal arbitrage risk exists", and "illegal arbitrage behavior exists", so as to achieve refined and standardized risk judgment.
[0051] In one example, the risk coefficient of enterprise user H on the insurance B-end platform is calculated by a neural network model (Graph Neural Network GCN) to be 0.85. The preset risk thresholds are low risk threshold = 0.4 and high risk threshold = 0.7. The risk coefficient of enterprise user H is 0.85, which is higher than the preset high risk threshold of 0.7, and is in the "high risk range". Therefore, the output identification result is that the target enterprise user has engaged in illegal arbitrage.
[0052] In one embodiment, identifying whether a target enterprise user is engaging in illegal arbitrage based on fused feature data and obtaining an identification result includes: retrieving the corresponding risk threshold based on the fused feature data, identifying whether the target enterprise user is engaging in illegal arbitrage based on the risk threshold and risk coefficient, and obtaining an identification result.
[0053] This embodiment refers to a refined identification scheme that, based on previously generated fusion feature data, first matches and retrieves the specific risk threshold (rather than a uniform fixed threshold) corresponding to the fusion feature data, then compares the risk coefficient output by the neural network model with the retrieved specific threshold, and finally determines whether the target enterprise user has engaged in illegal arbitrage behavior.
[0054] Traditional fixed thresholds are difficult to adapt to the risk differences of different types of enterprises (e.g., the arbitrage behavior characteristics of micro and small enterprises differ from those of large enterprises). In this embodiment, the risk threshold is associated with fused feature data, meaning different fused feature data corresponds to different risk thresholds. The core dimensions in the fused feature data (such as enterprise size, certification level, type of operation behavior, and degree of entity association) are the key basis for matching risk thresholds. For example, the fused feature of "uncertified micro and small enterprises + short-term concentrated operations" corresponds to one set of risk thresholds; the fused feature of "certified large enterprises + regular operations" corresponds to another set of more lenient risk thresholds. This embodiment can cover the identification needs of different enterprise types and different arbitrage models, and is especially suitable for B-end platforms with complex business scenarios and large enterprise differences. It solves the "one-size-fits-all" defect of traditional fixed thresholds. By designing different risk thresholds corresponding to different fused feature data, the risk judgment is more closely aligned with the specific risk characteristics of enterprises, further improving the accuracy and scenario adaptability of illegal arbitrage identification.
[0055] In one example, insurance B-end platform enterprise user K has the core dimensions of integrated feature data as "no advanced authentication + micro and small enterprise + batch registration of 3 sub-accounts + association with the same IP / device + receiving 5 subsidy coupons"; insurance B-end platform enterprise user L has the core dimensions of integrated feature data as "advanced authentication + large enterprise + routine operation of 2 sub-accounts + dispersed IP / devices + receiving 2 subsidy coupons"; Preset threshold storage rules (categorized by feature dimension combination): Threshold group A (suitable for "uncertified + micro and small enterprises + batch operations + high correlation"): low risk threshold = 0.3, high risk threshold = 0.6; Threshold group B (suitable for "certified + large enterprises + routine operations + low correlation"): low risk threshold = 0.5, high risk threshold = 0.8; According to the neural network model, the risk coefficient of firm K is 0.7, and the risk coefficient of firm L is 0.6. The core dimensions of the fusion feature data of enterprise user K are extracted: no advanced authentication, micro and small enterprises, batch operations, and high correlation. Based on the core dimensions, the corresponding threshold group A is matched and retrieved (low-risk threshold = 0.3, high-risk threshold = 0.6). The risk coefficient of enterprise user K = 0.7 > high-risk threshold 0.6. Therefore, the output identification result is that the target enterprise user has illegal arbitrage behavior. Extract the core dimensions of the fusion feature data of enterprise user L: advanced authentication, large enterprise, routine operation, and low correlation. Match and retrieve the corresponding threshold group B (low risk threshold = 0.5, high risk threshold = 0.8) based on the core dimensions. The risk coefficient of enterprise user L = 0.6 is between the low risk threshold of 0.5 and the high risk threshold of 0.8. Therefore, the output identification result is that the target enterprise user has potential illegal arbitrage risk and needs to be manually reviewed and confirmed.
[0056] In one embodiment, an illegal arbitrage detection device is provided. For example... Figure 3 As shown, the illegal arbitrage detection device includes a first acquisition module 31, a second acquisition module 32, a fusion module 33, and an identification module 34. Detailed descriptions of each functional module are as follows: The first acquisition module 31 is used to acquire the operational behavior data of the target enterprise users that need to be detected for illegal arbitrage. The second acquisition module 32 is used to acquire enterprise characteristic data of the enterprise to which the target enterprise user belongs; Fusion module 33 is used to fuse operational behavior data and enterprise characteristic data to obtain fused characteristic data; The identification module 34 is used to identify whether the target enterprise user has engaged in illegal arbitrage activities based on the fused feature data, and to obtain the identification result.
[0057] This invention provides a device for detecting illegal arbitrage, applicable to the financial and insurance sectors. The device comprises: a first acquisition module for acquiring operational behavior data of target enterprise users to be detected for illegal arbitrage; a second acquisition module for acquiring enterprise characteristic data of the enterprise to which the target enterprise user belongs; a fusion module for fusing the operational behavior data and enterprise characteristic data to obtain fused characteristic data; and an identification module for identifying whether the target enterprise user has engaged in illegal arbitrage based on the fused characteristic data, thus obtaining the identification result. This invention simultaneously acquires "target enterprise user operational behavior data" (such as arbitrage-related operations like batch registration and shared IP addresses) and "enterprise characteristic data of the enterprise to which the target enterprise belongs" (such as enterprise size, certification status, and compliance records), constructing a detection foundation from two dimensions to avoid identification bias caused by incomplete data. Fusing data from both dimensions directly reflects the correlation between the two types of data, and using the fused data as the identification basis solves the problem of missed detections caused by the inability to identify associated risks in existing technologies. Therefore, this invention solves the problem of inaccurate identification of illegal arbitrage by enterprise users in existing technologies.
[0058] Specific limitations regarding the illegal arbitrage detection device can be found in the limitations of the illegal arbitrage detection method described above, and will not be repeated here. Each module in the aforementioned illegal arbitrage detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0059] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for detecting illegal arbitrage.
[0060] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a client-side arbitrage detection method.
[0061] In one embodiment, a computer device is provided for use in the financial and insurance fields, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire operational behavior data of target enterprise users who need to be detected for illegal arbitrage; Obtain enterprise characteristic data of the target enterprise user's company; By integrating operational behavior data and enterprise characteristic data, integrated characteristic data is obtained. Based on the fused feature data, the system identifies whether target enterprise users are engaging in illegal arbitrage activities, and obtains the identification results.
[0062] In one embodiment, a computer-readable storage medium is provided, applicable to the financial and insurance fields, on which a computer program is stored, which, when executed by a processor, performs the following steps: Acquire operational behavior data of target enterprise users who need to be detected for illegal arbitrage; Obtain enterprise characteristic data of the target enterprise user's company; By integrating operational behavior data and enterprise characteristic data, integrated characteristic data is obtained. Based on the fused feature data, the system identifies whether target enterprise users are engaging in illegal arbitrage activities, and obtains the identification results.
[0063] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions of the computer device and external computer device in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0064] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0066] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for detecting illegal arbitrage, characterized in that, include: Acquire operational behavior data of target enterprise users who need to be detected for illegal arbitrage; Obtain the enterprise characteristic data of the enterprise to which the target enterprise user belongs; By integrating the operational behavior data and enterprise characteristic data, fused characteristic data is obtained; Based on the fused feature data, the system identifies whether the target enterprise user has engaged in illegal arbitrage activities, and obtains the identification result.
2. The method for detecting illegal arbitrage according to claim 1, characterized in that, The process of acquiring the operational behavior data and enterprise characteristic data includes: Obtain the preset window length and preset sliding step size corresponding to each type of data in the operation behavior data and enterprise feature data; The corresponding data is obtained based on the preset window length and preset sliding step size corresponding to each type of data.
3. The method for detecting illegal arbitrage according to claim 2, characterized in that, Before fusing the operational behavior data and enterprise characteristic data to obtain the fused characteristic data, the process includes: The weighted value of each type of data within a preset window length is calculated according to a preset time decay algorithm; The weighted data of each type are then merged to obtain the merged feature data.
4. The method for detecting illegal arbitrage according to claim 1, characterized in that, The fusion of the operational behavior data and enterprise characteristic data yields fused characteristic data, including: Extract core related entities from the operational behavior data and enterprise characteristic data; Each core associated entity is taken as a node, and the relationship between entities is taken as the associated edge between nodes. Based on the operation behavior data and enterprise feature data, feature vectors of each node are generated. The weights of the associated edges are calculated based on the operation behavior data, enterprise feature data and preset weight calculation rules. The feature vectors of each node and the weights of the associated edges are fused according to preset fusion rules to obtain the fused feature data.
5. The method for detecting illegal arbitrage according to claim 1, characterized in that, The step of identifying whether the target enterprise user has engaged in illegal arbitrage activities based on the fused feature data includes: The fused feature data is input into a preset neural network model to obtain a risk coefficient, and the target enterprise user is identified as having engaged in illegal arbitrage activities based on the risk coefficient.
6. The method for detecting corporate arbitrage according to claim 5, characterized in that, The step of identifying whether the target enterprise user has engaged in illegal arbitrage activities based on the fused feature data, and obtaining the identification result, includes: At least one risk threshold is preset to classify different risk levels. The risk coefficient is compared with each risk threshold. If the risk coefficient is higher than the preset high-risk threshold, the identification result is that the target enterprise user has engaged in illegal arbitrage. If the risk coefficient is lower than the preset low-risk threshold, the identification result is that the target enterprise user has no illegal arbitrage behavior; If the risk coefficient is between the preset low-risk threshold and the high-risk threshold, the identification result is that the target enterprise user has a potential risk of illegal arbitrage.
7. The method for detecting corporate arbitrage according to claim 6, characterized in that, The step of identifying whether the target enterprise user has engaged in illegal arbitrage activities based on the fused feature data, and obtaining the identification result, includes: Based on the fused feature data, the corresponding risk threshold is retrieved, and based on the risk threshold and risk coefficient, it is determined whether the target enterprise user has engaged in illegal arbitrage behavior, thus obtaining the identification result.
8. A device for detecting illegal arbitrage, characterized in that, include: The first acquisition module is used to acquire the operational behavior data of the target enterprise users that need to be detected for illegal arbitrage. The second acquisition module is used to acquire enterprise characteristic data of the enterprise to which the target enterprise user belongs; The fusion module is used to fuse the operational behavior data and enterprise characteristic data to obtain fused characteristic data; The identification module is used to identify whether the target enterprise user has engaged in illegal arbitrage activities based on the fused feature data, and to obtain the identification result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the illegal arbitrage detection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the illegal arbitrage detection method as described in any one of claims 1 to 7.