Data processing method, device and equipment
By constructing an entity relationship diagram and propagating field data along the propagation path, the attribute information of the target entity is automatically generated, and the problem of insufficient richness and objectivity of attribute information in the prior art is solved, and efficient and accurate attribute information generation is achieved.
Patent Information
- Application Number
- CN202510089620.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
When mining attribute information from tables in relational databases or warehousing, the prior art has problems of insufficient richness and objectivity, and relies on manual configuration items, which are inefficient and prone to errors.
By constructing an entity relationship diagram, the field data in the first and second relational data tables are used as entity attribute information, and propagate along the propagation path, and the attribute information of the target entity is automatically generated.
The generation of multi-dimensional attribute information is realized, the richness and interpretability of attribute information is improved, manual intervention is reduced, and efficiency and accuracy are improved.
Smart Images

Figure CN119988382A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a data processing method, device and equipment. Background Art
[0002] In many business scenarios, the characteristics of a subject or entity may not only be the characteristics or attributes currently carried, but may also include many other hidden or undiscovered characteristics or attributes, especially for newly emerged subjects or entities, whose characteristics or attributes are more likely to be blank, which is not conducive to the discovery of the subject or entity, the execution of the business and the accuracy of business execution is also poor. For example, in compliance performance scenarios such as anti-illegal financial activities, risk effectiveness and new risk discovery capabilities are weak.
[0003] Therefore, as people pay more and more attention to their privacy data, the business side has an urgent need to improve the validity and perception of the attributes (or features) of the subject or entity. Currently, many valuable data are often stored in tables in relational databases or data warehouses. The usual attribute information (or feature) mining is generated based on a large number of manual configuration items. However, the attribute information generated by the above method is relatively poor in richness and objectivity, and it also relies on expert experience, which is inefficient and subjective. It is prone to errors and cannot be quickly tiled across multiple sites. Based on this, it is necessary to provide a better solution for attribute information mining based on relational data tables to generate multi-dimensional, interpretable and effective attribute information. Summary of the invention
[0004] The purpose of the embodiments of this specification is to provide a better solution for attribute information mining based on relational data tables to generate multi-dimensional attribute information that is highly interpretable and effective.
[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided in an embodiment of the present specification comprises: receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relationship data table. In response to the acquisition request, a second relationship data table related to the target entity is obtained. An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, wherein the edges are association relationships between different entities, and wherein the entities include entity attribute information, wherein the entity attribute information is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. The propagation paths for different auxiliary entities contained in the entity relationship graph to reach the target entity are determined, and the target attribute information of the target entity is generated based on the entity attribute information of the auxiliary entities in each propagation path.
[0006] A data processing device provided in an embodiment of the present specification comprises: a request module, receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relationship data table. A response module, in response to the acquisition request, obtaining a second relationship data table related to the target entity. An entity graph construction module, constructing an entity relationship graph based on the first relationship data table and the second relationship data table, wherein the entity relationship graph comprises entities and edges, wherein the entities comprise the target entity and auxiliary entities constructed based on the second relationship data table, wherein the edges are associations between different entities, and wherein the entities comprise entity attribute information, wherein the entity attribute information is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. An attribute generation module, determining propagation paths for different auxiliary entities contained in the entity relationship graph to reach the target entity, and generating target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
[0007] A data processing device provided in an embodiment of the present specification includes: a processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor is caused to: receive a request for obtaining target attribute information for a target entity, wherein the target entity is constructed based on a first relationship data table. In response to the acquisition request, a second relationship data table related to the target entity is acquired. An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, wherein the edges are association relationships between different entities, and wherein the entities include entity attribute information, wherein the entity attribute information is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. The propagation paths for different auxiliary entities contained in the entity relationship graph to reach the target entity are determined, and the target attribute information of the target entity is generated based on the entity attribute information of the auxiliary entities in each propagation path.
[0008] The embodiment of the present specification also provides a storage medium, the storage medium is used to store computer executable instructions, and the executable instructions implement the following process when executed by the processor: receiving a request to obtain target attribute information for a target entity, the target entity is constructed based on a first relationship data table. In response to the acquisition request, a second relationship data table related to the target entity is obtained. An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, the entity relationship graph includes entities and edges, the entity includes the target entity and an auxiliary entity constructed based on the second relationship data table, the edge is an association relationship between different entities, the entity includes entity attribute information, and the entity attribute information is constructed based on the fields contained in the relationship data table corresponding to the entity and / or the field data corresponding to the fields. Determine the propagation path of different auxiliary entities contained in the entity relationship graph to reach the target entity, and generate the target attribute information of the target entity based on the entity attribute information of the auxiliary entity in each propagation path.
[0009] The embodiment of the present specification also provides a computer program product, including a computer program, which implements the following process when executed by a processor: receiving a request to obtain target attribute information for a target entity, the target entity being constructed based on a first relationship data table. In response to the acquisition request, obtaining a second relationship data table related to the target entity. An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, the entity relationship graph including entities and edges, the entities including the target entity and auxiliary entities constructed based on the second relationship data table, the edges being associations between different entities, the entities including entity attribute information, the entity attribute information being constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. Determine the propagation paths for different auxiliary entities contained in the entity relationship graph to reach the target entity, and generate the target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is an embodiment of a data processing method of this specification; Figure 2 This is a schematic diagram of the structure of an entity relationship diagram of this specification; Figure 3 A schematic diagram of a propagation path in an entity relationship diagram of this specification; Figure 4 Another data processing method embodiment of this specification; Figure 5 A schematic diagram of a propagation path and a propagation operator in an entity relationship diagram of this specification; Figure 6 This is another data processing method embodiment of the present specification; Figure 7 This is another data processing method embodiment of the present specification; Figure 8 This is an embodiment of a data processing device of the present specification; Fig. 9 This is an embodiment of a data processing device in this specification. DETAILED DESCRIPTION
[0011] The embodiments of this specification provide a data processing method, device and equipment.
[0012] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0013] The embodiments of this specification provide a mechanism for automatically generating attribute information or features from a relational data table. In many business scenarios, the features of a subject or entity may not only be the features or attributes currently carried, but may also include many other hidden or undiscovered features or attributes. In particular, the features or attributes of newly emerged subjects or entities are more likely to be blank, which makes it difficult for the subject or entity to be discovered, and is not conducive to the execution of business. The accuracy of business execution is also poor. For example, in compliance performance scenarios such as anti-illegal financial activities, risk effectiveness and new risk discovery capabilities are weak. Therefore, the business side has an urgent need to improve the validity and perception ability of the attributes (or features) of the subject or entity. At present, many valuable data are often stored in tables of relational databases or data warehouses. The usual attribute (or feature) mining can be generated by algorithms or based on a large number of manual configuration items. Among them, for the method of generating by algorithm, it is necessary to use a specified algorithm to generate the corresponding attribute information in a violent manner. Although this method can ensure the richness of the attribute information, the interpretability of the attribute information is relatively poor and difficult to debug. Generally, distributed computing is not possible for massive data; for the method of generating the corresponding attribute information based on a large number of manual configuration items, although this method has strong interpretability, the richness and objectivity of the attribute information are relatively poor, and it relies on expert experience. Manual operation is time-consuming, inefficient, and has subjective biases, and is prone to errors, and cannot be quickly tiled between multiple sites. Based on this, it is necessary to provide an automated solution to automatically generate multi-dimensional attribute information that is highly interpretable and effective. The embodiments of this specification provide an achievable method, which constructs an entity relationship diagram through the association relationship between relational data tables. The field data in the relational data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of attribute information. At the same time, the definition and calculation of attribute information are decoupled, and the attribute information protocol is designed to define the calculation logic of the corresponding attribute information. Different attribute information calculation implementations share the same attribute information protocol. For specific processing, please refer to the specific content in the following embodiments.
[0014] like Figure 1 As shown, an embodiment of this specification provides a data processing method, and the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.), etc., wherein the server can be an independent server, or a server cluster composed of multiple servers, etc., and the server can be a background server for financial services or online shopping services, or a background server for an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server situation processing, which will not be repeated here. The method can specifically include the following steps: In step S102, a request for obtaining target attribute information of a target entity is received, and the target entity is constructed based on a first relational data table.
[0015] Among them, the target entity can be any entity, wherein the entity can be an object that can be distinguished in the real world, and the entity can be a specific or abstract representation of a person, place, thing, event or a concept. In this embodiment, the entity can correspond to a relational data table, that is, the target entity corresponds to the first relational data table. The relational data table may include multiple types, for example, a relational data table composed of different departments within a certain organization, or a relational data table of employee-related information contained in the above-mentioned organization, etc., which can be set according to actual conditions. The target attribute information can be the specified entity attribute information, wherein the entity attribute information can be used to represent the characteristics or features of the entity, and the entity attribute information can include multiple types. The entity attribute information can be the field data of a field in the relational data table. For example, if the entity is an employee, the entity attribute information may include employee ID, position, contact information, email address, department, etc. The relational data table may include one or more different fields, and the field data corresponding to each field. Specifically, as shown in Table 1 below, the relational data table of "employee" may include fields such as employee ID, position, contact information, email address, department, etc. Table 1
[0016] Based on Table 1, the field data corresponding to the same field of different employees can be the same or different. For example, the employee ID of one employee is 0101, and the employee ID of another employee is 0011. The position of one employee is technical research and development, and the position of another employee can also be technical research and development. The department to which one employee belongs is the research and development department, and the department to which another employee belongs is the security department, etc. The specific settings can be made according to the actual situation.
[0017] In implementation, when a relational data table (i.e., the first relational data table) lacks field data corresponding to a specified field, or when no field or corresponding field data exists in the first relational data table, a target entity can be constructed based on the first relational data table. For example, a corresponding identifier can be set for the first relational data table, and the target entity can be represented by the identifier. In addition, relevant information of the target attribute information to be generated can be set. For example, if the total amount of online payments made by a user in the last month needs to be determined, the attribute item (or field) corresponding to the target attribute information can be "the total amount of online payments made in the last month", and the total amount of online payments made in the last month can be set as relevant information of the target attribute information. A request for obtaining target attribute information for the target entity can be generated based on the identifier of the first relational data table, relevant information of the target attribute information, etc., and the server can receive the request for obtaining target attribute information.
[0018] In actual applications, the above-mentioned acquisition request can be executed after passive triggering, or it can be executed through active triggering. For example, when a new entity or relationship data table appears during the operation (or execution) of a certain business or system, at this time, a corresponding identifier can be set for the first relationship data table, and the target entity can be represented by the identifier, and the relevant information of the target attribute information to be generated (such as the characteristics of a certain risk, etc.) can be set. An acquisition request for the target attribute information of the target entity can be generated based on the identifier of the first relationship data table, the relevant information of the target attribute information, etc., and the acquisition request can be provided to the server, and the server can receive the acquisition request, which can also be set according to actual conditions.
[0019] In step S104, in response to the above acquisition request, a second relationship data table related to the target entity is acquired.
[0020] The second relational data table may be different from the first relational data table. For example, the first relational data table may be a relational data table of employee-related information contained in a certain organization, and the second relational data table may be a relational data table composed of different departments within the organization, etc., which may be set specifically according to actual conditions. The second relational data table may also include one or more different fields, and field data corresponding to each field. Some fields in the second relational data table may be the same as fields in the second relational data table, and fields in the second relational data table may be completely the same as fields in the second relational data table, which may be set specifically according to actual conditions.
[0021] In implementation, after receiving the above acquisition request, a relationship data table that has a direct specified association relationship with the target entity or an indirect association with the target entity can be acquired from a specified database in response to the acquisition request, and the acquired relationship data table can be used as the second relationship data table.
[0022] In actual applications, the business system can record relevant data involved in the process of different users performing designated businesses (such as online payment, etc.), and can build a relational data table based on the recorded data. After receiving the above acquisition request, the relationship data table related to the target entity can be obtained from the relationship data table built based on the recorded data in the above business system in response to the acquisition request, and the obtained relationship data table can be used as the second relationship data table.
[0023] In step S106, an entity relationship graph is constructed based on the first relationship data table and the second relationship data table. The entity relationship graph includes entities and edges. The entity includes a target entity and an auxiliary entity constructed based on the second relationship data table. The edge is an association relationship between different entities. The entity includes entity attribute information. The entity attribute information is constructed based on the fields contained in the relationship data table corresponding to the entity and / or the field data corresponding to the field.
[0024] In implementation, the relationship data table can be used as an entity, and each entity can be used as a node (or vertex), and the association relationship between different relationship data tables can be used as an edge to construct an entity relationship graph. Based on this, the entity relationship graph can be composed of entities as nodes and edges. After obtaining the first relationship data table and the related second relationship data table in the above manner, the first relationship data table is the target entity, and the second relationship data table can be an auxiliary entity, which is used to assist the target entity in generating corresponding attribute information. There can be an association relationship between the target entity and different auxiliary entities, and there can also be an association relationship between different auxiliary entities, so that a corresponding entity relationship graph can be constructed, such as Figure 2 As shown, the entity relationship diagram includes entities and edges, wherein the entities include entity A, entity B, entity C, entity D and entity F, and different entities are connected by edges. Figure 2 It includes 5 edges, namely edge R1, edge R2, edge R3, edge R4 and edge R5. In addition, the entities in the above entity relationship diagram include target entities and auxiliary entities. For example, entity B can be the target entity, and the remaining entity A, entity C, entity D and entity F are all auxiliary entities, or entity B and entity D can be the target entities, and the remaining entity A, entity C and entity F are all auxiliary entities, etc. The specific settings can be made according to actual conditions. In addition, each entity can include entity attribute information, and the entity attribute information can be constructed based on the fields contained in the relationship data table corresponding to the entity and / or the field data corresponding to the field. For example, the entity attribute information of entity C can be constructed by the fields contained in the relationship data table corresponding to entity C and / or the field data corresponding to the field, etc. In addition, since the main purpose of constructing an entity relationship diagram is to determine the attribute information of a certain entity, the various entity attribute information contained in the entity can be set to the corresponding entity in the entity relationship diagram, such as Figure 2As shown, entity A contains 3 attribute information ( Figure 2 In order to simplify the representation, only one attribute information is represented by a solid circle with "attribute". Entity C contains three attribute information, and entity F contains three attribute information. The entity attribute information of different entities can be different ( Figure 2 In order to simplify the representation, different attribute information is not distinguished and all are represented by the same solid circle with "attribute". Through the above entity relationship diagram, the relationship between different entities will become very intuitive, which is convenient for extracting relevant information and performing related calculations.
[0025] It should be noted that the entity attribute information of the auxiliary entity may be all the attribute information of the auxiliary entity, or it may be the attribute information only related to the target attribute information. For example, if the target attribute information is "the total amount of online payments in the past month", the entity attribute information of the auxiliary entity may include attribute information related to "the amount of online payments in the past month", "the amount of online payments", "online transactions", etc. The specific setting can be based on actual conditions, and the embodiments of this specification do not limit this.
[0026] In step S108, the propagation paths from different auxiliary entities included in the entity relationship graph to the target entity are determined, and the target attribute information of the target entity is generated based on the entity attribute information of the auxiliary entities in each propagation path.
[0027] In implementation, after obtaining the entity relationship graph in the above manner, the paths from different auxiliary entities to the target entity can be calculated through the entity relationship graph. Figure 2 Take the entity relationship diagram as an example. If the target entity is entity B, you can use Figure 2 , the calculated paths from different auxiliary entities to the target entity include 6, such as Figure 3 As shown, starting from entity A, it reaches entity B through edge R1; starting from entity C, it reaches entity B through edge R2; starting from entity D, it reaches entity B through edge R3; starting from entity D, it reaches entity B through edge R4; starting from entity F, it reaches entity B through edge R5, entity D, edge R3 in turn; starting from entity F, it reaches entity B through edge R5, entity D, edge R4 in turn. The path calculated above can be used as the propagation path from different auxiliary entities in the entity relationship graph to the target entity.
[0028] Since each propagation path can reach the target entity, it means that the auxiliary entities on each propagation path can assist the target entity in generating target attribute information, or, for generating target attribute information, the auxiliary entities on each propagation path are related thereto. Therefore, the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path. Specifically, a variety of different data processing rules (such as mapping relationships, maximum or minimum value extraction rules, etc.) and / or algorithms (such as weighted summation algorithm, addition, multiplication, etc.) can be pre-set according to actual conditions. For a certain propagation path, the entity attribute information of the auxiliary entities on the propagation path can be collected. Based on the collected entity attribute information and target attribute information, a data processing rule that can realize the conversion from entity attribute information to target attribute information can be selected from the above-mentioned data processing rules, and / or an algorithm that can realize the conversion from entity attribute information to target attribute information can be selected from the above-mentioned algorithms. For example, a certain propagation path can be calculated using a weighted summation algorithm, or a certain propagation path can be calculated using addition, or a certain propagation path can be processed using a maximum value extraction rule, etc. Based on the above-selected data processing rules and / or algorithms, the corresponding results can be calculated based on the entity attribute information of the corresponding auxiliary entity, so as to obtain the calculation result of the propagation path. Through the above method, the calculation result of each propagation path can be calculated, and the calculation results of each propagation path can be fused to obtain a fusion result, which can be used as the target attribute information of the target entity.
[0029] The embodiment of the present specification provides a data processing method, by receiving a request to obtain target attribute information for a target entity, responding to the request to obtain a second relationship data table related to the target entity, and then, an entity relationship graph can be constructed based on the first relationship data table and the second relationship data table corresponding to the target entity, the entity relationship graph including entities and edges, the entities including the target entity and auxiliary entities constructed based on the second relationship data table, the edges being association relationships between different entities, the entities including entity attribute information, the entity attribute information being constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields, and finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path, so that the entity relationship graph is constructed through the association relationship between the relationship data tables, and the relationship between the auxiliary entities is constructed based on the entity attribute information of the auxiliary entities. The field data in the relationship data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relationship data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and will not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability, and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0030] In practical applications, the specific processing methods for generating the target attribute information of the target entity based on the entity attribute information of the auxiliary entity in each propagation path in the above step S108 can be various. An optional processing method is provided below, such as Figure 4 As shown, the processing may specifically include the following steps S10802 to S10810.
[0031] In step S10802, based on the attribute items corresponding to the target attribute information of the target entity, the first attribute information is selected from the entity attribute information of the auxiliary entity in each propagation path.
[0032] In implementation, the attribute items corresponding to the target attribute information of the target entity can be analyzed, and the attribute information matching the target attribute information can be extracted from the entity attribute information of the auxiliary entity in each propagation path, and the acquired attribute information can be used as the first attribute information.
[0033] In step S10804, one or more different attribute conditions are constructed based on the selected different first attribute information.
[0034] In implementation, for example, the first attribute information includes information 1, information 2, and information 3 (wherein information 1, information 2, and information 3 are different types of information). Information 1 and information 2 can be combined into one attribute condition, or information 1 and information 3 can be combined into one attribute condition, or information 2 and information 3 can be combined into one attribute condition, or information 1, information 2, and information 3 can be combined into one attribute condition, etc., so that multiple different attribute conditions can be obtained. Specifically, for example, the first attribute information includes the payment amount being less than 10 yuan, the payment amount being between 10 and 50 yuan, and the payment amount being less than 10 yuan. If the amount is greater than 50 yuan, the payment time is between 5:00 and 9:00, and the payment time is between 9:00 and 12:00, the attribute conditions may include the payment amount is less than 10 yuan and the payment time is between 5:00 and 9:00; the payment amount is between 10-50 yuan and the payment time is between 5:00 and 9:00; the payment amount is greater than 50 yuan and the payment time is between 5:00 and 9:00; the payment amount is less than 10 yuan and the payment time is between 9:00 and 12:00; the payment amount is between 10-50 yuan and the payment time is between 9:00 and 12:00; the payment amount is greater than 50 yuan and the payment time is between 9:00 and 12:00.
[0035] It should be noted that when constructing attribute conditions, it can also be further determined in combination with relevant information of the target attribute information to be generated. For example, the target attribute information indicates that it is necessary to determine the proportion of transactions in which users pay an amount between 10 and 50 yuan between 5 a.m. and 12 p.m., then the attribute conditions may include a payment amount between 10 and 50 yuan and a payment time between 5 a.m. and 9 a.m.; a payment amount between 10 and 50 yuan and a payment time between 9 a.m. and 12 p.m. The specific conditions can also be set according to actual conditions, and the embodiments of this specification do not limit this.
[0036] In step S10806, a propagation operator of each propagation path is determined based on the first attribute information on each propagation path, the constructed attribute condition, and the attribute item corresponding to the target attribute information of the target entity.
[0037] In implementation, a variety of different propagation operators can be pre-set according to actual conditions. The propagation operator can be constructed through the above-mentioned data processing rules and / or algorithms. For example, a propagation operator can be constructed through a specified mapping relationship to map the data, or a propagation operator can be constructed through a maximum extraction rule to extract the maximum value in the data, or a propagation operator can be constructed through a weighted summation algorithm to calculate and sum the data, etc. After obtaining the first attribute information and the constructed attribute conditions on each propagation path in the above manner, the first attribute information and the constructed attribute conditions on each propagation path, as well as the attribute items corresponding to the target attribute information to be generated, etc., can be combined to select a suitable propagation operator from the above-mentioned propagation operators (i.e., an appropriate propagation operator that can generate information related to the target attribute information (or the attribute items corresponding to the target attribute information) based on the first attribute information and the constructed attribute conditions on the propagation path), such as Figure 5 As shown, the selected propagation operator can be set on the corresponding propagation path.
[0038] In step S10808, path attribute information of the above attribute items on each propagation path is generated based on the first attribute information on each propagation path, the constructed attribute conditions and the propagation operator.
[0039] In implementation, Figure 5 As shown, through the above processing, each propagation path contains the first attribute information, the constructed attribute conditions and the propagation operator, so that the corresponding result can be obtained by calculation based on the above information, that is, for any propagation path, the first attribute information and the constructed attribute conditions on the propagation path can be used as the path attribute information of the above attribute items on the propagation path. In the above manner, the path attribute information of the above attribute items on each propagation path can be calculated.
[0040] In step S10810, target attribute information of the target entity is constructed based on the path attribute information of the above attribute items on each propagation path.
[0041] In implementation, the path attribute information of the above attribute items on each propagation path may be fused, and finally fused information may be obtained, and the fused information may be used as target attribute information of the target entity.
[0042] It should be noted that the processing of the above steps S10802 to S10810 is implemented when the target attribute information is relatively simple or does not require complex (complexity exceeds a preset threshold) calculations to obtain the information. For example, the target attribute information may be information that can be obtained through a single hop. Specifically, if the attribute item corresponding to the target attribute information is the user's full-day transaction amount, then the corresponding target attribute information can be obtained by adding up the amounts of all transactions of the user within 1 day. In actual applications, the target attribute information may also be relatively complex or require very complex (complexity exceeds a preset threshold) calculations to obtain, that is, the target attribute information needs to go through multiple hops (such as 2 hops, 3 hops, etc.) to obtain. In the above step S108, the specific processing method of generating the target attribute information of the target entity based on the entity attribute information of the auxiliary entity in each propagation path can also be implemented in the following way, such as Figure 6 As shown, the processing may specifically include the following steps S10812 to S10818.
[0043] In step S10812, the target attribute information is disassembled to obtain a reasoning chain for constructing the target attribute information, wherein the reasoning chain includes a plurality of sub-attribute information arranged in a reasoning order for constructing the target attribute information.
[0044] In implementation, the attribute items and other related information corresponding to the target attribute information to be generated can be analyzed in advance. If it is determined that the complexity of the target attribute information to be generated is higher than a preset threshold, the target attribute information can be disassembled to obtain a reasoning chain for constructing the target attribute information. Based on the above example, the target attribute information indicates (or the attribute item corresponding to the target attribute information is) that it is necessary to determine the proportion of transactions in which users pay an amount between 10 and 50 yuan between 5 a.m. and 12 p.m., then to determine the target attribute information, it is necessary to first calculate the amount of transactions in which users pay an amount between 10 and 50 yuan between 5 a.m. and 12 p.m., then calculate the total amount of the user's transactions, and finally calculate the ratio of the above-calculated amount to the total amount, so as to obtain the target attribute information. In this way, the target attribute information is determined. The target attribute information needs to go through 3 jumps to be obtained. The above target attribute information can be disassembled and the reasoning thinking chain for constructing the target attribute information can be "1. Calculate the amount of transactions between 10-50 yuan paid by users between 5:00 and 12:00; 2. Calculate the total amount of user transactions; 3. Calculate the ratio of the above calculated amount to the total amount". A sub-attribute information (or sub-attribute items corresponding to the sub-attribute information, i.e. "the amount of transactions between 10-50 yuan paid by users between 5:00 and 12:00", "the total amount of user transactions", "the ratio of the above calculated amount to the total amount") can be constructed based on the data of each step in the reasoning process provided in the reasoning thinking chain, and then multiple sub-attribute information arranged in the order of reasoning thinking can be obtained.
[0045] In step S10814, for the first sub-attribute information among the multiple sub-attribute information, path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path is generated based on the entity attribute information of the auxiliary entity in each propagation path, and the first sub-attribute information is constructed based on the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path.
[0046] The first sub-attribute information may be any one of the multiple sub-attribute information.
[0047] The specific processing of generating the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path based on the entity attribute information of the auxiliary entity in each propagation path, and constructing the first sub-attribute information based on the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path can refer to the relevant content in the aforementioned step S108. In addition, the specific processing of generating the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path based on the entity attribute information of the auxiliary entity in each propagation path, and constructing the first sub-attribute information based on the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path can also be implemented by the above-mentioned steps S10802 to S10810. Now, based on the sub-attribute items corresponding to the first sub-attribute information, the second attribute information is selected from the entity attribute information of the auxiliary entity in each propagation path; one or more different target attribute conditions are constructed based on the selected different second attribute information; based on the second attribute information on each propagation path, the constructed target attribute conditions and the sub-attribute items corresponding to the first sub-attribute information, the target propagation operator of each propagation path is determined; based on the second attribute information on each propagation path, the constructed target attribute conditions and the target propagation operator, the path attribute information of the above sub-attribute items on each propagation path is generated; based on the path attribute information of the above sub-attribute items on each propagation path, the first sub-attribute information is constructed. The specific processing process can be referred to the above-mentioned related content, which will not be repeated here.
[0048] In step S10816, each sub-attribute information in the plurality of sub-attribute information except the first sub-attribute information is constructed based on the construction method of the first sub-attribute information in the plurality of sub-attribute information.
[0049] In implementation, the same processing method as that of constructing the first sub-attribute information may be used to construct each sub-attribute information except the first sub-attribute information in the multiple sub-attribute information. For details, please refer to the relevant content of the above step S10814, which will not be repeated here.
[0050] In step S10818, based on the reasoning chain of each sub-attribute information in the constructed multiple sub-attribute information and the target attribute information, the target attribute information of the target entity is constructed.
[0051] It should be noted that after the processing of the above steps S10812 to S10816, the sub-attribute information obtained may already contain the target attribute information. At this time, the specified sub-attribute information can be directly selected as the target attribute information from the multiple sub-attribute information constructed above according to the reasoning chain of the target attribute information. If the sub-attribute information obtained after the processing of the above steps S10812 to S10816 does not contain the target attribute information, and the target attribute information needs to be further calculated based on multiple sub-attribute information, the target attribute information of the target entity can be further calculated based on the reasoning chain of each sub-attribute information and the target attribute information in the multiple sub-attribute information constructed, which can be set according to the actual situation, and this embodiment of the specification does not limit this.
[0052] In practical applications, the various attribute information or sub-attribute information generated above can be added to the corresponding entity according to actual conditions as the entity attribute information of the entity. For details, please refer to the following content: each sub-attribute information in the multiple sub-attribute information is added as the attribute information of the target entity to the entity attribute information of the target entity in the entity relationship diagram; in addition, the target attribute information of the target entity can also be added to the entity attribute information of the target entity in the entity relationship diagram.
[0053] In practical applications, the above-mentioned attribute information generation mechanism can be applied to the generation of risk characteristics of preset risks (such as fraud risk or illegal financial activity risk, etc.), and the details can be referred to the processing of the following steps A2 and A4.
[0054] In step A2, based on the target attribute information of the target entity, risk characteristic information of the preset risk corresponding to the target entity is determined.
[0055] In implementation, for example, the target attribute information generated may be the amount of money received and spent by the user's account within a certain period of time. The amount of money received and spent by the user's account within a certain period of time may be used to determine whether the user's account has received large amounts of funds. Then, all or most of the received funds may be dispersed into multiple different accounts, thereby obtaining risk characteristic information on whether the user is at risk of illegal financial activities, in order to determine whether the user is at risk of illegal financial activities.
[0056] In step A4, risk prevention and control processing is performed on the target entity based on the risk characteristic information of the preset risk corresponding to the target entity.
[0057] In implementation, if the target entity has a preset risk, risk prevention and control measures can be taken for the target entity, for example, freezing the account of the target entity or limiting the collection or payment authority of the account, etc., which can be set according to actual conditions. If the target entity does not have a preset risk, no measures are required.
[0058] In practical applications, in the case where the preset risk includes the risk of illegal financial activities, the specific processing methods of the above step S104 can be various. An optional processing method is provided below, such as Figure 7 As shown, the processing may specifically include the following steps S1042 to S1046.
[0059] In step S1042, in response to the above acquisition request, a set of relational data tables constructed based on transaction data of financial transaction businesses corresponding to the risk of illegal financial activities is acquired.
[0060] In step S1044, an optional relationship data table having a preset association relationship with the target entity is obtained from the set of relationship data tables.
[0061] In step S1046, a second relationship data table related to the target entity is determined based on the acquired optional relationship data table.
[0062] In practical applications, propagation operators include transfer operators and aggregation operators. The transfer operator is an operator used to map input data based on a preset mapping relationship, and the aggregation operator is an operator used to output results in a manner less than the amount of input data.
[0063] Among them, the transfer operator can make the input data and the output data have a one-to-one mapping relationship. For example, the transfer operator can be used to implement "extracting the domain name information of the email address". The aggregation operator can make multiple groups or multiple input data be output in a smaller number or a single number. For example, the aggregation operator can be used to implement summation calculation and maximum value extraction.
[0064] In practical applications, in a relational data table containing a large number of numerical values, in addition to field data, a lot of other related data can be extended based on the field data. The extended data is often more critical information. Based on this, after obtaining the first relational data table and the second relational data table, the two relational data tables can be analyzed to determine other key information therein. For details, please refer to the following content: Analyze the first relational data table and the second relational data table respectively to obtain the key information contained in the first relational data table and the second relational data table. The key information includes one or more of the fields, the field data corresponding to the fields, the association relationship between different relational data tables, the type of the field data, and the value range of the field data.
[0065] Based on the above analysis and processing of the relationship data table, the specific processing method of the above step S106 can be varied. An optional processing method is provided below, which may specifically include the following: constructing an entity relationship diagram based on the key information contained in the first relationship data table and the key information contained in the second relationship data table.
[0066] In implementation, it has been explained above when constructing the entity relationship diagram that the entity attribute information of the entity is constructed by one or more different fields and / or field data corresponding to each field. Based on the above processing, in addition to the fields and / or field data corresponding to each field, the entity attribute information may also include one or more items of the type of the field data and the value range of the field data. The specific setting can be based on actual conditions, and the embodiments of this specification do not limit this.
[0067] The embodiment of the present specification provides a data processing method, by receiving a request to obtain target attribute information for a target entity, responding to the request to obtain a second relationship data table related to the target entity, and then, an entity relationship graph can be constructed based on the first relationship data table and the second relationship data table corresponding to the target entity, the entity relationship graph including entities and edges, the entities including the target entity and auxiliary entities constructed based on the second relationship data table, the edges being association relationships between different entities, the entities including entity attribute information, the entity attribute information being constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields, and finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path, so that the entity relationship graph is constructed through the association relationship between the relationship data tables, and the relationship between the auxiliary entities is constructed based on the entity attribute information of the auxiliary entities. The field data in the relationship data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relationship data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and will not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability, and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0068] In addition, the definition and calculation of attribute information are decoupled, and an attribute information protocol is designed to define the calculation logic of the corresponding attribute information. Different attribute information calculation implementations share the same attribute information protocol.
[0069] The above is a data processing method provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a data processing device, such as Figure 8 shown.
[0070] The data processing device comprises: a request module 801, a response module 802, an entity graph construction module 803 and an attribute generation module 804, wherein: Request module 801, receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; A response module 802, in response to the acquisition request, acquires a second relationship data table related to the target entity; An entity graph construction module 803 constructs an entity relationship graph based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are associations between different entities. The entities include entity attribute information, and the entity attribute information is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; The attribute generation module 804 determines the propagation paths from different auxiliary entities included in the entity relationship diagram to the target entity, and generates target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path. In the embodiment of this specification, the attribute generation module 804 includes: A selection unit, based on the attribute item corresponding to the target attribute information of the target entity, selects the first attribute information from the entity attribute information of the auxiliary entity in each propagation path; A condition construction unit, constructing one or more different attribute conditions based on the selected different first attribute information; An operator determination unit determines a propagation operator for each propagation path based on the first attribute information on each propagation path, the constructed attribute condition, and the attribute item corresponding to the target attribute information of the target entity; An information generating unit, which generates path attribute information of the attribute item on each propagation path based on the first attribute information on each propagation path, the constructed attribute condition and the propagation operator; The first attribute construction unit constructs target attribute information of the target entity based on the path attribute information of the attribute item on each propagation path.
[0071] In the embodiment of this specification, the attribute generation module 804 includes: A disassembling unit disassembles the target attribute information to obtain a reasoning chain for constructing the target attribute information, wherein the reasoning chain includes a plurality of sub-attribute information arranged in a reasoning order for constructing the target attribute information; a first sub-attribute construction unit, for first sub-attribute information among the plurality of sub-attribute information, generating path attribute information of a sub-attribute item corresponding to the first sub-attribute information on each propagation path based on entity attribute information of an auxiliary entity in each propagation path, and constructing the first sub-attribute information based on the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path; a second sub-attribute constructing unit, constructing each sub-attribute information among the plurality of sub-attribute information except the first sub-attribute information based on the constructing manner of the first sub-attribute information among the plurality of sub-attribute information; The second attribute construction unit constructs the target attribute information of the target entity based on the reasoning chain of each sub-attribute information in the constructed multiple sub-attribute information and the target attribute information.
[0072] In the embodiment of this specification, the device further includes: A first adding module, adding each sub-attribute information of the plurality of sub-attribute information as attribute information of the target entity to the entity attribute information of the target entity in the entity relationship diagram; The second adding module adds the target attribute information of the target entity to the entity attribute information of the target entity in the entity relationship diagram.
[0073] In the embodiment of this specification, the device further includes: A risk characteristic determination module, which determines risk characteristic information of a preset risk corresponding to the target entity based on the target attribute information of the target entity; The risk prevention and control module performs risk prevention and control processing on the target entity based on risk feature information of preset risks corresponding to the target entity.
[0074] In the embodiment of this specification, the preset risk includes the risk of illegal financial activities, and the response module 804 includes: A response unit, in response to the acquisition request, acquires a set of relational data tables constructed based on transaction data of financial transaction businesses corresponding to the risk of illegal financial activities; An acquisition unit, which acquires an optional relationship data table having a preset association relationship with the target entity from the set of relationship data tables; A data table determining unit determines a second relationship data table related to the target entity based on the acquired optional relationship data table.
[0075] In the embodiment of this specification, the propagation operator includes a transfer operator and an aggregation operator. The transfer operator is an operator used to map the input data based on a preset mapping relationship, and the aggregation operator is an operator used to output the result in a manner less than the amount of input data.
[0076] In the embodiment of this specification, the device further includes: an analysis module, analyzing the first relational data table and the second relational data table respectively, to obtain key information contained in the first relational data table and the second relational data table, wherein the key information includes one or more of a field, field data corresponding to the field, an association relationship between different relational data tables, a type to which the field data belongs, and a value range of the field data; The entity graph construction module 803 constructs an entity relationship graph based on the key information contained in the first relationship data table and the key information contained in the second relationship data table.
[0077] An embodiment of the present specification provides a data processing device, which obtains a second relationship data table related to the target entity in response to a request for obtaining target attribute information of a target entity upon receiving the request, and then constructs an entity relationship graph based on the first relationship data table and the second relationship data table corresponding to the target entity, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. Finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path. In this way, the entity relationship graph is constructed through the association relationship between the relationship data tables, and the relationship between the auxiliary entities is constructed based on the entity attribute information of the auxiliary entities. The field data in the relationship data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relationship data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and will not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability, and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0078] In addition, the definition and calculation of attribute information are decoupled, and an attribute information protocol is designed to define the calculation logic of the corresponding attribute information. Different attribute information calculation implementations share the same attribute information protocol.
[0079] The above is a data processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Fig. 9 shown.
[0080] The data processing device may provide a terminal device or a server, etc. for the above-mentioned embodiments.
[0081] The data processing device may have relatively large differences due to different configurations or performances, and may include one or more processors 901 and memory 902, and the memory 902 may store one or more storage applications or data. Among them, the memory 902 may be a short-term storage or a persistent storage. The application stored in the memory 902 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the data processing device. Furthermore, the processor 901 may be configured to communicate with the memory 902 and execute a series of computer executable instructions in the memory 902 on the data processing device. The data processing device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input and output interfaces 905, and one or more keyboards 906.
[0082] Specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs are configured to be executed by one or more processors, including computer executable instructions for performing the following: Receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; In response to the acquisition request, acquiring a second relationship data table related to the target entity; An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; Determine the propagation paths for different auxiliary entities included in the entity relationship graph to reach the target entity, and generate target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
[0083] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0084] An embodiment of the present specification provides a data processing device, which obtains a second relationship data table related to the target entity in response to a request for obtaining target attribute information of a target entity upon receiving the request, and then constructs an entity relationship graph based on the first relationship data table and the second relationship data table corresponding to the target entity, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. Finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path. In this way, the entity relationship graph is constructed through the association relationship between the relationship data tables. The field data in the relationship data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relationship data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and will not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability, and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0085] Furthermore, based on the above Figures 1 to 7 In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; In response to the acquisition request, acquiring a second relationship data table related to the target entity; An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; Determine the propagation paths for different auxiliary entities included in the entity relationship graph to reach the target entity, and generate target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
[0086] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0087] An embodiment of the present specification provides a storage medium, which obtains a second relationship data table related to the target entity in response to a request for obtaining target attribute information of a target entity upon receiving the request, and then an entity relationship graph can be constructed based on the first relationship data table and the second relationship data table corresponding to the target entity, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. Finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path. In this way, the entity relationship graph is constructed through the association relationships between the relationship data tables, and the relationship The field data in the data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relational data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and does not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0088] Furthermore, based on the above Figures 1 to 7 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; In response to the acquisition request, acquiring a second relationship data table related to the target entity; An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; Determine the propagation paths for different auxiliary entities included in the entity relationship graph to reach the target entity, and generate target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
[0089] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0090] The embodiment of the present specification provides a computer program product, which obtains a second relationship data table related to the target entity in response to a request for obtaining target attribute information of a target entity upon receiving the request, and then constructs an entity relationship graph based on the first relationship data table and the second relationship data table corresponding to the target entity, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields. Finally, the propagation paths of different auxiliary entities contained in the entity relationship graph to reach the target entity can be determined, and the target attribute information of the target entity can be generated based on the entity attribute information of the auxiliary entities in each propagation path. In this way, the entity relationship graph is constructed through the association relationship between the relationship data tables, and the relationship between the auxiliary entities is constructed based on the entity attribute information of the auxiliary entities. The field data in the relationship data table is propagated along the propagation path in the entity relationship diagram as entity attribute information, and the attribute information of the target entity (such as the risk characteristics of a certain risk, etc.) is automatically generated through the attribute information on the propagation path, achieving a good balance between the richness and interpretability of the attribute information. In addition, only a small amount of meta-information such as the entity relationship of the relationship data table needs to be input, and there is no need to list a lot of templates or base classes for feature generation. The meta-information such as the entity relationship is relatively stable and will not change frequently. It is easy to get started and is convenient for daily operation and maintenance. Moreover, the attribute information (or features) is generated by broadcasting based on the entity relationship diagram, which can clearly understand the complete calculation process from the original data to the final attribute information or features. It has strong interpretability, and also provides some entity attribute information of auxiliary entities to control the complexity and quantity of the final attribute information generation, achieving a good balance between interpretability, richness, objectivity, etc.
[0091] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0093] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0094] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0095] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0096] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0097] The embodiments of this specification are described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0099] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0100] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0101] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0102] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0103] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0104] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0106] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0107] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.
Claims
1. A data processing method, the method comprising: Receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; In response to the acquisition request, acquiring a second relationship data table related to the target entity; An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; Determine the propagation paths for different auxiliary entities included in the entity relationship graph to reach the target entity, and generate target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
2. The method according to claim 1, wherein generating target attribute information of the target entity based on entity attribute information of the auxiliary entity in each propagation path comprises: Based on the attribute item corresponding to the target attribute information of the target entity, selecting first attribute information from the entity attribute information of the auxiliary entity in each propagation path; Constructing one or more different attribute conditions based on the selected different first attribute information; Determine a propagation operator for each propagation path based on the first attribute information on each propagation path, the constructed attribute condition, and the attribute item corresponding to the target attribute information of the target entity; Generate path attribute information of the attribute item on each propagation path based on the first attribute information on each propagation path, the constructed attribute condition and the propagation operator; The target attribute information of the target entity is constructed based on the path attribute information of the attribute items on each propagation path.
3. The method according to claim 2, wherein generating target attribute information of the target entity based on entity attribute information of the auxiliary entity in each propagation path comprises: Decomposing the target attribute information to obtain a reasoning chain for constructing the target attribute information, wherein the reasoning chain includes a plurality of sub-attribute information arranged in a reasoning order for constructing the target attribute information; For the first sub-attribute information among the multiple sub-attribute information, generate path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path based on the entity attribute information of the auxiliary entity in each propagation path, and construct the first sub-attribute information based on the path attribute information of the sub-attribute item corresponding to the first sub-attribute information on each propagation path; constructing each sub-attribute information among the plurality of sub-attribute information except the first sub-attribute information based on the construction method of the first sub-attribute information among the plurality of sub-attribute information; Based on the reasoning chain of each sub-attribute information in the constructed multiple sub-attribute information and the target attribute information, the target attribute information of the target entity is constructed.
4. The method according to claim 3, further comprising: adding each sub-attribute information of the plurality of sub-attribute information as attribute information of the target entity to the entity attribute information of the target entity in the entity relationship diagram; The target attribute information of the target entity is added to the entity attribute information of the target entity in the entity relationship diagram.
5. The method according to any one of claims 1 to 4, further comprising: Based on the target attribute information of the target entity, determining risk characteristic information of a preset risk corresponding to the target entity; Based on the risk characteristic information of the preset risk corresponding to the target entity, risk prevention and control processing is performed on the target entity.
6. The method according to claim 5, wherein the preset risk includes illegal financial activity risk, and the step of obtaining, in response to the acquisition request, a second relationship data table related to the target entity comprises: In response to the acquisition request, acquiring a set of relational data tables constructed based on transaction data of financial transaction businesses corresponding to the illegal financial activity risk; Acquire an optional relationship data table having a preset association relationship with the target entity from the set of relationship data tables; A second relationship data table related to the target entity is determined based on the acquired optional relationship data table.
7. According to the method of claim 5, the propagation operator includes a transfer operator and an aggregation operator, the transfer operator is an operator used to map the input data based on a preset mapping relationship, and the aggregation operator is an operator used to output the result in a manner less than the amount of input data.
8. The method according to claim 7, further comprising: Analyze the first relationship data table and the second relationship data table respectively to obtain key information contained in the first relationship data table and the second relationship data table, wherein the key information includes one or more of a field, field data corresponding to the field, an association relationship between different relationship data tables, a type to which the field data belongs, and a value range of the field data; The constructing an entity relationship diagram based on the first relationship data table and the second relationship data table includes: An entity relationship diagram is constructed based on the key information contained in the first relationship data table and the key information contained in the second relationship data table.
9. A data processing device, comprising: A request module receives a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; A response module, in response to the acquisition request, acquires a second relationship data table related to the target entity; an entity graph construction module, which constructs an entity relationship graph based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, and the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entity and / or field data corresponding to the fields; The attribute generation module determines the propagation paths from different auxiliary entities included in the entity relationship diagram to the target entity, and generates target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.
10. A data processing device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Receiving a request for obtaining target attribute information of a target entity, wherein the target entity is constructed based on a first relational data table; In response to the acquisition request, acquiring a second relationship data table related to the target entity; An entity relationship graph is constructed based on the first relationship data table and the second relationship data table, wherein the entity relationship graph includes entities and edges, wherein the entities include the target entity and auxiliary entities constructed based on the second relationship data table, the edges are association relationships between different entities, and the entities include entity attribute information, which is constructed based on fields contained in the relationship data table corresponding to the entities and / or field data corresponding to the fields; Determine the propagation paths for different auxiliary entities included in the entity relationship graph to reach the target entity, and generate target attribute information of the target entity based on the entity attribute information of the auxiliary entities in each propagation path.