A data processing method and device, electronic equipment, storage medium and product

CN116166663BActive Publication Date: 2026-09-22CCB FINTECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211643933.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2026-09-22
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

[0003]目前,集团派系关系划分大部分使用集团与企业的关系,而集团与企业关系又大多通过客户经理来认定以及维护,取决于客户经理本身认知程度及经验,需要耗费大量人力物力,同时受限于人工,该集团派系识别存在识别延迟

Benefits of technology

[0021]本发明实施例的技术方案,通过识别并提取预先创建的金融数据表中目标金融机构所关联的目标主体种类的第一对象标识集合,识别并提取风险数据表中目标金融机构所关联的第二类型目标对象以及第二类型目标对象包含的第二对象标识集合,在预先创建的对象关联关系表中识别并查找第二类型目标对象所在的第三类型目标对象以及第三类型目标对象包含的第三对象标识集合,根据第一对象标识集合与第三对象标识集合的比对结果,或者,第一对象标识集合与第二对象标识集合的对比结果确定数据监管质量。实现了根据金融数据表、风险数据表与对象关联关系表确定数据监管质量,提高了数据监管质量的准确性,减少通过人工进行数据监管的成本,提升用户的使用体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166663B_ABST
    Figure CN116166663B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and device, electronic equipment, storage medium and product, and relates to the technical field of data processing. The method comprises the following steps: identifying and extracting a first object identifier set of a target subject category associated with a target financial institution in a pre-created financial data table; identifying and extracting a second type target object associated with the target financial institution in a risk data table and a second object identifier set contained in the second type target object; identifying and searching for a third type target object in which the second type target object is located and a third object identifier set contained in the third type target object in a pre-created object association relationship table; and determining data supervision quality according to a comparison result of the first object identifier set and the third object identifier set or a comparison result of the first object identifier set and the second object identifier set. According to the embodiment of the application, the confirmation of the data supervision quality is realized, the artificial cost is reduced, and the accuracy of the data supervision quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data processing method, apparatus, electronic device, storage medium, and product. Background Technology

[0002] With financial innovation and increased financial diversity, the pressure to prevent financial risks has also increased, necessitating improvements in data supervision levels and efficiency. In daily regulatory work, determining factional relationships within groups requires significant time and effort from regulatory personnel. Furthermore, the growing social interactions between groups and between groups and their subsidiaries have led to numerous complex factional relationships within these groups.

[0003] Currently, the classification of group factions largely relies on the relationship between the group and its subsidiaries. This relationship is primarily identified and maintained by account managers, which depends heavily on their knowledge and experience, requiring significant human and material resources. Furthermore, this manual approach leads to delays in faction identification. For banking regulators, accurately identifying group relationships can help detect issues with the quality of data reported by banking institutions. Summary of the Invention

[0004] This invention provides a data processing method, apparatus, electronic device, storage medium, and product to enable quality inspection of data from financial institutions such as banks and improve the accuracy of data supervision.

[0005] According to one aspect of the present invention, a data processing method is provided, wherein the method includes:

[0006] Identify and extract the first set of object identifiers for the target entity types associated with the target financial institution from a pre-created financial data table;

[0007] Identify and extract the second type of target object associated with the target financial institution in the risk data table, as well as the set of second object identifiers contained in the second type of target object;

[0008] Identify and locate the third type of target object containing the second type of target object in the pre-created object association table, as well as the set of third object identifiers contained in the third type of target object;

[0009] The quality of data supervision is determined based on the comparison results between the first set of object identifiers and the third set of object identifiers, or the comparison results between the first set of object identifiers and the second set of object identifiers.

[0010] According to another aspect of the present invention, a data processing apparatus is provided, wherein the apparatus comprises:

[0011] The first set determination module is used to identify and extract the first set of object identifiers of the target entity types associated with the target financial institution in the pre-created financial data table;

[0012] The second set determination module is used to identify and extract the second type of target object associated with the target financial institution in the risk data table and the second object identifier set contained in the second type of target object;

[0013] The third set determination module is used to identify and find the third type target object where the second type target object is located and the third object identifier set contained in the third type target object in the pre-created object association table;

[0014] The data quality determination module is used to determine the data supervision quality based on the comparison results between the first object identifier set and the third object identifier set, or the comparison results between the first object identifier set and the second object identifier set.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory that is communicatively connected to at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the data processing method of any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the data processing method of any embodiment of the present invention.

[0020] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the data processing method of any embodiment of the present invention.

[0021] The technical solution of this invention identifies and extracts a first set of object identifiers for target entities associated with a target financial institution from a pre-created financial data table; identifies and extracts a second type of target object associated with a target financial institution from a risk data table, as well as a second set of object identifiers contained within the second type of target object; identifies and searches for a third type of target object containing the second type of target object and a third set of object identifiers contained within the third type of target object in a pre-created object association table; and determines data regulatory quality based on the comparison results of the first set of object identifiers and the third set of object identifiers, or the comparison results of the first set of object identifiers and the second set of object identifiers. This achieves the determination of data regulatory quality based on financial data tables, risk data tables, and object association tables, improving the accuracy of data regulatory quality, reducing the cost of manual data supervision, and enhancing the user experience.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a data processing method provided according to an embodiment of the present invention;

[0025] Figure 2 This is a flowchart of a data processing method provided according to an embodiment of the present invention;

[0026] Figure 3 This is a flowchart for determining similarity and coverage according to an embodiment of the present invention;

[0027] Figure 4 This is a flowchart of a data processing method provided according to an embodiment of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of a data processing device according to an embodiment of the present invention;

[0029] Figure 6 This is a schematic diagram of the structure of an electronic device that implements a data processing method according to an embodiment of the present invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0033] In one embodiment, Figure 1 This is a flowchart of a data processing method according to an embodiment of the present invention. This embodiment is applicable to determining the regulatory quality of data in financial data tables. The method can be executed by a data processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0034] S110. Identify and extract the first set of object identifiers for the target entity types associated with the target financial institution in a pre-created financial data table.

[0035] The financial data table can refer to a table created based on financial data-related data reported by various financial institutions, which may include, but are not limited to, banks. The financial data table may contain attribute information of the first type of object, for example, including the name of the first type of object, the type of credit entity, and the organization code. In one embodiment, the financial data table may include, but is not limited to, the east data table. The target financial institution can refer to the financial institution to be identified. Since the data in the financial data table is reported by multiple financial institutions, the target financial institution can be one of the reporting financial institutions. The entity type can refer to the type of credit entity, used to characterize the affiliation of the first type of object. In one embodiment, when the first type of object is an enterprise, the type of credit entity can include single legal person credit and group credit, etc. The target entity type can include single legal person credit. The first object identifier set can refer to a collection of enterprises whose target entity type is single legal person credit.

[0036] In one embodiment, a pre-created financial data table can be extracted. Fields related to the target financial institution can be identified within the financial data table, and data related to the target financial institution can be extracted based on these fields. The entity type information associated with the target financial institution is then searched, and the entity type information is identified as a first type of object of the target entity type. The extracted first type objects are then aggregated into a first object identifier set. In one embodiment, information associated with a specific financial institution can be extracted from a pre-created financial institution table. Based on the obtained relevant information, enterprises with single-legal-person credit granting types can be identified, and the identified enterprises are aggregated into a first object identifier set.

[0037] S120. Identify and extract the second type of target object associated with the target financial institution in the risk data table, as well as the set of second object identifiers contained in the second type of target object.

[0038] The risk data table can refer to a table that records risk information for first-type and second-type objects. In practice, the risk data table can record second-type target objects and a set of second-object identifiers belonging to each second-type target object. The second-type target objects may include first-type objects. The set of second-object identifiers can refer to a collection of first-type objects belonging to the second-type target objects. In one embodiment, when the second-type target objects are the groups recorded in the risk data table, the set of second-object identifiers can refer to the enterprises belonging to each group.

[0039] In one embodiment, a pre-stored risk data table can be extracted, and information on second-type target objects associated with the target financial institution can be identified and extracted from the risk data table. Based on the second-type target object, the first-type objects it contains can be located, and the first-type objects belonging to the same second-type target object can be aggregated into a second object identifier set. In another embodiment, data on the target financial institution can be searched in the risk data table to obtain the corresponding associated second-type target objects. The first-type objects contained in the second-type target objects can be extracted to generate a second object identifier set. In one embodiment, the credit granting entity type of the first-type objects in the second object identifier set is group credit.

[0040] S130. Identify and search for the third type of target object where the second type of target object is located, as well as the set of third object identifiers contained in the third type of target object, in the pre-created object association table.

[0041] The object association information table can be a pre-generated information table showing the relationship between third-type target objects and first-type target objects. The third-type target object can refer to an object generated by merging the attribute information of second-type target objects. In practice, a third-type target object can contain one or more second-type target objects, therefore, a third-type target object can contain first-type objects. The third object identifier set can refer to a set generated by summarizing the first-type objects contained within the third-type target objects.

[0042] In one embodiment, information about a second type of target object can be searched in a pre-created object association table to identify a third type of target object to which the second type of target object belongs. The first type of objects contained within the third type of target object are then aggregated into a third object identifier set. In another embodiment, the third type of target group can be found in the object association table by searching for the first type of objects contained within the second type of target object. Once the third type of target group is determined, the first type of objects contained within the third type of target group can be extracted to generate a third object identifier set.

[0043] S140. Determine the data supervision quality based on the comparison results between the first object identifier set and the third object identifier set, or the comparison results between the first object identifier set and the second object identifier set.

[0044] Among them, the quality of data supervision can refer to the degree of good or bad data supervision by the target financial institution, which can be determined based on the comparison results of the first object identifier set and the third object identifier set, or the comparison results of the first object identifier set and the second object identifier set.

[0045] In this embodiment, since the first object identifier set, the second object identifier set, and the third object identifier set all contain corresponding first-type objects, the data supervision quality can be determined by comparing the first object identifier set and the third object identifier set to see if they contain first-type objects, or by comparing the first object identifier set and the second object identifier set to see if they contain first-type objects. In actual operation, since the credit subject type of the first-type objects in the first object identifier set is single-legal-person credit, and the credit subject type of the first-type objects in the second and third object identifier sets is group credit, when the first object identifier set and the third object identifier set do not match, or when the second object identifier set and the third object identifier set should match, the data supervision quality can be considered good. In one embodiment, when the comparison result between the first object identifier set and the third object identifier set is a mismatch, or the comparison result between the second object identifier set and the third object identifier set is a match, it can be considered that the data uploaded by the target financial institution matches the actual data, and the data supervision quality is considered good; when the comparison result between the first object identifier set and the third object identifier set is a match, or the comparison result between the second object identifier set and the third object identifier set is a mismatch, it can be considered that the data uploaded by the target financial institution does not match the actual data, and the data supervision quality is considered poor.

[0046] In this embodiment of the invention, by identifying and extracting a first set of object identifiers for target entities associated with a target financial institution in a pre-created financial data table, identifying and extracting a second type of target object associated with a target financial institution in a risk data table, and a second set of object identifiers contained in the second type of target object, and identifying and searching for a third type of target object in a pre-created object association table, and a third set of object identifiers contained in the third type of target object, the data supervision quality is determined based on the comparison result between the first set of object identifiers and the third set of object identifiers, or the comparison result between the first set of object identifiers and the second set of object identifiers. This achieves convenient confirmation of data supervision quality, improves the accuracy of data supervision quality, reduces the cost of manual data supervision, and enhances the user experience.

[0047] In some embodiments, Figure 2 This is a flowchart of a data processing method provided according to an embodiment of the present invention. This embodiment further illustrates a data processing method based on the above embodiments. Figure 2 As shown, the method includes:

[0048] S2010: Obtain the pre-created initial object collection.

[0049] The initial object set can refer to a collection of unprocessed second-type objects.

[0050] In this embodiment, the initial object set may be pre-created, and the pre-created initial object set may be retrieved locally on the electronic device to obtain the pre-created initial object set.

[0051] S2020. Perform a culling operation based on the second object attribute information of each second type object in the initial object set to obtain the corresponding target object set.

[0052] The second object attribute information may be parameter information representing the second type of object. The second object attribute information may include, but is not limited to, the number of first type objects contained in each second type object, and the first object attribute information corresponding to each first type object.

[0053] In one embodiment, second object attribute information that meets the removal rules can be removed from the initial object set based on the second object attribute information of the second type of object, and the initial object set after removal can be used as the first object attribute information. In one embodiment, the second type of object that meets the removal rules may include, but is not limited to, the second type of object having the name of the target name and the first object identifier being empty. By performing the removal operation on the initial object set, the corresponding target object set is obtained.

[0054] In one embodiment, S2020 includes:

[0055] S2021. Obtain the name of the second object corresponding to each second type object in the initial object set, and the identifier of the first object corresponding to the first type object contained therein.

[0056] Wherein, the second object name can refer to the name of a second type of object; the first object identifier can refer to the identifier of a first type of object. When the second type of object is a group, the second object name can be the group name, the first type of objects contained in the second type of object can be enterprises contained in the group, and the first object identifier can be an identifier used to identify the first type of object. In one embodiment, the first object identifier may include the name of the first type of object, organization code, etc.

[0057] In this embodiment, information about each second-type object can be obtained from the initial object set, and the second object name field corresponding to the second-type object information can be extracted to obtain the second object name corresponding to each second-type object. After obtaining the information about each second-type object, the first-type object contained in each second-type object can be obtained, and the first object identifier field contained in the first-type object can be extracted to obtain the first object identifier corresponding to the first-type object.

[0058] S2022. Remove the second type object, whose second object name is the target name and whose first object identifier is empty, from the initial object set to obtain the corresponding target object set.

[0059] The target name can refer to the name to be removed. In practice, the target name may include encrypted names that cannot be displayed. For example, the target name can be empty, or it may be ***, etc. The target object set can be the set generated after removing the initial set of second-type data objects whose second object name is the target name and whose first object identifier is empty.

[0060] In an embodiment, a second type of object with the second object name being the target name and the first object identifier being empty can be found in the initial data object. The data corresponding to the second type of data with the second object name being the target name and the first object identifier being empty can be cleared, and the initial object set after removing the corresponding second type of object can be used as the target object set.

[0061] S2030. Obtain the second object attribute information of each second type object in the pre-created target object set; wherein, the second object attribute information includes: the number of first type objects contained in each second type object, and the first object attribute information corresponding to each first type object.

[0062] The second object attribute information may be parameter information that includes the relationship between the second type object and the first type object; the first object attribute information may refer to the parameter information contained in the first type object. For example, the attribute information of the first type object may include information such as the first object identifier.

[0063] In this embodiment, information about each second-type object can be extracted from a pre-created set of target objects to determine the attribute information of each second-type object. The attribute information of the second objects may include, but is not limited to, the number of first-type objects contained in each second-type object, and the attribute information corresponding to each first-type object. The number of first-type objects contained in each second-type object can be different. For example, when the second-type object is a group, the first-type objects can be enterprises belonging to the group, and the number of enterprises contained in each group can be different. After determining the first-type objects contained in each second-type object, the attribute information of each first-type object can be extracted. For example, the attribute information of the first objects may include, but is not limited to, the name and organization code of the first-type object.

[0064] S2040. Determine the similarity and coverage between every two second-type objects based on the second object attribute information.

[0065] Here, similarity can refer to the degree of similarity between any two second-type objects; coverage can refer to the degree of coverage between any two second-type objects.

[0066] In one embodiment, attribute information for every two second objects can be extracted, and the similarity and coverage of the second-type objects can be determined based on the attribute information of the second-type objects. In actual operation, the similarity and coverage between every two types of objects can be determined based on the number of first-type objects contained in each second-type object and the attribute information of the first objects corresponding to each first-type object. In one embodiment, the number of first-type objects contained in every two second-type objects can be obtained, and the number of common first-type objects contained in the two second-type objects can be determined based on the attribute information of the first-type objects. The maximum value of the number of common first-type objects in the two second-type objects can be determined, and the similarity between the two second-type objects can be determined by dividing the number of common first-type objects by the maximum value of the two second-type objects. In another embodiment, the number of common first-type objects contained in the two second-type objects can be determined based on the attribute information of the first-type objects, and the minimum value of the number of common first-type objects in the two second-type objects can be determined. The coverage between the two second-type objects can be determined by dividing the number of common first-type objects by the minimum value of the two second-type objects.

[0067] In one embodiment, S2040 includes:

[0068] S2041. Classify the second type objects in the target object set based on the number of first type objects contained in the second type objects to obtain the corresponding object groups.

[0069] Here, an object group refers to a group of second-type objects in a target object set, divided according to the first-type objects contained within the second-type objects. The number of object groups can be multiple, for example, 4, 5, 6, etc. In practice, the second-type objects in the target object set can be divided into equal-frequency groups based on the number of first-type objects contained within each second-type object. The interval boundaries of the equal-frequency stratification can be used to divide the second-type objects containing different numbers of first-type objects into different object groups. The interval boundary values ​​for equal-frequency stratification can be set based on the experience of business personnel to ensure that each object group contains approximately the same number of second-type objects. In one embodiment, the interval boundary values ​​for equal-frequency stratification can include [1,2], [3,6], [7,24], [25,99], and [100, ) etc.

[0070] S2042. Determine the similarity between every two second-type objects in the same object group based on the number of first-type objects contained in the second-type objects.

[0071] In one embodiment, the number of first-type objects contained in two second-type objects and a first object identifier can be extracted from the same object group. The number of identical first-type objects contained in the two second-type objects is determined based on the first object identifier. The similarity between each pair of second-type objects in the same group is determined based on the number of identical first-type objects and the total number of second-type objects. In one embodiment, the maximum number of two second-type objects in the same object group can be determined. The similarity between the two second-type objects is determined by dividing the number of identical first-type objects contained in these two second-type objects by the maximum number of the two second-type objects. In one embodiment, the similarity between each pair of second-type objects in the same object group can be determined separately.

[0072] S2043. Merge the second type of objects whose similarity reaches the preset similarity threshold to obtain the corresponding third type of initial object.

[0073] The third type of initial object can refer to the object generated by merging second-type objects whose similarity reaches a preset similarity threshold. The preset similarity threshold can be a pre-set threshold for similarity; when the similarity between two second-type objects reaches the preset similarity threshold, the two second-type objects can be merged. In actual operation, the preset similarity thresholds corresponding to different object groups can be different. The preset similarity threshold can be set by business personnel based on experience, or it can be determined based on the interval boundary values ​​of the equal-frequency stratification. For example, when the preset similarity threshold is determined based on the interval boundary of the equal-frequency stratification, and when the interval of the equal-frequency stratification includes [1,2], [3,6], [7,24], [25,99], and [100,), the preset similarity thresholds can be 1, 1 / 2, 7 / 24, 25 / 99, and 1 / 30, respectively.

[0074] In one embodiment, two second-type objects whose similarity reaches a preset similarity threshold can be merged. Once it is determined that all two second-type objects in the same group whose similarity reaches the preset similarity threshold have been merged, the merged second-type objects can be used as the initial objects of the third type. The initial objects of the third type can include all first-type objects contained in the second-type objects in the same group. The number of first-type objects contained in the initial objects of the third type can be the number of first-type objects contained in the second-type objects before merging minus the number of identical first-type objects in each second-type object in the same group. In one embodiment, the number of corresponding initial objects of the third type in the same group is not limited, as long as all second-type objects whose similarity reaches the preset similarity threshold are merged.

[0075] S2044. Determine the coverage between every two third-type initial objects from different object groups, the third-type initial objects and the second-type objects, or every two second-type objects, based on the number of first-type objects contained in the third-type initial objects and / or the second-type objects.

[0076] In this embodiment, second-type objects with a similarity threshold within the same object group are merged to generate third-type initial objects, while second-type objects with a similarity threshold are not merged. The coverage ratio between each pair of third-type initial objects, between a third-type initial object and a second-type object, or between two second-type objects can be determined separately. In actual operation, the number of first-type objects and their identifiers within each pair of third-type initial objects, between a third-type initial object and a second-type object, or between two second-type objects can be obtained separately. Based on the first object identifiers, the number of identical first-type objects within each pair of third-type initial objects, between a third-type initial object and a second-type object, or between two second-type objects can be determined. The coverage ratio between each pair of third-type initial objects, between a third-type initial object and a second-type object, or between each pair of second-type objects can be determined based on the number of first-type objects and the number of identical first-type objects within each pair of third-type initial objects, between a third-type initial object and a second-type object, or between each pair of second-type objects. In one embodiment, the minimum number of first-type objects contained in each pair of third-type initial objects and second-type objects, or in each pair of second-type objects, can be determined. The coverage between each pair of third-type initial objects, third-type initial objects and second-type objects, or in each pair of second-type objects from different object groups is determined by dividing the number of first-type objects contained in each pair of third-type initial objects and second-type objects, or in each pair of second-type objects, by the minimum number of first-type objects contained in each pair of third-type initial objects and second-type objects.

[0077] S2050. Merge the second type of objects based on similarity and coverage to generate the corresponding object association table.

[0078] The object association table can refer to the association table between the merged second type of object and the first type of object.

[0079] In this embodiment, when the coverage between any two third-type initial objects from different object groups, between a third-type initial object and a second-type object, or between any two second-type objects reaches a preset threshold coverage, the two third-type initial objects, the third-type initial object and the second-type object, or the two second-type objects can be merged. The preset coverage can be pre-set based on the experience of business personnel. A corresponding object association table can be generated based on the merged second-type objects and third-type initial objects. In one embodiment, when two second-type objects and / or third-type initial objects are merged, the corresponding first-type objects they contain can be attributed to the merged second-type objects and / or third-type initial objects. A corresponding object association table can be established based on the merged second-type objects, third-type initial objects, and the contained first-type objects.

[0080] S2060. Identify and extract the first set of object identifiers for the target entity types associated with the target financial institution in a pre-created financial data table.

[0081] S2070. Identify and extract the second type of target object associated with the target financial institution in the risk data table, and the set of second object identifiers contained in the second type of target object.

[0082] S2080. Identify and search for the third type target object where the second type target object is located, as well as the set of third object identifiers contained in the third type target object, in the pre-created object association table.

[0083] S2090. Determine the matching situation between the identifier of the first type of target object in the first object identifier set and the identifier of the first type of object in the third object identifier set, and take it as the first matching situation.

[0084] In this embodiment, the identifier of a first target object in a first object identifier set and the identifier of a first type object in a third object identifier set can be obtained respectively. The identifiers of the first type target object in the first object identifier set and the first type object in the third object identifier set are matched to determine whether the identifiers of the first target object in the first object identifier set and the first type object in the third object identifier set are the same, and the matching result is taken as the first matching result. The first matching result can include matching and non-matching.

[0085] S2100. Determine the matching situation between the identifier of the first type of target object in the first object identifier set and the identifier of the first type of object in the second object identifier set, and take it as the second matching situation.

[0086] In this embodiment, the identifier of a first target object in a first object identifier set and the identifier of a first type object in a second object identifier set can be obtained respectively. The identifiers of the first type target object in the first object identifier set and the first type objects in the second object identifier set are matched to determine whether the identifiers of the first target object in the first object identifier set and the first type objects in the second object identifier set are the same. The matching result is taken as the second matching result. The second matching result can include matching and non-matching.

[0087] S2110. Determine the quality of data supervision based on the first matching result or the second matching result.

[0088] In this embodiment, after obtaining the first matching result or the second matching result, the data supervision quality can be determined based on the first matching result or the second matching result. When the first matching result is a mismatch or the second matching result is a match, the data uploaded by the target financial institution can be considered to match the actual data, and the data supervision quality is considered to be good. When the first matching result is a match or the second matching result is a mismatch, the data uploaded by the target financial institution can be considered to mismatch the actual data, and the data supervision quality is considered to be poor.

[0089] In this embodiment of the invention, a target object set is obtained by acquiring a pre-created initial object set and performing a removal operation based on the second object attribute information of each second type object in the initial object set. The target object set is then acquired by acquiring the second object attribute information of each second type object in the pre-created target object set, determining the similarity and coverage between any two second type objects based on the second object attribute information, and merging the second type objects according to the similarity and coverage to generate a corresponding object association table. This system identifies and extracts a first set of object identifiers associated with target financial institutions from a pre-created financial data table, and identifies and extracts a second type of target object associated with the target financial institution from a risk data table, as well as a second set of object identifiers contained within the second type of target object. Based on the pre-created object association relationships, it determines the matching status between the identifiers of the first type of target object in the first object identifier set and the identifiers of the first type of target object in the third object identifier set, using this as the first matching status. It then determines the matching status between the identifiers of the first type of target object in the first object identifier set and the identifiers of the first type of target object in the second object identifier set, using this as the second matching status. Based on either the first or second matching status, it determines the data regulatory quality. This system achieves the determination of similarity and coverage between every two second type of objects through second object attribute information, and merges second type of objects based on similarity and coverage. This enables accurate judgment and merging of second type of object relationships, facilitating the identification of quality issues in the data reported by target financial institutions, thereby achieving accurate judgment of data regulatory quality and improving the user experience.

[0090] In one embodiment, Figure 3 This is a flowchart for determining similarity and coverage according to an embodiment of the present invention. This embodiment is a further explanation of determining the similarity and coverage between every two second-type objects based on the second object attribute information, based on the above embodiment. Figure 3 As shown, the method includes:

[0091] S310. Determine the comparison result between the number of first-type objects contained in the second-type object and the pre-configured threshold for the number of each object.

[0092] The object quantity threshold can refer to the threshold for the number of first-type objects contained in a second-type object, or it can be a threshold used to divide object groups. The object quantity threshold can be set by business personnel based on experience, or it can be set based on the number of first-type objects contained in each second-type object. There can be one or more object quantity thresholds, as long as the number of second-type objects within each threshold range is basically evenly distributed.

[0093] In this embodiment, pre-configured object quantity thresholds can be extracted, and the number of first-type objects contained in each second-type object can be compared with each object quantity threshold in turn to determine the comparison result between the number of first-type objects contained in the second-type object and each pre-configured object quantity threshold.

[0094] S320. Based on the comparison results, classify the second type of objects in the target object set to obtain the corresponding object groups.

[0095] In an embodiment, the second type of objects in the target object set can be classified according to the comparison results. The second type of objects whose number of first type objects meets the threshold range can be classified into the same object group to obtain the corresponding object group.

[0096] S330. Determine the number of identical first-type objects contained in every two second-type objects in the same object group, and use this as the first quantity.

[0097] In one embodiment, the number of identical first-type objects contained in every two second-type objects within the same object group can be obtained as a fourth quantity. In another embodiment, two second-type objects can be extracted from the same object group, and the number of identical first-type objects contained in the two second-type objects can be determined based on the first object identification information, which is then used as a first quantity. For example, when one second-type object contains 4 first-type objects, another second-type object contains 5 first-type objects, and the two second-type objects contain 3 identical first-type objects, then 3 can be used as the first quantity.

[0098] S340. Determine the number of first-type objects contained in each pair of second-type objects in the same object group, and use them as the second quantity and the third quantity.

[0099] In one embodiment, the number of first-type objects contained in every two second-type objects within the same object data can be determined, and these can be used as the second and third quantities. In another embodiment, two second-type objects can be extracted from the same object group, and the number of first-type objects contained in every two second-type objects can be determined, and these can be used as the second and third quantities. For example, if one second-type object contains 4 first-type objects and the other second-type object contains 5 first-type objects, then the second quantity can be considered to be 4, and the third quantity to be 5.

[0100] S350. Determine the similarity between every two second-type objects in the same object group based on the maximum value of the second and third quantities, and the first quantity.

[0101] In an embodiment, the size of a second quantity and a third quantity can be compared, and the similarity between every two second-type objects in the same object group can be determined based on the maximum value of the second and third quantities and the first quantity. In practice, the similarity between every two second-type objects can be calculated by dividing the first quantity by the maximum value of the second and third quantities. For example, when the first quantity is 3, the second quantity is 4, and the third quantity is 5, the similarity between two second-type objects can be 3 / 5.

[0102] S360. Merge the second type of objects whose similarity reaches the preset similarity threshold to obtain the corresponding third type of initial object.

[0103] S370. Determine the coverage between every two third-type initial objects from different object groups, the third-type initial objects and the second-type objects, or every two second-type objects, based on the number of first-type objects contained in the third-type initial objects and / or the second-type objects.

[0104] In one embodiment, when determining the coverage between every two third-type initial objects from different object groups based on the number of first-type objects contained in the third-type initial object, S370 accordingly includes:

[0105] S371. Determine the number of identical first-type objects contained in every two third-type initial objects from different object groups, as the fourth quantity.

[0106] In one embodiment, the number of identical first-type objects contained in every two third-type initial objects from different object groups can be determined as a fourth quantity. In another embodiment, two third-type initial objects can be extracted from different object groups, and the number of identical first-type objects contained in the two third-type initial objects can be determined based on the first object identification information, as the fourth quantity. For example, when one third-type initial object contains 7 first-type objects, another third-type initial object contains 8 first-type objects, and the two third-type initial objects contain 2 identical first-type objects, then 2 can be used as the first quantity.

[0107] S372. Determine the number of first-type objects contained in each of the two third-type initial objects from different object groups, as the fifth and sixth quantities.

[0108] In one embodiment, the number of first-type objects contained in every two third-type initial objects in different object groups can be determined, and these can be designated as the fifth and sixth quantities. In another embodiment, two third-type initial objects can be extracted from different object groups, and the number of first-type objects contained in every two third-type initial objects can be determined, and these can be designated as the fifth and sixth quantities. For example, if one third-type initial object contains 7 first-type objects and the other third-type initial object contains 8 first-type objects, the fifth quantity can be considered to be 7 and the sixth quantity to be 8.

[0109] S373. Determine the coverage between every two initial objects of the third type from different object groups based on the minimum of the fifth and sixth quantities and the fourth quantity.

[0110] In this embodiment, the coverage between two initial objects of the third type from different object groups can be determined by comparing the fifth and sixth quantities and using the minimum of the fifth and sixth quantities along with the fourth quantity. In practice, the coverage between two initial objects of the third type can be calculated by dividing the fourth quantity by the minimum of the fifth and sixth quantities. For example, when the fourth quantity is 2, the fifth quantity is 7, and the sixth quantity is 8, the coverage between two initial objects of the third type can be 2 / 7.

[0111] In one embodiment, when determining the coverage ratio between third-type initial objects and second-type objects from different object groups based on the number of first-type objects contained in the third-type initial objects and the second-type objects, S370 accordingly includes:

[0112] S374. Determine the number of identical first-type objects contained in the third-type initial objects and the second-type objects from different object groups, as the seventh quantity.

[0113] In one embodiment, the number of identical first-type objects contained in different third-type initial objects and second-type objects can be determined as the seventh quantity. In another embodiment, third-type initial objects and second-type objects can be extracted from different object groups, and the number of identical first-type objects contained in the third-type initial objects and second-type objects can be determined based on the first object identification information, as the seventh quantity. For example, when the number of first-type objects contained in the third-type initial objects is 8, the number of first-type objects contained in the second-type objects is 7, and the number of identical first-type objects contained in the third-type initial objects and second-type objects is 3, then 3 can be used as the seventh quantity.

[0114] S375. Determine the number of first-type objects contained in the third-type initial objects and the second-type objects from different object groups, respectively, as the eighth and ninth quantities.

[0115] In one embodiment, the number of first-type objects contained in the third-type initial objects and the second-type objects from different object groups can be determined, and these can be designated as the eighth and ninth quantities. In another embodiment, the third-type initial objects and the second-type objects can be extracted from different object groups, and the number of first-type objects contained in the third-type initial objects and the second-type objects can be determined, and these can be designated as the eighth and ninth quantities. For example, if the number of first-type objects contained in the third-type initial objects is 8, and the number of first-type objects contained in the second-type objects is 7, then the eighth quantity can be considered to be 8, and the ninth quantity to be 7.

[0116] S376. Determine the coverage between the third type initial object and the second type object from different object groups based on the minimum of the eighth and ninth quantities and the seventh quantity.

[0117] In this embodiment, the coverage ratio between the third type initial objects and the second type objects from different object groups can be determined by comparing the size of the eighth and ninth quantities and using the seventh quantity as the basis for calculation. In practice, the coverage ratio between the third type initial objects and the second type objects can be calculated by dividing the seventh quantity by the minimum of the eighth and ninth quantities. For example, when the seventh quantity is 3, the eighth quantity is 8, and the ninth quantity is 7, the coverage ratio between the third type initial objects and the second type objects can be 3 / 7.

[0118] In one embodiment, when determining the coverage between every two second-type initial objects from different object groups based on the number of first-type objects contained in the second-type initial objects, S370 accordingly includes:

[0119] S377. Determine the number of identical first-type objects contained in every two second-type initial objects from different object groups, as the tenth number.

[0120] In one embodiment, the number of identical first-type objects contained in every two second-type initial objects from different object groups can be determined as the tenth number. In another embodiment, two second-type initial objects can be extracted from different object groups, and the number of identical first-type objects contained in the two second-type initial objects can be determined based on the first object identification information, which is then used as the tenth number. For example, when one second-type initial object contains 5 first-type objects, the other second-type initial object contains 4 first-type objects, and the two second-type initial objects contain 1 identical first-type object, then 1 can be used as the tenth number.

[0121] S378. Determine the number of first-type objects contained in each of the two second-type initial objects from different object groups, as the eleventh and twelfth quantities.

[0122] In one embodiment, the number of first-type objects contained in every two second-type initial objects in different object groups can be determined, and designated as the eleventh and twelfth quantities. In another embodiment, two second-type initial objects can be extracted from different object groups, and the number of first-type objects contained in every two second-type initial objects can be determined, designated as the eleventh and twelfth quantities. For example, if one second-type initial object contains 5 first-type objects and the other contains 4, the eleventh quantity can be considered to be 5 and the twelfth quantity to be 4.

[0123] S379. Determine the coverage between every two second-type initial objects from different object groups based on the minimum of the eleventh and twelfth quantities and the tenth quantity.

[0124] In an embodiment, the coverage between two second-type initial objects from different object groups can be determined by comparing the eleventh and twelfth quantities, and using the minimum of the eleventh and twelfth quantities along with the tenth quantity. In practice, the coverage between two second-type initial objects can be calculated by dividing the tenth quantity by the minimum of the eleventh and twelfth quantities. For example, when the tenth quantity is 1, the eleventh quantity is 5, and the twelfth quantity is 4, the coverage between two second-type initial objects can be 1 / 4.

[0125] S371-S373, S374-S376, and S377-S379 are parallel schemes for determining the coverage of every two third-type initial objects from different object groups, the third-type initial objects and second-type objects, or every two second-type objects, based on the number of first-type objects contained in the third-type initial objects and / or the second-type objects. There is no specific execution order.

[0126] In this embodiment of the invention, by determining the comparison result between the number of first-type objects contained in the second-type objects and a pre-configured threshold for the number of each object, the second-type objects in the target object set are classified according to the comparison result to obtain corresponding object groups. The similarity of second-type objects within the same object group is calculated, making the calculation of the similarity between every two second-type objects more reasonable. By determining the coverage between every two initial third-type objects from different object groups, the coverage between initial third-type objects and second-type objects, and the coverage between every two second-type objects, it is possible to merge every two initial third-type objects, every initial third-type object and every second-type object, and every two second-type objects, thereby improving the accuracy of merging second-type objects.

[0127] In one embodiment, Figure 4 This is a flowchart of a data processing method provided by an embodiment of the present invention. This embodiment, based on the above embodiments, uses a first type of group as an enterprise, a second type of object as a group, a third type of object as a group merged based on enterprise information, and a target financial institution as a target bank, to illustrate a specific implementation of a data processing method. Figure 4 As shown, the method includes:

[0128] S4010. Remove classified groups and groups whose organizational codes are empty from the initial object set, and generate the target object set.

[0129] The name of the classified group can be the target name, for example, it can be "*********".

[0130] In one embodiment, groups whose target name is "Group" and whose organizational codes for the companies they contain are empty can be removed from the initial object set to obtain the corresponding target object set. In another embodiment, all company names and their organizational codes within the same group can be stored separately in newly added data frame columns "Company Name Set" and "Organization Code Set Column".

[0131] S4020. Based on the number of enterprises contained in the group, perform equal-frequency stratification to obtain the corresponding object groups. Groups within the same object group are merged based on the similarity of the enterprises.

[0132] In this embodiment, the group can be divided into multiple object groups based on the number of enterprises it contains, with each object group maintaining a relatively balanced number of enterprises. The groups in the target object set can be stratified by frequency according to the number of enterprises they contain, and groups containing different numbers of enterprises can be divided into different object groups based on the interval boundaries of the frequency stratification. The interval boundary values ​​for frequency stratification can be set based on the experience of business personnel to ensure that each object group contains approximately the same number of groups. In one embodiment, the interval boundary values ​​for frequency stratification may include [1,2], [3,6], [7,24], [25,99], and [100,), etc., and the preset similarity thresholds can be 1, 1 / 2, 7 / 24, 25 / 99, and 1 / 30, respectively. After dividing the object groups, groups within the same object group can be merged based on enterprise similarity. In one embodiment, the similarity can be determined based on the number of enterprises in the group and the enterprise identifiers. In practice, enterprise identifiers may include organization codes. In one embodiment, the similarity can be calculated using the following formula:

[0133] Among them G i Let len(G) represent the set of enterprises contained in group i. i ) represents G i The number of companies, len(G) i ∩G j ) represents G i With G j The length of the intersection of the two groups is given by θ, which represents the similarity between the two groups. Groups that meet the preset similarity threshold are considered to be related and are merged into the same group.

[0134] S4030. Groups in different object groups are merged based on enterprise coverage.

[0135] In one embodiment, groups can be extracted from different object groups and merged based on the coverage of the enterprises contained within each group. In one embodiment, similarity can be determined based on the number of enterprises contained in the group and the enterprise identifiers. In one embodiment, coverage can be calculated using the following formula:

[0136] Among them G i Let len(G) represent the set of enterprises contained in group i. i ) represents G i The number of companies, len(G) i ∩G j ) represents G i With G jThe length of the intersection of the two groups, f, represents the coverage rate of the two groups. When the coverage rate reaches a preset coverage rate threshold, the two groups are merged. The preset coverage rate threshold can be a threshold pre-set by business personnel based on experience; for example, it may include, but is not limited to, 0.4, 0.5, 0.6, etc. In one embodiment, groups in different object groups include groups that have completed mergers within the same object group and groups that have not merged with other groups. Therefore, merging groups in different object groups based on enterprise coverage rate can include merging every two groups that have completed mergers within the same object group, merging every two groups that have not merged with other groups, and merging groups that have completed mergers within the same object group with groups that have not merged with other groups. Two groups that meet the preset coverage rate threshold are considered to have a relationship and can be merged into the same group. In one embodiment, the name of the group with the most merged enterprises can be taken as the name of the final merged group, and the reporting bank corresponding to the group with the most participating merged enterprises can be taken as the reporting bank for the final merge result.

[0137] S4040: Groups in different object groups are merged based on enterprise coverage.

[0138] S4050. Construct an object association table based on the information of the merged group and the enterprises contained within the group.

[0139] S4060. Construct a financial data table based on the authorization information table and the corporate client information table.

[0140] In this embodiment, the east data reported by banking institutions includes a credit information table, which contains fields such as customer unique identifier, customer name, credit agreement number, and type of credit entity (e.g., single legal entity credit, group credit, etc.). The east data also includes a corporate customer information table, whose main fields include customer unique identifier, customer name, and organization code. The credit information table and the corporate customer information table can be linked through unified customer information to obtain a wide table containing key fields such as customer name, organization code, and type of credit entity. This wide table can then be used as the financial data table. In one embodiment, the enterprises in the financial data table represent single legal entities that do not belong to any group.

[0141] S4070 Extract the set of enterprises in the financial data table whose credit granting entities are associated with the target bank institution and whose credit granting type is a single legal person.

[0142] S4080 Extract the target bank institution's associated groups and the first set of enterprises contained in the group from the object association table.

[0143] S4090. Identify and extract the groups associated with the target banking institution and the second set of enterprises contained within the groups from the risk data table.

[0144] S4100. Determine the quality of data supervision based on the comparison results between the set of enterprises with single legal person credit and the second set of enterprises, or the comparison results between the first set of enterprises and the second set of enterprises.

[0145] In this embodiment, when the comparison result between the single-legal-entity credit enterprise set and the second enterprise set is a mismatch, or the comparison result between the first enterprise set and the second enterprise set is a match, it can be considered that the data uploaded by the target bank institution matches the actual data, and the data supervision quality is considered good. Conversely, when the comparison result between the single-legal-entity credit enterprise set and the second enterprise set is a match, or the comparison result between the first enterprise set and the second enterprise set is a mismatch, it can be considered that the data uploaded by the target bank institution does not match the actual data, and the data supervision quality is considered poor. In other words, if the enterprises reported by the same bank institution are both single-legal-entity enterprises and belong to the merged group, there is a high probability that the data reported by the bank has data quality problems. It is possible that enterprises that should be included in the group's credit reporting were not included in the group's credit reporting, but were mistakenly reported as separate enterprises, thus revealing a problem with the bank's regulatory data quality.

[0146] In one embodiment, taking Shanghai Bank as the target bank as an example, the organizational codes x of enterprises reported by Shanghai Bank in the financial data table with a credit subject type of single legal person credit are filtered out. Group a and its organizational codes reported by Shanghai Bank in the risk data table are filtered out. In the object association table after group identification, the large group C to which Group a belongs and all organizational codes of large group C are found. If the organizational code of a single legal person credit exists in group C but not in group a, it indicates that Shanghai Bank did not include the organizational code x of the single legal person credit in group a and reported the data in the form of a group, but incorrectly reported the data as a single legal entity, resulting in an incorrect identification of the enterprise group relationship. In one embodiment, Group a and its organizational codes reported by Shanghai Bank in the risk data table are filtered out. In the object association table after group identification, the large group C to which Group a belongs and all organizational codes of C are found. If the enterprise organization code x is found to exist in group a, Shanghai Bank will report x as a single legal entity and also as a group in group a. This indicates that there is an error in the data reporting of either the east data or the risk data table.

[0147] In one embodiment, Figure 5 This is a schematic diagram of the structure of a data processing apparatus according to an embodiment of the present invention. Figure 5 As shown, the device includes: a first set determination module 51, a second set determination module 52, a third set determination module 53, and a data quality determination module 54.

[0148] The first set determination module 51 is used to identify and extract the first object identifier set of the target entity types associated with the target financial institution in the pre-created financial data table.

[0149] The second set determination module 52 is used to identify and extract the second type of target object associated with the target financial institution in the risk data table and the second object identifier set contained in the second type of target object.

[0150] The third set determination module 53 is used to identify and find the third type target object where the second type target object is located and the third object identifier set contained in the third type target object in the pre-created object association table.

[0151] The data quality determination module 54 is used to determine the data supervision quality based on the comparison results between the first object identifier set and the third object identifier set, or the comparison results between the first object identifier set and the second object identifier set.

[0152] In this embodiment of the invention, a first set determination module identifies and extracts a first set of object identifiers for target entities associated with a target financial institution from a pre-created financial data table. A second set determination module identifies and extracts a second type of target object associated with a target financial institution from a risk data table, as well as a second set of object identifiers contained within the second type of target object. A third set determination module identifies and searches for a third type of target object containing the second type of target object in a pre-created object association table, as well as a third set of object identifiers contained within the third type of target object. A data quality determination module determines the data supervision quality based on the comparison results of the first set of object identifiers and the third set of object identifiers, or the comparison results of the first set of object identifiers and the second set of object identifiers. This enables convenient confirmation of data supervision quality, improves the accuracy of data supervision quality, reduces the cost of manual data supervision, and enhances the user experience.

[0153] In one embodiment, a data processing apparatus further includes:

[0154] Obtain the second object attribute information for each second type object in the pre-created target object set; wherein, the second object attribute information includes: the number of first type objects contained in each second type object, and the first object attribute information corresponding to each first type object.

[0155] The similarity coverage determination module is used to determine the similarity and coverage between every two second-type objects based on the attribute information of the second object.

[0156] The second type of objects are merged based on similarity and coverage to generate a corresponding object association table.

[0157] In one embodiment, the similarity coverage determination module includes:

[0158] The object group generation unit is used to classify the second type objects in the target object set based on the number of first type objects contained in the second type objects, and obtain the corresponding object groups.

[0159] A similarity determination unit is used to determine the similarity between every two second-type objects in the same object group based on the number of first-type objects contained in the second-type object.

[0160] The third type of initial object generation unit is used to merge second type objects whose similarity reaches a preset similarity threshold to obtain the corresponding third type of initial object.

[0161] The coverage determination unit is used to determine the coverage between every two third-type initial objects from different object groups, between third-type initial objects and second-type objects, or between every two second-type objects, based on the number of first-type objects contained in the third-type initial objects and / or the second-type objects.

[0162] In one embodiment, the object group generation unit includes:

[0163] The comparison result determination unit is used to determine the comparison result between the number of first-type objects contained in the second-type object and a pre-configured threshold for the number of each object.

[0164] The group generation unit is used to classify the second type of objects in the target object set according to the comparison results, and obtain the corresponding object groups.

[0165] In one embodiment, the similarity determination unit includes:

[0166] The first quantity determination unit is used to determine the number of identical first-type objects contained in every two second-type objects in the same object group, which is taken as the first quantity.

[0167] The second and third quantity determination unit is used to determine the number of first type objects contained in each of two second type objects in the same object group, as the second quantity and the third quantity.

[0168] The object similarity determination unit is used to determine the similarity between every two second-type objects in the same object group based on the maximum value of the second and third quantities, and the first quantity.

[0169] In one embodiment, the coverage determination unit includes:

[0170] The fourth quantity determination unit is used to determine the number of identical first-type objects contained in every two third-type initial objects from different object groups, as the fourth quantity.

[0171] The fifth and sixth quantity determination units are used to determine the number of first-type objects contained in each of two third-type initial objects from different object groups, as the fifth and sixth quantities.

[0172] The first coverage determination unit is used to determine the coverage between every two initial objects of the third type from different object groups based on the minimum of the fifth and sixth quantities, and the fourth quantity.

[0173] In one embodiment, the coverage determination unit includes:

[0174] The seventh quantity determination unit is used to determine the number of identical first-type objects contained in the third-type initial objects and the second-type objects from different object groups, as the seventh quantity.

[0175] The eighth and ninth quantity determination unit is used to determine the number of first-type objects contained in the third-type initial objects and the second-type objects from different object groups, respectively, as the eighth quantity and the ninth quantity.

[0176] The second coverage determination unit is used to determine the coverage between the third type initial object and the second type object from different object groups based on the minimum of the eighth and ninth quantities and the seventh quantity.

[0177] In one embodiment, the coverage determination unit includes:

[0178] The tenth quantity determination unit is used to determine the number of identical first-type objects contained in every two second-type initial objects from different object groups, as the tenth quantity.

[0179] The eleventh and twelfth quantity determination unit is used to determine the number of first-type objects contained in each of two second-type initial objects from different object groups, as the eleventh quantity and the twelfth quantity.

[0180] The third coverage determination unit is used to determine the coverage between every two second-type initial objects from different object groups based on the minimum of the eleventh and twelfth quantities, and the tenth quantity.

[0181] In one embodiment, a data processing apparatus further includes:

[0182] The initial object collection creation module is used to obtain a pre-created initial object collection.

[0183] The target object set generation module is used to perform a culling operation based on the second object attribute information of each second type object in the initial object set to obtain the corresponding target object set.

[0184] In one embodiment, the target object collection generation module further includes:

[0185] The name identifier acquisition unit is used to obtain the name of the second object corresponding to each second type object in the initial object set, and the first object identifier corresponding to the first type object contained therein.

[0186] The target object set generation unit is used to remove second-type objects whose second object name is the target name and whose first object identifier is empty from the initial object set to obtain the corresponding target object set.

[0187] In one embodiment, the data quality determination module 54 includes:

[0188] The first matching situation determination unit is used to determine the matching situation between the identifier of the first type target object in the first object identifier set and the identifier of the first type object in the third object identifier set, as the first matching situation.

[0189] The second matching determination unit is used to determine the matching situation between the identifier of the first type target object in the first object identifier set and the identifier of the first type object in the second object identifier set, as the second matching situation.

[0190] The data quality determination unit is used to determine the data supervision quality based on the first matching condition or the second matching condition.

[0191] The data processing apparatus provided in this embodiment of the invention can execute a data processing method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0192] In one embodiment, Figure 6 This is a schematic diagram of the structure of an electronic device 10 implementing a data processing method according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0193] like Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0194] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0195] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a data processing method.

[0196] In some embodiments, a data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of a data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a data processing method by any other suitable means (e.g., by means of firmware).

[0197] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0198] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0199] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0200] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0201] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0202] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0203] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the data processing method of any embodiment of the present invention.

[0204] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0205] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0206] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized in that, include: Obtain the second object attribute information for each second type object in the pre-created target object set; wherein, the second object attribute information includes: the number of first type objects contained in each second type object, and the first object attribute information corresponding to each first type object; The similarity and coverage between every two objects of the second type are determined based on the attribute information of the second object. The second type of objects are merged based on the similarity and the coverage to generate a corresponding object association table; wherein, the object association table refers to the association table between the merged second type of objects and the first type of objects; Identify and extract a first set of object identifiers for the target entity types associated with the target financial institution in a pre-created financial data table; wherein, the first set of object identifiers refers to a group of enterprises whose target entity types are single legal entities with credit lines; Identify and extract the second type of target object associated with the target financial institution in the risk data table and the second object identifier set contained in the second type of target object; wherein, the second type of target object is the group recorded in the risk data table, and the second object identifier set refers to the enterprise belonging to the group; Identify and locate the third type target object containing the second type target object and the third object identifier set contained in the third type target object in the pre-created object association table; wherein, the third object identifier set refers to the set generated by summarizing the first type objects contained in the third type target object, and the third type target object refers to the object generated by merging the attribute information of the second type target object; Determine the matching situation between the identifier of the first type of target object in the first object identifier set and the identifier of the first type of object in the third object identifier set, and take it as the first matching situation; Determine the matching situation between the identifier of the first type of target object in the first object identifier set and the identifier of the first type of object in the second object identifier set, and use it as the second matching situation; If the first matching condition is a non-match or the second matching condition is a match, then the data supervision quality is determined to be good. If the first matching condition is a match or the second matching condition is a non-match, then the data supervision quality is determined to be poor. The step of determining the similarity and coverage between every two objects of the second type based on the second object attribute information includes: Based on the number of first-type objects contained in the second-type objects, the second-type objects in the target object set are classified to obtain the corresponding object groups; The similarity between every two second-type objects in the same object group is determined based on the number of first-type objects contained in the second-type object. The second type of objects whose similarity reaches a preset similarity threshold are merged to obtain the corresponding third type of initial object; different object groups correspond to different preset similarity thresholds; The coverage ratio between each pair of the third type initial objects and the second type objects, or between each pair of the second type objects, is determined based on the number of first type objects contained in the third type initial objects and / or the second type objects.

2. The method according to claim 1, characterized in that, The step of classifying the second type of objects in the target object set based on the number of first type objects contained in the second type of objects to obtain corresponding object groups includes: Determine the comparison result between the number of first-type objects contained in the second-type object and a pre-configured threshold for the number of each object; Based on the comparison results, the second type of objects in the target object set are classified to obtain the corresponding object groups.

3. The method according to claim 1, characterized in that, Determining the similarity between every two second-type objects in the same object group based on the number of first-type objects contained in the second-type objects includes: The number of identical first-type objects contained in every two second-type objects within the same object group is determined as the first quantity; Determine the number of first-type objects contained in each pair of second-type objects within the same object group, and use these as the second quantity and the third quantity; The similarity between every two second-type objects in the same object group is determined based on the maximum value of the second and third quantities, and the first quantity.

4. The method according to claim 1, characterized in that, The step of determining the coverage between every two third-type initial objects from different object groups based on the number of first-type objects contained in the third-type initial object includes: The number of identical first-type objects contained in every two third-type initial objects from different object groups is determined as the fourth quantity; The number of first-type objects contained in each of every two third-type initial objects from different object groups is determined as the fifth and sixth quantities; The coverage between every two of the third type of initial objects from different object groups is determined based on the minimum of the fifth and sixth quantities, and the fourth quantity.

5. The method according to claim 1, characterized in that, The step of determining the coverage ratio between the third type initial object and the second type object from different object groups based on the number of first type objects contained in the third type initial object and the second type object includes: The number of identical first-type objects contained in the third-type initial objects and the second-type objects from different object groups is determined as the seventh quantity; The number of first-type objects contained in the third-type initial objects and the second-type objects from different object groups are determined as the eighth and ninth numbers, respectively. The coverage ratio between the third type of initial objects and the second type of objects from different object groups is determined based on the minimum of the eighth and ninth quantities, and the seventh quantity.

6. The method according to claim 1, characterized in that, Determining the coverage between every two second-type objects from different object groups based on the number of first-type objects contained in the second-type objects includes: Determine the number of identical first-type objects contained in every two second-type objects from different object groups, as the tenth number; Determine the number of first-type objects contained in each pair of second-type objects from different object groups, as the eleventh and twelfth quantities; The coverage between every two second-type objects from different object groups is determined based on the minimum of the eleventh and twelfth quantities, and the tenth quantity.

7. The method according to claim 1, characterized in that, Before obtaining the second object attribute information of each second type object in the pre-created target object set, the method further includes: Get the pre-created initial object collection; Perform a culling operation based on the second object attribute information of each second type object in the initial object set to obtain the corresponding target object set.

8. The method according to claim 7, characterized in that, The step of performing a culling operation based on the second object attribute information of each second type object in the initial object set to obtain the corresponding target object set includes: Obtain the name of the second object corresponding to each second type object in the initial object set, and the identifier of the first object corresponding to the first type object contained therein; The second type of object, whose second object name is the target name and whose first object identifier is empty, is removed from the initial object set to obtain the corresponding target object set.

9. A data processing apparatus, characterized in that, include: The second object attribute information acquisition module is used to acquire the second object attribute information of each second type object in the pre-created target object set; wherein, the second object attribute information includes: the number of first type objects contained in each second type object, and the first object attribute information corresponding to each first type object; The similarity and coverage determination module is used to determine the similarity and coverage between every two objects of the second type based on the attribute information of the second object; merge the objects of the second type based on the similarity and coverage to generate a corresponding object association table; wherein, the object association table refers to the association table between the merged objects of the second type and the objects of the first type; The first set determination module is used to identify and extract the first object identifier set of the target entity type associated with the target financial institution in the pre-created financial data table; wherein, the first object identifier set refers to the collective of enterprises whose target entity type is a single legal person with credit. The second set determination module is used to identify and extract the second type of target objects associated with the target financial institution in the risk data table and the second object identifier set contained in the second type of target objects; wherein, the second type of target objects are the groups recorded in the risk data table, and the second object identifier set refers to the enterprises belonging to the groups; The third set determination module is used to identify and search for the third type target object where the second type target object is located and the third object identifier set contained in the third type target object in a pre-created object association table; wherein, the third object identifier set refers to the set generated by summarizing the first type objects contained in the third type target object, and the third type target object refers to the object generated by merging the attribute information of the second type target object; The data quality determination module includes: The first matching situation determination unit is used to determine the matching situation between the identifier of the first type target object in the first object identifier set and the identifier of the first type object in the third object identifier set, as the first matching situation; The second matching situation determination unit is used to determine the matching situation between the identifier of the first type target object in the first object identifier set and the identifier of the first type object in the second object identifier set, as the second matching situation; The data quality determination unit is used to determine that the data supervision quality is good when the first matching condition is a mismatch or the second matching condition is a match; and to determine that the data supervision quality is poor when the first matching condition is a match or the second matching condition is a mismatch. The similarity coverage determination module includes: An object group generation unit is used to classify the second type of objects in the target object set based on the number of first type objects contained in the second type of objects, so as to obtain the corresponding object groups; A similarity determination unit is used to determine the similarity between every two second-type objects in the same object group based on the number of first-type objects contained in the second-type object; The third type of initial object generation unit is used to merge the second type objects whose similarity reaches a preset similarity threshold to obtain the corresponding third type of initial object; different object groups correspond to different preset similarity thresholds; The coverage determination unit is configured to determine the coverage between every two third-type initial objects from different object groups and the second-type objects, or between every two second-type objects, based on the number of first-type objects contained in the third-type initial objects and / or the second-type objects.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data processing method according to any one of claims 1-8.

12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data processing method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Supervision data quality verification method and device, electronic equipment and storage medium

    CN112597165A