Data processing method, device and equipment and computer readable storage medium

By diffusing and merging the relationship data of abnormal groups, related abnormal groups are identified and merged, which solves the problems of low audit efficiency and low resource utilization in the existing technology and achieves more efficient audit and resource allocation.

CN120765375APending Publication Date: 2025-10-10腾安基金销售(深圳)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410383099.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies are unable to efficiently review a large number of abnormal object identifiers, resulting in low resource utilization.

Method used

By obtaining the relationship data of multiple abnormal groups, diffusion processing and merging processing are performed, abnormal diffusion groups with associated relationships are identified and merged to form larger-scale abnormal merged groups.

Benefits of technology

It improves the review efficiency of abnormal object identification, rationally allocates resources, and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765375A_ABST
    Figure CN120765375A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring a plurality of abnormal groups; obtaining relational data comprising a source object identifier and an associated object identifier; performing diffusion processing on the plurality of abnormal groups through the relational data to obtain a plurality of abnormal diffusion groups; the diffusion object identifier included in one abnormal diffusion group belongs to the associated object identifier; performing object identifier identification processing on every two abnormal diffusion groups to obtain an abnormal diffusion group with the same diffusion object identifier; and merging the two abnormal diffusion groups in the abnormal diffusion group according to the same diffusion object identifiers contained in the abnormal diffusion group to obtain an abnormal merged group. By adopting the method and the device, the auditing efficiency of the abnormal object identifier can be improved, and the resource utilization rate can also be improved. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a data processing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] In existing technologies, business transactions of payment accounts bound to object identifiers are inspected to determine the object identifier's identification attributes. For example, if a business transaction is abnormal, the object identifier has abnormal attributes; if a business transaction is normal, the object identifier has normal attributes. However, with the increasing number of abnormal object identifiers, existing technologies are unable to efficiently review a large number of abnormal object identifiers. Furthermore, by reviewing each abnormal object identifier, audit resources cannot be effectively allocated, thereby reducing resource utilization. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, apparatus, device, and computer-readable storage medium, which can not only improve the review efficiency of abnormal object identification, but also improve resource utilization.

[0004] On the one hand, an embodiment of the present application provides a data processing method, including:

[0005] Acquire multiple abnormal groups; the multiple abnormal groups all refer to groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other;

[0006] Acquire relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier;

[0007] Through relational data, diffusion processing is performed on multiple abnormal groups respectively to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by diffusion processing an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group belong to the associated object identifiers;

[0008] Performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups;

[0009] According to the same diffusion object identifier contained in the abnormal diffusion group, two abnormal diffusion groups in the abnormal diffusion group are merged to obtain an abnormal merged group.

[0010] In one aspect, an embodiment of the present application provides a data processing device, including:

[0011] An acquisition module is used to acquire multiple abnormal groups; the multiple abnormal groups are groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other;

[0012] The acquisition module is further used to acquire relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier;

[0013] The processing module is used to perform diffusion processing on multiple abnormal groups respectively through relational data to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group are associated object identifiers;

[0014] The processing module is further configured to perform object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups;

[0015] The processing module is further configured to merge two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group.

[0016] In a possible implementation, the plurality of abnormal groups include abnormal group E f , f is a positive integer, and f is less than or equal to the total number of multiple abnormal groups; abnormal group E f Including the exception object identifier G h , h is a positive integer, and h is less than or equal to the abnormal population E f The total number of abnormal object identifiers in the ; the total number of source object identifiers and the total number of associated object identifiers are both Z, where Z is a positive integer; there is an association relationship between a source object identifier and an associated object identifier;

[0017] The processing module performs diffusion processing on multiple abnormal groups through relational data to obtain multiple abnormal diffusion groups, which are used to perform the following operations:

[0018] Identify the abnormal object G h Match with Z source object identifiers. If there is an exception object identifier G in the Z source object identifiers, h If the source object ID is the same, then among the Z associated object IDs, obtain the one with the exception object ID G h The identifier of the associated object with which the association relationship exists;

[0019] will be identified with the exception object G hThe associated object ID with an associated relationship is determined as the abnormal object ID G h The diffusion object identifier;

[0020] Identify the abnormal object G h The diffusion object identifier is added to the abnormal group E f , will add the abnormal object identifier G h The abnormal group E identified by the diffusion object f , identified as abnormal group E f The corresponding abnormal diffusion group.

[0021] In a possible implementation, the multiple abnormal diffusion groups include a first abnormal diffusion group and a second abnormal diffusion group;

[0022] The processing module performs object identification processing on every two abnormal diffusion groups in the multiple abnormal diffusion groups to obtain an abnormal diffusion group group with the same diffusion object identification, which is used to perform the following operations:

[0023] comparing the diffusion object identifiers in the first abnormal diffusion group with the diffusion object identifiers in the second abnormal diffusion group;

[0024] If the first abnormal diffusion group and the second abnormal diffusion group include the same diffusion object identifier, the first abnormal diffusion group and the second abnormal diffusion group are determined as an abnormal diffusion group group;

[0025] The same diffusion object identifiers included in the first abnormal diffusion group and the second abnormal diffusion group are determined as the same diffusion object identifiers of the abnormal diffusion group group.

[0026] In a possible implementation, the plurality of abnormal diffusion groups include Y object identifiers; Y is a positive integer greater than 1;

[0027] The data processing device further includes:

[0028] A determination module, configured to determine identification scores corresponding to each of the Y object identifications;

[0029] In a possible implementation, the processing module merges two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group for performing the following operations:

[0030] Determine the same diffusion object identifiers contained in the abnormal diffusion group as the common object identifier of the abnormal diffusion group;

[0031] Among the Y identification scores, obtain the public identification score of the public object identification;

[0032] According to the public identification score, two abnormal diffusion groups in the abnormal diffusion group are merged to obtain an abnormal merged group.

[0033] In a possible implementation, the processing module merges two abnormal diffusion groups in the abnormal diffusion group according to the common identification score to obtain an abnormal merged group, which is used to perform the following operations:

[0034] Perform mean processing on the public identification scores to obtain the public scores corresponding to the abnormal diffusion group;

[0035] Compare the common score with the merge score threshold. If the common score is equal to or greater than the merge score threshold, merge the two abnormal diffusion groups in the abnormal diffusion group to obtain an abnormal merged diffusion group.

[0036] The diffusion object identifier in the abnormal merged diffusion group is deleted, and the abnormal merged diffusion group after the diffusion object identifier is deleted is determined as the abnormal merged group.

[0037] In a possible implementation, the total number of abnormal diffusion group groups is at least two, the at least two abnormal diffusion group groups include a first abnormal diffusion group group and a second abnormal diffusion group group different from the first abnormal diffusion group group; the common score includes a first common score corresponding to the first abnormal diffusion group group and a second common score corresponding to the second abnormal diffusion group group;

[0038] If the common score is equal to or greater than the merge score threshold, the processing module merges the two abnormal diffusion groups in the abnormal diffusion group to obtain an abnormal merged diffusion group, which is used to perform the following operations:

[0039] If the first common score is equal to or greater than the merge score threshold, and the second common score is equal to or greater than the merge score threshold, comparing the two abnormal diffusion groups in the first abnormal diffusion group and the two abnormal diffusion groups in the second abnormal diffusion group;

[0040] If there is a same abnormal diffusion group in the first abnormal diffusion group and the second abnormal diffusion group, the two abnormal diffusion groups in the first abnormal diffusion group and the two abnormal diffusion groups in the second abnormal diffusion group are merged to obtain an abnormal merged diffusion group.

[0041] In a possible implementation, the determination module determines identification scores corresponding to the Y object identifications, respectively, to perform the following operations:

[0042] Construct an original topological map based on multiple abnormal groups;

[0043] According to the multiple abnormal diffusion groups, the original topology map is diffused to obtain a diffusion topology map; the diffusion topology map includes Y nodes; one node is used to represent one object identifier among the Y object identifiers;

[0044] Perform node score processing on the diffusion topology graph to obtain the node scores corresponding to Y nodes;

[0045] The Y node scores are determined as the identification scores corresponding to the Y object identifications respectively.

[0046] In a possible implementation, the plurality of abnormal groups include Z abnormal object identifiers; the Z abnormal object identifiers belong to Y object identifiers; Z is a positive integer greater than 1;

[0047] The determination module constructs an original topology map based on multiple abnormal groups, which is used to perform the following operations:

[0048] Generate Z nodes for representing Z abnormal object identifiers; wherein one node is used to represent one abnormal object identifier;

[0049] Determine an edge score between every two nodes in the Z nodes according to an association relationship between every two abnormal object identifiers in the Z abnormal object identifiers;

[0050] Construct the original topology graph based on Z nodes and the edge scores between every two nodes.

[0051] In a possible implementation, the Z abnormal object identifiers include a first abnormal object identifier and a second abnormal object identifier different from the first abnormal object identifier;

[0052] The determination module determines the edge score between every two nodes in the Z nodes based on the association relationship between every two abnormal object identifiers in the Z abnormal object identifiers, and is used to perform the following operations:

[0053] Obtaining first registration information corresponding to the first abnormal object identifier and second registration information corresponding to the second abnormal object identifier;

[0054] Performing similarity processing on the first registration information and the second registration information to obtain a similarity value between the first registration information and the second registration information;

[0055] The similarity value between the first registration information and the second registration information is determined as the edge score between the first node and the second node; the first node is used to represent the first abnormal object identifier; the second node is used to represent the second abnormal object identifier; the first node and the second node both belong to Z nodes.

[0056] In a possible implementation, the relationship data further includes an association weight for characterizing the degree of association between the source object identifier and the associated object identifier; the plurality of abnormal diffusion groups include the abnormal object identifier I j , and the exception object identifier I j Diffusion object identifier K j ; j is a positive integer, and j is less than or equal to the total number of diffusion object identifiers in multiple abnormal diffusion groups; diffusion object identifier K j Belongs to the associated object identifier; abnormal object identifier I j Belongs to the source object identifier;

[0057] The determination module performs diffusion processing on the original topology map based on multiple abnormal diffusion groups to obtain a diffusion topology map, which is used to perform the following operations:

[0058] Generate the identifier K for characterizing the diffusion object j The third node of

[0059] In relational data, obtain the identifier I used to characterize the abnormal object j And the diffusion object identifier K j The association weight of the degree of association between them;

[0060] Identify the exception object as I j And the diffusion object identifier K j The association weight between them is determined as the edge score between the third node and the fourth node; the fourth node refers to the identifier I used to characterize the abnormal object in the original topology graph. j Node;

[0061] The third node and the edge score are added to the original topology graph, and the original topology graph with the third node and the edge score added is determined as a diffusion topology graph.

[0062] In a possible implementation, the data processing apparatus further includes:

[0063] The acquisition module is further configured to acquire first detailed data corresponding to a plurality of resource management identifiers, and merge first detailed data with the same object identifier among the plurality of first detailed data to obtain a plurality of second detailed data; wherein the object identifiers corresponding to the plurality of second detailed data are different from each other;

[0064] The processing module is further configured to divide the plurality of object identifiers into A groups according to the object names respectively included in the plurality of second detailed data, wherein A is a positive integer; and the object identifiers respectively included in the A groups are different from each other;

[0065] The processing module is further used to perform identification processing on the A groups in multiple information dimensions to obtain group attributes corresponding to the A groups;

[0066] The determination module is used to determine, among the A groups, a group with a group attribute that is an abnormal group attribute as an abnormal group.

[0067] In a possible implementation, the processing module divides the multiple object identifiers according to the object names respectively included in the multiple second detailed data to obtain A groups, which are used to perform the following operations:

[0068] obtaining a text recognition model, and inputting the object names respectively included in the plurality of second detailed data into the text recognition model;

[0069] Through the text recognition model, feature extraction is performed on multiple object names respectively to obtain multiple text vectors; wherein, a text vector is obtained by extracting features from an object name through the text recognition model;

[0070] Obtain a vector clustering model, input multiple text vectors into the vector clustering model respectively, and cluster the multiple text vectors through the vector clustering model to obtain A text vector clusters;

[0071] According to A text vector clusters, multiple object identifiers are divided and processed to obtain A groups; wherein, the text vector corresponding to the object identifier in a group belongs to a text vector cluster.

[0072] In one possible implementation, group A includes group B c , c is a positive integer, and c is less than or equal to A;

[0073] The processing module identifies and processes A groups separately in multiple information dimensions to obtain group attributes corresponding to A groups, which are used to perform the following operations:

[0074] In multiple information dimensions, group B c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions;

[0075] If there are abnormal recognition results among multiple recognition results, the abnormal group attribute is determined to be group B c Corresponding group attributes;

[0076] If multiple recognition results are all normal recognition results, the normal group attribute is determined to be group B c The corresponding group attributes.

[0077] In one possible implementation, the plurality of information dimensions include a name information dimension, and the plurality of recognition results include a name recognition result;

[0078] The processing module processes group B in multiple information dimensions. cPerform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0079] Among the plurality of second detailed data, obtain group B c The object name corresponding to the object identifier in the object name is determined as an object name set;

[0080] Get the reference name with the exception attribute, match the reference name with the object name in the object name collection;

[0081] If there is an object name in the object name set that successfully matches the reference name, the abnormal recognition result is determined as the name recognition result;

[0082] If there is no object name in the object name set that successfully matches the reference name, the normal recognition result is determined as the name recognition result.

[0083] In one possible implementation, the plurality of information dimensions include a registration information dimension, and the plurality of recognition results include a registration recognition result;

[0084] The processing module processes group B in multiple information dimensions. c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0085] Among the plurality of second detailed data, obtain group B c The registration information corresponding to the object identifier in the object identifier is determined as a registration information set;

[0086] In the registration information set, obtaining a maximum number of identical registration information, and determining a ratio between the maximum number and the total number of registration information in the registration information set as an abnormal information ratio;

[0087] Compare the abnormal information ratio with the abnormal information ratio threshold. If the abnormal information ratio is greater than the abnormal information ratio threshold, the abnormal recognition result is determined as the registered recognition result.

[0088] If the abnormal information ratio is equal to or less than the abnormal information ratio threshold, the normal recognition result is determined as the registered recognition result.

[0089] In one possible implementation, the multiple information dimensions include a business information dimension, and the multiple identification results include a business identification result;

[0090] The processing module processes group B in multiple information dimensions. c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0091] Among the plurality of second detailed data, obtain group B c The business scope corresponding to the object identifier in the object is determined as a business scope set;

[0092] Perform similarity processing on every two business scopes in the business scope set to obtain D similarities corresponding to the business scope set; D is a positive integer;

[0093] Determine the mean similarity value corresponding to the D similarities, and compare the mean similarity value with the similarity threshold;

[0094] If the similarity mean is equal to or greater than the similarity threshold, the anomaly identification result is determined as the business identification result;

[0095] If the similarity mean is less than the similarity threshold, the normal recognition result is determined as the business recognition result.

[0096] On one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0097] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface, wherein the above-mentioned network interface is used to provide data communication function, the above-mentioned memory is used to store computer programs, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.

[0098] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded by a processor and executing the method in the embodiment of the present application.

[0099] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.

[0100] In the embodiment of the present application, the computer device acquires a plurality of abnormal groups; each of the plurality of abnormal groups refers to a group comprising an abnormal object identifier; the abnormal object identifiers comprised by the plurality of abnormal groups are different from each other; relationship data comprising a source object identifier and an associated object identifier are acquired; the source object identifier belongs to the abnormal object identifiers comprised by the plurality of abnormal groups; the associated object identifier refers to an object identifier having an association relationship with the source object identifier; through the relationship data, the plurality of abnormal groups can be respectively subjected to diffusion processing to obtain a plurality of abnormal diffusion groups; the diffusion object identifiers comprised by one abnormal diffusion group belong to the associated object identifiers; each of the plurality of abnormal diffusion groups can be subjected to object identifier identification processing to obtain an abnormal diffusion group group having the same diffusion object identifier; the abnormal diffusion group group comprises two different abnormal diffusion groups; according to the same diffusion object identifier contained by the abnormal diffusion group group, the two abnormal diffusion groups in the abnormal diffusion group group can be subjected to merging processing to obtain an abnormal merging group. As can be seen from the above, through group diffusion, the abnormal groups having an association relationship can be identified, and through merging the abnormal groups having an association relationship, small groups can be converted into large groups, so that the auditing efficiency of the abnormal object identifier can be improved, and through determining larger-scale abnormal groups, large-scale abnormal groups can be processed preferentially, so that the auditing resources can be more reasonably allocated, and thus the resource utilization rate can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0101] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0102] Figure 1 is a system architecture schematic diagram provided by an embodiment of the present application;

[0103] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;

[0104] Figure 3 is a data processing scene schematic diagram provided by an embodiment of the present application Figure 1 ;

[0105] Figure 4 is a data processing scene schematic diagram provided by an embodiment of the present application Figure 2 ;

[0106] Figure 5 is a flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;

[0107] Figure 6 This is a data processing scenario provided by the embodiment of the present application. Figure 3 ;

[0108] Figure 7 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 3 ;

[0109] Figure 8 This is a data processing scenario provided by the embodiment of the present application. Figure 4 ;

[0110] Figure 9 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0111] Figure 10 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0112] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0113] See Figure 1 , Figure 1 This is a schematic diagram of a system architecture provided by an embodiment of the present application. Figure 1 As shown, the system may include a service server 100 and a terminal device cluster. The terminal device cluster may include: terminal device 200a, terminal device 200b, terminal device 200c, ..., terminal device 200n. It is understandable that the above system may include one or more terminal devices, and this application does not limit the number of terminal devices.

[0114] Among them, there can be communication connections between terminal device clusters, for example, there is a communication connection between terminal device 200a and terminal device 200b, and there is a communication connection between terminal device 200a and terminal device 200c. At the same time, any terminal device in the terminal device cluster can have a communication connection with the business server 100, for example, there is a communication connection between terminal device 200a and business server 100. The above-mentioned communication connection is not limited to the connection method, and can be directly or indirectly connected through wired communication, directly or indirectly connected through wireless communication, or through other methods, and this application does not impose any restrictions on this.

[0115] like Figure 1 Each terminal device in the terminal device cluster shown can run an application management client for a business application. The business application can be a video application, payment application, navigation application, music application, shopping application, electronic map application, browser, or other application client with resource transfer capabilities. The business application can be a standalone client or an embedded sub-client integrated into a client (for example, a video client or a travel client), without limitation.

[0116] The application management client for business applications is intended for end users such as application developers and other application managers. Through the application management client, the application manager can trigger operations to identify the identification attributes of object identifiers. Server 100, which can be an application management server for business applications, is used to execute the data processing method provided in embodiments of the present application. This method can not only identify group attributes of groups including object identifiers, but also merge and process multiple abnormal groups that have associated relationships.

[0117] It is understandable that in the specific implementation of this application, related data such as user information (such as an object identifier bound to one or more resource management identifiers, and second detailed data) are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of the relevant regions.

[0118] To facilitate subsequent understanding and explanation, the embodiments of the present application can be Figure 1 A terminal device example is selected from the terminal device cluster shown for description, for example, terminal device 200a is used as an example for description. When the application management client receives a group merging instruction for multiple abnormal groups, the terminal device 200a can generate a group merging request for the multiple abnormal groups and send the group merging request to the business server 100. Among them, multiple abnormal groups all refer to groups including abnormal object identifiers, and the abnormal object identifiers included in the multiple abnormal groups are different from each other. For example, multiple abnormal groups include abnormal group 1 and abnormal group 2, abnormal group 1 includes object identifier 1-object identifier 100, and abnormal group 2 includes object identifier 101-object identifier 150, wherein object identifier 1-object identifier 150 are all abnormal object identifiers. An abnormal group can be understood as a group with abnormal attributes, and an abnormal object identifier refers to an object identifier with abnormal attributes; an object identifier refers to the unique identification information of an object in a business application, and its composition can be set according to the actual application scenario, for example, a string of numbers, or a combination of a string of numbers and letters.

[0119] The business server 100 obtains the group merging request sent by the terminal device 200a, and obtains multiple abnormal groups according to the group merging request. The embodiment of the present application does not limit the way in which the business server 100 obtains multiple abnormal groups, and can be set according to the actual application scenario. One feasible way to obtain is that the terminal device 200a generates a group merging request including multiple abnormal groups, so the business server 100 can obtain multiple abnormal groups from the group merging request; another feasible way to obtain is that the business server 100 obtains multiple abnormal groups from the local database according to the group merging request; another feasible way to obtain is that the group merging request sent by the terminal device 200a carries the storage addresses of multiple abnormal groups, so the business server 100 can obtain multiple abnormal groups according to the storage addresses.

[0120] The business server 100 obtains relationship data including a source object identifier and an associated object identifier; the source object identifier belongs to an abnormal object identifier included in multiple abnormal groups; the associated object identifier refers to an object identifier that has an associated relationship with the source object identifier. For example, multiple abnormal groups include object identifiers 1-object identifier 150, and the number of relationship data is 2, namely relationship data 1 and relationship data 2; relationship data 1 includes object identifier 1 (which is a source object identifier) ​​and object identifier a (which is an associated object identifier) ​​that has an associated relationship with object identifier 1, and relationship data 2 includes object identifier 2 (which is a source object identifier) ​​and object identifier b (which is an associated object identifier) ​​that has an associated relationship with object identifier 2. The embodiment of the present application does not limit the number of relationship data, which should be set according to the actual application scenario and can be one or more.

[0121] Using the relational data, the business server 100 can perform diffusion processing on multiple abnormal groups separately to obtain multiple abnormal diffusion groups. An abnormal diffusion group is obtained by performing diffusion processing on an abnormal group using the relational data, and the diffusion object identifiers included in an abnormal diffusion group are associated object identifiers. In this embodiment of the present application, the associated object identifier added to the abnormal group is referred to as the diffusion object identifier of the abnormal group, and the abnormal group to which the diffusion object identifier is added is referred to as the abnormal diffusion group.

[0122] Furthermore, the business server 100 performs object identification processing on every two abnormal diffusion groups in the multiple abnormal diffusion groups, and can obtain an abnormal diffusion group group with the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups; for example, abnormal diffusion group 1 includes diffusion object identification 1, and abnormal diffusion group 2 includes diffusion object identification 1, then the business server 100 can combine abnormal diffusion group 1 and abnormal diffusion group 2 into an abnormal diffusion group group.

[0123] According to the same diffusion object identifier contained in the abnormal diffusion group, the business server 100 merges two abnormal diffusion groups in the abnormal diffusion group to obtain an abnormal merged group.

[0124] Subsequently, the service server 100 sends the abnormal merged groups corresponding to the multiple abnormal groups to the terminal device 200 a , and the terminal device 200 a may display the abnormal merged groups on its corresponding screen.

[0125] Optionally, the business server 100 returns the abnormal diffusion group to the terminal device 200a, so the terminal device 200a can merge two abnormal diffusion groups in the abnormal diffusion group according to the same diffusion object identifier contained in the abnormal diffusion group to obtain an abnormal merged group.

[0126] Optionally, the business server 100 can return multiple abnormal diffusion groups to the terminal device 200a, so the terminal device 200a can perform object identification processing on every two abnormal diffusion groups in the multiple abnormal diffusion groups to obtain an abnormal diffusion group group with the same diffusion object identification; the subsequent processing is the same as the above description, so it will not be repeated.

[0127] Optionally, the business server 100 may return the relationship data to the terminal device 200a, so the terminal device 200a may perform diffusion processing on multiple abnormal groups respectively through the relationship data to obtain multiple abnormal diffusion groups; the subsequent processing is the same as the above description, so it will not be repeated.

[0128] Optionally, if the terminal device 200a holds relational data, the terminal device 200a can perform diffusion processing on multiple abnormal groups separately through the relational data when obtaining the group merging instruction to obtain multiple abnormal diffusion groups; the subsequent processing is the same as the above description, so it is not repeated.

[0129] From the above, it can be seen that the embodiment of the present application can identify abnormal groups with related relationships through group diffusion, and can transform small groups into large groups by merging abnormal groups with related relationships, so the audit efficiency of abnormal object identification can be improved. By determining larger-scale abnormal groups, large-scale abnormal groups can be processed first, so audit resources can be allocated more reasonably, thereby improving resource utilization.

[0130] It should be noted that the above-mentioned business server 100, terminal device 200a, terminal device 200b, terminal device 200c..., terminal device 200n can all be blockchain nodes in the blockchain network, and the data described in the full text (for example, an object identifier bound to one or more resource management identifiers, and the second detailed data) can be stored. The storage method can be a method in which the blockchain node generates a block based on the data and adds the block to the blockchain for storage.

[0131] Blockchain is a novel application model of computer technologies, integrating distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. It primarily organizes data in chronological order and encrypts it into a ledger, rendering it tamper-proof and forgery-proof. It also enables data verification, storage, and updates. Blockchain is essentially a decentralized database, where each node stores an identical blockchain. Blockchain networks categorize nodes as core nodes, data nodes, and light nodes. Core nodes, data nodes, and light nodes collectively constitute blockchain nodes. Core nodes are responsible for consensus across the entire blockchain network, effectively serving as consensus nodes within the blockchain network.

[0132] The process of writing transaction data in the blockchain network into the ledger can be as follows: the data node or light node in the blockchain network obtains the transaction data and passes the transaction data in the blockchain network (that is, the node passes the transaction data in a relay manner) until the consensus node receives the transaction data. The consensus node then packages the transaction data into a block, performs consensus on the block, and writes the transaction data into the ledger after the consensus is completed. Here, an object identifier bound to one or more resource management identifiers and second detailed data are used as an example of transaction data. After reaching a consensus on the transaction data, the business server 100 (blockchain node) generates a block based on the transaction data and stores the block in the blockchain network. As for reading the transaction data (that is, the object identifier bound to one or more resource management identifiers and the second detailed data), the blockchain node can obtain the block containing the transaction data in the blockchain network and further obtain the transaction data in the block.

[0133] It is understandable that the method provided in the embodiments of the present application can be executed by a computer device, including but not limited to a terminal device or a business server. Among them, the business server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Among them, the terminal device and the business server can be directly or indirectly connected by wired or wireless means, and the embodiments of the present application are not limited here.

[0134] Further, see Figure 2 , Figure 2 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 1 The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, autonomous driving, etc. The embodiments of the present application can be applied to business scenarios such as resource management identification scenarios, object identification scenarios, object identification scenarios, resource management identification review scenarios, object identification review scenarios, and object review scenarios. Specific business scenarios will not be listed one by one here.

[0135] The data processing method can be performed by a business server (for example, Figure 1 The service server 100 shown in FIG. 1 may also be executed by a terminal device (for example, the above Figure 1 The terminal device 200a shown in the figure can also be executed by the business server and the terminal device interactively. For ease of understanding, the embodiment of the present application takes the method executed by the business server as an example for explanation, that is, the business server executes the data processing method as a computer device. Figure 2 As shown, the data processing method may at least include the following steps S101 to S105.

[0136] Step S101 , obtaining multiple abnormal groups; the multiple abnormal groups all refer to groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other.

[0137] Specifically, an abnormal group refers to a group whose object identifiers are all abnormal object identifiers. The embodiment of this application does not limit the number of abnormal object identifiers in the abnormal group. It can be one or more and should be set according to the actual application scenario. If an abnormal group includes multiple abnormal object identifiers, then the multiple abnormal object identifiers have similarities. For details, please see below Figure 7 Description in the corresponding embodiment.

[0138] An object identifier refers to information that can uniquely identify an object. For example, if the object is a merchant, the object identifier can be the merchant identifier in the business application. The resource management identifier described below refers to the merchant account in the business application. A merchant can apply for one or more merchant accounts using one merchant identifier, that is, one object identifier can be bound to one or more resource management identifiers.

[0139] Please also see Figure 3 , Figure 3 This is a data processing scenario provided by the embodiment of the present application. Figure 1 . Figure 3 The example plurality of abnormal populations include an abnormal population 31 a , an abnormal population 32 a , and an abnormal population 33 a . Figure 3 The abnormal object identification in the abnormal group is abbreviated as identification. The abnormal group 31a includes 4 abnormal object identifications, Figure 3 Examples are identification 1, identification 2, identification 3, and identification 4; the abnormal group 32a includes 3 abnormal object identifications, Figure 3 Examples are identification 5, identification 6, and identification 7; the abnormal group 33a includes 2 abnormal object identifications, Figure 3 Examples are ID 8 and ID 9.

[0140] Step S102 , acquiring relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier.

[0141] Specifically, the relationship data also includes an association weight used to characterize the degree of association between the source object identifier and the associated object identifier, and the degree of association between the source object identifier and the associated object identifier is proportional to the association weight, that is, the greater the degree of association, the greater the association weight, and the smaller the degree of association, the smaller the association weight.

[0142] It is understandable that the source object identifier and the associated object identifier can be regarded as a group of object identifier pairs. The embodiment of the present application does not limit the number of object identifier pairs in the relationship data, and can include one or more groups of object identifier pairs. In actual applications, the business server can use all abnormal object identifiers in multiple abnormal groups as candidate source object identifiers. If a candidate source object identifier has an object identifier with which it has an association relationship (different from the candidate source object identifier), the candidate source object identifier is determined as the source object identifier, and the business server determines the object identifier with which the source object identifier has an association relationship as the associated object identifier of the source object identifier.

[0143] For the convenience of description and understanding, Figure 3The example identifier 1 is described as the source object identifier. The business server obtains the resource management identifier applied for through identifier 1. For example, there are 3 resource management identifiers corresponding to identifier 1, and obtains the registration information corresponding to the 3 resource management identifiers respectively; the business server obtains the resource management identifier applied for through identifier a (which can be any object identifier different from identifier 1). For example, there are 4 resource management identifiers corresponding to identifier 2, and obtains the registration information corresponding to the 4 resource management identifiers respectively. The business server matches the 3 registration information corresponding to identifier 1 with the 4 registration information corresponding to identifier 2.

[0144] For example, if the registration information includes contact information, region, and bank account number, the specific matching process can be as follows: the business server matches the three contact information corresponding to identifier 1 (the three contact information can be exactly the same or different) with the four contact information corresponding to identifier 2. Assuming that there are three resource management identifiers with the same contact information between identifier 1 and identifier 2, then in the contact information dimension, the association weight between identifier 1 and identifier 2 can be set to 3; the business server matches the three regions corresponding to identifier 1 with the four regions corresponding to identifier 2. Assuming that there are two resource management identifiers with the same region between identifier 1 and identifier 2, then in the region dimension, the association weight between identifier 1 and identifier 2 can be set to 2; the business server matches the three bank account numbers corresponding to identifier 1 with the four bank account numbers corresponding to identifier 2. Assuming that there are three resource management identifiers with the same bank account number between identifier 1 and identifier 2, then in the bank account dimension, the association weight between identifier 1 and identifier 2 can be set to 3; combined with the above three information dimensions, the business server can determine that the association weight between identifier 1 and identifier 2 is 3+2+3=8.

[0145] The above description uses registration information to determine the association weight between two object identifiers. In actual application, the association weight between two object identifiers can also be determined based on a direct transaction between the two object identifiers, or based on a common transaction object identifier between the two object identifiers. The embodiments of this application do not limit the method for determining the association relationship between the source object identifier and the associated object identifier, and can be set according to the needs of the actual application scenario.

[0146] Please see again Figure 3 , Figure 3The example relationship data 30b includes 6 object identification pairs, the object identification pair with serial number 1 includes the source object identification generated by identification 1 and the associated object identification generated by identification a, and the association weight between identification 1 and identification a is 6; the object identification pair with serial number 2 includes the source object identification generated by identification 1 and the associated object identification generated by identification b, and the association weight between identification 1 and identification b is 7; the object identification pair with serial number 3 includes the source object identification generated by identification 5 and the associated object identification generated by identification a, and the association weight between identification 5 and identification a is 4; the object identification pair with serial number 4 includes the source object identification generated by identification 7 and the associated object identification generated by identification b, and the association weight between identification 7 and identification b is 3; the object identification pair with serial number 5 includes the source object identification generated by identification 8 and the associated object identification generated by identification c, and the association weight between identification 8 and identification c is 4; the object identification pair with serial number 6 includes the source object identification generated by identification 9 and the associated object identification generated by identification b, and the association weight between identification 9 and identification b is 3.

[0147] Obviously, Figure 3 The example source object identifiers (including identifier 1, identifier 5, identifier 7, identifier 8, and identifier 9) all belong to abnormal object identifiers included in three abnormal groups.

[0148] Step S103 , performing diffusion processing on multiple abnormal groups respectively through relational data to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group belong to associated object identifiers.

[0149] Specifically, the multiple abnormal groups include abnormal group E f , f is a positive integer, and f is less than or equal to the total number of multiple abnormal groups; abnormal group E f Including the exception object identifier G h , h is a positive integer, and h is less than or equal to the abnormal population E f The total number of abnormal object identifiers in the; the total number of source object identifiers and the total number of associated object identifiers are Z, Z is a positive integer; there is an association relationship between a source object identifier and an associated object identifier; the abnormal object identifier G h Match with Z source object identifiers. If there is an exception object identifier G in the Z source object identifiers, h If the source object ID is the same, then among the Z associated object IDs, obtain the one with the exception object ID G h The associated object identifier with an associated relationship; will be associated with the abnormal object identifier G h The associated object ID with an associated relationship is determined as the abnormal object ID Gh Diffusion object identification; abnormal object identification G h The diffusion object identifier is added to the abnormal group E f , will add the abnormal object identifier G h The abnormal group E identified by the diffusion object f , identified as abnormal group E f The corresponding abnormal diffusion group.

[0150] Please see again Figure 3 , Figure 3 In this example, the total number Z of source object identifiers is 6. The business server performs diffusion processing on abnormal group 31a through relationship data 30b. The specific diffusion process is as follows: the business server matches each abnormal object identifier in abnormal group 31a with each source object identifier in relationship data 30b. For example, the business server matches identifier 1 with each of the six source object identifiers in relationship data 30b. Through the matching, the business server can determine that the first and second source object identifiers in relationship data 30b match identifier 1, and therefore adds the first and second associated object identifiers to abnormal group 31a. Similarly, the business server matches identifier 2 with the six source object identifiers in the relationship data 30b respectively. Through matching, it can be determined that there is no source object identifier matching identifier 2 in the relationship data 30b; the business server matches identifier 3 with the six source object identifiers in the relationship data 30b respectively. Through matching, it can be determined that there is no source object identifier matching identifier 3 in the relationship data 30b; therefore, the business server adds the first associated object identifier and the second associated object identifier to the abnormal group 31a, and obtains the abnormal diffusion group 31c corresponding to the abnormal group 31a, as shown in FIG. Figure 3 shown.

[0151] According to the above processing, the service server can generate an abnormal diffusion group 32c corresponding to the abnormal group 32a, such as Figure 3 As shown, the diffusion object identifiers in the abnormal diffusion group 32c include identifier a that matches identifier 5, and identifier b that matches identifier 7. The business server can generate an abnormal diffusion group 33c corresponding to the abnormal group 33a, as shown in FIG. Figure 3 As shown, the diffusion object identifiers in the abnormal diffusion group 33c include identifier c that matches identifier 8 and identifier b that matches identifier 9.

[0152] Step S104 , performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups.

[0153] Specifically, the multiple abnormal diffusion groups include a first abnormal diffusion group and a second abnormal diffusion group; the diffusion object identifiers in the first abnormal diffusion group are compared with the diffusion object identifiers in the second abnormal diffusion group; if the first abnormal diffusion group and the second abnormal diffusion group contain the same diffusion object identifiers, the first abnormal diffusion group and the second abnormal diffusion group are determined as an abnormal diffusion group group; and the same diffusion object identifiers contained in the first abnormal diffusion group and the second abnormal diffusion group are determined as the same diffusion object identifiers of the abnormal diffusion group group.

[0154] It can be understood that the first abnormal diffusion group is any one of the multiple abnormal diffusion groups, and the second abnormal diffusion group is any one of the multiple abnormal diffusion groups and different from the first abnormal diffusion group.

[0155] Please refer to Figure 3 , it is assumed that the first abnormal diffusion group is Figure 3 the abnormal diffusion group 31c of the example, and the second abnormal diffusion group is Figure 4 the abnormal diffusion group 32c of the example. The diffusion object identifiers in the abnormal diffusion group 31c include the first associated object identifier (i.e., identifier a) and the second associated object identifier (i.e., identifier b) in the associated data 30b, and the diffusion object identifiers in the abnormal diffusion group 32c include the third associated object identifier (i.e., identifier a) and the fourth associated object identifier (i.e., identifier b) in the associated data 30b.

[0156] The business server compares the diffusion object identifiers in the abnormal diffusion group 31c and the diffusion object identifiers in the abnormal diffusion group 32c, and can determine that the abnormal diffusion group 31c and the abnormal diffusion group 32c both include identifier a and identifier b, so the abnormal diffusion group 31c and the abnormal diffusion group 32c are determined as an abnormal diffusion group group (1, 2), where 1 in the bracket can represent the abnormal diffusion group 31c, and 2 in the bracket can represent the abnormal diffusion group 32c; and the business server determines identifier a and identifier b as the same diffusion object identifiers of the abnormal diffusion group group (1, 2).

[0157] In the same processing, the business server compares the diffusion object identifiers in abnormal diffusion group 31c with the diffusion object identifiers in abnormal diffusion group 33c, and can determine that abnormal diffusion group 31c and abnormal diffusion group 33c both include identifier b. Therefore, abnormal diffusion group 31c and abnormal diffusion group 33c are determined to be abnormal diffusion group group (1, 3), where the 3 in the brackets can represent abnormal diffusion group 33c; the business server determines identifier b as the same diffusion object identifier of abnormal diffusion group group (1, 3). The business server compares the diffusion object identifier in abnormal diffusion group 33c with the diffusion object identifier in abnormal diffusion group 32c, and can determine that abnormal diffusion group 33c and abnormal diffusion group 32c both include identifier b. Therefore, abnormal diffusion group 33c and abnormal diffusion group 32c are determined to be abnormal diffusion group group (2, 3), and the business server determines identifier b as the same diffusion object identifier of abnormal diffusion group group (2, 3).

[0158] Step S105 : merging two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group.

[0159] Specifically, multiple abnormal diffusion groups include Y object identifiers; Y is a positive integer greater than 1; the identification scores corresponding to the Y object identifiers are determined; the same diffusion object identifiers contained in the abnormal diffusion group group are determined as the common object identifier of the abnormal diffusion group group; among the Y identification scores, the common identification score of the common object identifier is obtained; according to the common identification score, two abnormal diffusion groups in the abnormal diffusion group group are merged to obtain an abnormal merged group.

[0160] Among them, according to the public identification score, the two abnormal diffusion groups in the abnormal diffusion group group are merged to obtain the abnormal merged group. The specific process may include: averaging the public identification score to obtain the public score corresponding to the abnormal diffusion group group; comparing the public score with the merge score threshold, if the public score is equal to or greater than the merge score threshold, merging the two abnormal diffusion groups in the abnormal diffusion group group to obtain an abnormal merged diffusion group; deleting the diffusion object identification in the abnormal merged diffusion group, and determining the abnormal merged diffusion group after the diffusion object identification is deleted as the abnormal merged group.

[0161] Among them, the total number of abnormal diffusion group groups is at least two, and the at least two abnormal diffusion group groups include a first abnormal diffusion group group and a second abnormal diffusion group group different from the first abnormal diffusion group group; the common score includes a first common score corresponding to the first abnormal diffusion group group and a second common score corresponding to the second abnormal diffusion group group; if the common score is equal to or greater than the merge score threshold, the two abnormal diffusion groups in the abnormal diffusion group group are merged, and the specific process of obtaining the abnormal merged diffusion group may include: if the first common score is equal to or greater than the merge score threshold, and the second common score is equal to or greater than the merge score threshold, the two abnormal diffusion groups in the first abnormal diffusion group group and the two abnormal diffusion groups in the second abnormal diffusion group group are compared; if there is an identical abnormal diffusion group in the first abnormal diffusion group group and the second abnormal diffusion group group, the two abnormal diffusion groups in the first abnormal diffusion group group and the two abnormal diffusion groups in the second abnormal diffusion group group are merged to obtain the abnormal merged diffusion group.

[0162] Please also see Figure 4 as well as Figure 2 , Figure 3 This is a data processing scenario provided by the embodiment of the present application. Figure 3 . Figure 6 The example multiple abnormal diffusion groups include abnormal diffusion group 31c, abnormal diffusion group 32c and abnormal diffusion group 33c. The three abnormal diffusion groups include 9 abnormal object identifiers (identifier 1 to identifier 9) and 3 diffusion object identifiers (identifier a, identifier b and identifier c). Figure 4 In this example, Y is 12. The business server determines the identification scores corresponding to the 12 object identifiers. This embodiment of the application does not describe the process of determining the identification scores of the object identifiers. Please refer to the following for details. Figure 4 Description in the corresponding embodiment.

[0163] Figure 4 The identity fraction is abbreviated as fraction, such as Figure 4 As shown, the score of identifier 1 is 1, the scores of identifiers 2, 8 and b are all 3, the scores of identifiers 3, 5, 7 and a are all 2, the scores of identifiers 4, 9 and c are all 4, and the score of identifier 6 is 7. The business server can determine the same diffusion object identifiers contained in the abnormal diffusion group (1, 2) as the common object identifiers of the abnormal diffusion group (1, 2), such as Figure 4 The logo a and logo b are shown. Figure 4 Among the 12 identification scores in the example, the business server obtains the score of identification a, such as Figure 4 2 of the example, and the score of the identifier b, such asFigure 4 Example 4. In the embodiment of the present application, the identification score of a public object identification is referred to as a public identification score.

[0164] Furthermore, the business server performs mean processing on the public identification scores. For example, the public identification scores 2 and 4 of the abnormal diffusion group (1, 2) are averaged to obtain a public score 3 corresponding to the abnormal diffusion group (1, 2). The business server compares the public score 3 with the combined score threshold. Figure 4 The example merge score threshold is 2, so Figure 3 The common score 3 of the example abnormal diffusion group (1, 2) is greater than the merge score threshold. At this time, the business server merges the two abnormal diffusion groups in the abnormal diffusion group (1, 2) to obtain the abnormal merged diffusion group 30d, which includes all object identifiers of the abnormal diffusion group 31c and the abnormal diffusion group 32c. It can be understood that the business server merges the abnormal groups, so after obtaining the abnormal merged diffusion group 30d, it will delete the object identifiers in the abnormal merged diffusion group 30d that do not belong to the abnormal group 31a and the abnormal group 32a, that is, delete the diffusion object identifiers, such as Figure 4 As shown, the business server can obtain the abnormal merged group 30e, that is, all abnormal object identifiers including the abnormal group 31a and the abnormal group 32a.

[0165] Likewise, the business server Figure 4 The public identification score of the abnormal diffusion group (1, 3) of the example (such as Figure 3 The example 3) is processed by averaging to obtain the common score 3 of the abnormal diffusion group (1, 3). Since the common score 3 of the abnormal diffusion group (1, 3) is greater than Figure 4 The merge score threshold of the example is set, so the business server merges the two abnormal diffusion groups in the abnormal diffusion group (1, 3). As can be seen from the above, the abnormal diffusion group (1, 2) also meets the merging conditions, that is, the common score 3 of the abnormal diffusion group (1, 2) is greater than the merge score threshold. Therefore, at this time, the business server can add the abnormal diffusion group 33c to the abnormal merged diffusion group 30d, that is, the embodiment of the present application can not only merge an abnormal diffusion group that meets the merging conditions, but also merge multiple abnormal diffusion group groups that meet the merging conditions and have a common abnormal diffusion group. For example Figure 3 as well as Figure 5 The abnormal diffusion group (1, 3) and the abnormal diffusion group (1, 2) of the example both meet the merging conditions, and the abnormal diffusion group (1, 3) and the abnormal diffusion group (1, 2) have a common abnormal diffusion group, that is, Figure 5 Example of abnormal diffusion population 33c.

[0166] It is understandable that the processing process of the business server for the abnormal diffusion group (2, 3) is the same as the processing process described above, so it will not be repeated here.

[0167] The embodiments of the present application can be applied to suspicious transaction detection scenarios to identify suspicious transactions and merchants. Specifically, this method is used to screen all objects (such as merchants), identify suspicious objects, generate standard cases through the system, and push them to auditors for review or due diligence. Auditors will feedback the audit conclusions on the system. If it is determined that there is indeed suspicious transaction behavior, a suspicious transaction report needs to be generated and reported to a specific detection department. The embodiments of the present application merge abnormal groups in order to expand the size of the same group as much as possible and avoid pushing cases separately, so as to facilitate the audit work.

[0168] From the above, it can be seen that the embodiment of the present application can identify abnormal groups with related relationships through group diffusion, and can transform small groups into large groups by merging abnormal groups with related relationships, so the audit efficiency of abnormal object identification can be improved. By determining larger-scale abnormal groups, large-scale abnormal groups can be processed first, so audit resources can be allocated more reasonably, thereby improving resource utilization.

[0169] Further, see Figure 2 , Figure 1 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 1 This method can be performed by a business server (for example, Figure 5 The service server 100 shown in FIG. 1 may also be executed by a terminal device (for example, the above Figure 2 The terminal device 200a shown in the figure can also be executed by the business server and the terminal device interactively. For ease of understanding, the embodiment of the present application takes the method executed by the business server as an example for explanation, that is, the business server executes the data processing method as a computer device. Figure 3 As shown, the method may at least include the following steps S201 to S209.

[0170] Step S201 , obtaining multiple abnormal groups; the multiple abnormal groups all refer to groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other.

[0171] For the specific implementation process of step S201, please refer to the above Figure 4 Step S101 in the corresponding embodiment is not described in detail here.

[0172] Step S202 , acquiring relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier.

[0173] Specifically, the implementation strategy for merging anomaly groups is to diffuse each anomaly group based on relational data. The new object identifiers obtained through diffusion are referred to as associated object identifiers in the relational data. In this embodiment, the associated object identifiers added to an anomaly group are referred to as diffused object identifiers. Duplicate diffused object identifiers may exist in different anomaly diffusion groups, which are also referred to as identical diffused object identifiers and common object identifiers in this embodiment. Therefore, the common identifier scores of the common object identifiers can be calculated to merge different anomaly diffusion groups.

[0174] The embodiment of the present application uses relational data to perform a one-degree diffusion on each abnormal group.

[0175] Step S203 , performing diffusion processing on multiple abnormal groups respectively through relational data to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group belong to associated object identifiers.

[0176] Specifically, the business server can construct an original topology map based on multiple abnormal groups. This step does not describe the process of constructing the original topology map in detail; please refer to the description of step S205 below. The business server can use the original topology map and relationship data as input data and generate output data using the diffusion algorithm shown in Table 1. The output data includes multiple abnormal diffusion groups and a diffusion topology map.

[0177] Table 1

[0178]

[0179] Among them, G = (V, E) in Table 1 represents the original topological graph, V represents the node set, a node in the node set is used to represent an abnormal object identifier in multiple abnormal groups, E represents the edge set, an edge in the edge set is used to represent the association relationship between two abnormal object identifiers; R represents the relationship data, src i Indicates the i-th source object identifier, dst i Indicates the i-th associated object identifier, weight i Indicates src i and dst iG′=(V′,E′) represents the diffusion topology graph, V′ represents the node set after diffusion, which includes not only the above V but also the nodes used to represent the diffusion object identifier (referred to as neighbor nodes in Table 1), and E′ represents the edge set after diffusion, which includes not only the above E but also the association weights in the relational data (representing the degree of association between the diffusion object identifier and the abnormal object identifier associated with the diffusion object identifier); T i represents the i-th abnormal group, T represents all abnormal groups, T i ′ represents the abnormal diffusion group corresponding to the i-th abnormal group, and T′ represents all abnormal diffusion groups.

[0180] The plurality of abnormal diffusion groups includes Y object identifiers; Y is a positive integer greater than 1.

[0181] Step S204 , performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups.

[0182] Specifically, in conjunction with the description of step S203 , after generating the diffusion topology graph, the service server may determine the common object identifiers in different abnormal diffusion groups by using the set intersection algorithm shown in Table 2.

[0183] Table 2

[0184]

[0185] Among them, C(T i ′,T j ′) represents the common object identifier between the i-th abnormal diffusion group and the j-th abnormal diffusion group, that is, the same diffusion object identifier.

[0186] Step S205: constructing an original topology map based on the multiple abnormal groups.

[0187] Specifically, the multiple abnormal groups include Z abnormal object identifiers; the Z abnormal object identifiers belong to Y object identifiers; Z is a positive integer greater than 1; Z nodes are generated for representing the Z abnormal object identifiers; wherein one node is used to represent one abnormal object identifier; based on the association relationship between every two abnormal object identifiers in the Z abnormal object identifiers, an edge score between every two nodes in the Z nodes is determined; and based on the Z nodes and the edge score between every two nodes, an original topological graph is constructed.

[0188] The Z abnormal object identifiers include a first abnormal object identifier and a second abnormal object identifier different from the first abnormal object identifier; the specific process of determining the edge score between each two nodes in the Z nodes based on the association relationship between each two abnormal object identifiers in the Z abnormal object identifiers may include: obtaining first registration information corresponding to the first abnormal object identifier and second registration information corresponding to the second abnormal object identifier; performing similarity processing on the first registration information and the second registration information to obtain a similarity value between the first registration information and the second registration information; determining the similarity value between the first registration information and the second registration information as the edge score between the first node and the second node; the first node is used to represent the first abnormal object identifier; the second node is used to represent the second abnormal object identifier; and the first node and the second node both belong to the Z nodes.

[0189] The business server may execute steps S205 and S206 after determining the abnormal diffusion group, or may integrate steps S205 and S206 into step S203 for execution, which is not limited in this embodiment of the present application.

[0190] This step mainly describes the construction process of the original topology map. Please also refer to Figure 6 、 Figure 6 as well as Figure 3 , Figure 3 This is a data processing scenario provided by the embodiment of the present application. Figure 3 .like Figure 6 For example, multiple abnormal groups include 9 abnormal object identifiers, namely Figure 2 The example ID 1 to ID 9, that is Figure 6 In the example, Z is 9. The business server builds a node for each abnormal object ID, such as Figure 6 As shown, node 1 is constructed for identifier 1, node 2 is constructed for identifier 2, node 3 is constructed for identifier 3, node 4 is constructed for identifier 4, node 5 is constructed for identifier 5, node 6 is constructed for identifier 6, node 7 is constructed for identifier 7, node 8 is constructed for identifier 8, and node 9 is constructed for identifier 9.

[0191] It is understandable that the way in which the business server determines the association relationship between two abnormal object identifiers is the same as the way in which the business server determines the degree of association between the source object identifier and the associated object identifier, so it is not described here in detail. Please refer to the above. Figure 6 Similarly, the embodiment of the present application does not limit the method for determining the association relationship between two abnormal object identifiers, and can be set according to the actual application scenario, and the method for determining the degree of association between the source object identifier and the associated object identifier can be the same.

[0192] Please see again Figure 6, for the convenience of description and understanding, Figure 6 In the example, the association relationship between identifier 1 and identifier 2 is a value of 3, so the business server can construct an edge with a score of 3 between node 1 and node 2. In this embodiment of the application, the edge score is referred to as the edge score. Figure 6 For example, the association between identifier 1 and identifier 3 is a value of 2, so the business server can build an edge with a score of 2 between node 1 and node 3. For example, the association between identifier 1 and identifier 4 is a value of 2, so the business server can build an edge with a score of 2 between node 1 and node 4. Figure 3 For example, the association between identifier 2 and identifier 3 is a value of 3, so the business server can build an edge with a score of 3 between nodes 2 and 3. For example, the association between identifier 3 and identifier 4 is a value of 4, so the business server can build an edge with a score of 4 between nodes 3 and 4. Figure 4 The association between example identifiers 4 and 5 is value 1, so the business server can build an edge with a score of 1 between nodes 4 and 5. The association between example identifiers 5 and 6 is value 1, so the business server can build an edge with a score of 1 between nodes 5 and 6. Figure 6 For example, the association between IDs 5 and 7 is a value of 3, so the service server can construct an edge with a score of 3 between nodes 5 and 7. For example, the association between IDs 8 and 9 is a value of 5, so the service server can construct an edge with a score of 5 between nodes 8 and 9. It can be understood that the absence of an edge between two nodes is because there is no association between the corresponding two abnormal object IDs.

[0193] It needs to be emphasized that Figure 6 、 Figure 6 as well as Figure 3 The specific values ​​are for ease of description and understanding of the examples and do not represent the values ​​in actual application scenarios.

[0194] Please see again Figure 6 , the business server builds a Figure 3 The original topology diagram 60a in .

[0195] Step S206 , performing diffusion processing on the original topology map according to the multiple abnormal diffusion groups to obtain a diffusion topology map; the diffusion topology map includes Y nodes; one node is used to represent one object identifier among the Y object identifiers.

[0196] Specifically, the relationship data also includes an association weight for characterizing the degree of association between the source object identifier and the associated object identifier; the plurality of abnormal diffusion groups include the abnormal object identifier I j , and the exception object identifier I j Diffusion object identifier K j ; j is a positive integer, and j is less than or equal to the total number of diffusion object identifiers in multiple abnormal diffusion groups; diffusion object identifier K j Belongs to the associated object identifier; abnormal object identifier I j Belongs to the source object identifier; generates the identifier K used to characterize the diffusion object j The third node; in the relational data, obtain the identifier I used to characterize the abnormal object j And the diffusion object identifier K j The association weight of the degree of association between j And the diffusion object identifier K j The association weight between them is determined as the edge score between the third node and the fourth node; the fourth node refers to the identifier I used to characterize the abnormal object in the original topology graph. j ; adding the third node and the edge score to the original topology graph, and determining the original topology graph with the third node and the edge score added as the diffusion topology graph.

[0197] The business server generates nodes for representing the diffusion object identifiers in multiple abnormal diffusion groups, including the third node described above. Please refer to Figure 6 as well as Figure 6 , the three abnormal diffusion groups include three diffusion object identifiers, namely identifier a, identifier b and identifier c. Figure 6 The example relationship data 30b shows that the identifier a is associated with the identifier 1 in the abnormal diffusion group 31c, and the association weight is 8, so the business server can Figure 6 In the original topology graph 60a, node a is added and an edge with a score of 8 is constructed between node 1 and node a. Identifier a is associated with identifier 5 in the abnormal diffusion group 32c, and the association weight is 4. Therefore, the business server constructs an edge with a score of 4 between node 5 and node a. Identifier b is associated with identifier 1 in the abnormal diffusion group 31c, and the association weight is 7. Therefore, the business server can Figure 6Node b is added to the original topology graph 60a in the example, and an edge with a score of 7 is constructed between node 1 and node b. Identifier b is associated with identifier 7 in the abnormal diffusion group 32c, and the association weight is 3, so the business server constructs an edge with a score of 3 between node 7 and node b. Identifier b is associated with identifier 9 in the abnormal diffusion group 33c, and the association weight is 3, so the business server constructs an edge with a score of 3 between node 9 and node b. Identifier c is associated with identifier 8 in the abnormal diffusion group 33c, and the association weight is 4, so the business server can Figure 6 Add node c to the original topology graph 60a in the figure, and build an edge with a score of 4 between node 8 and node c; after the above processing, the service server can obtain Figure 3 Diffusion topology in Figure 60b.

[0198] Step S207 : performing node score processing on the diffusion topology graph to obtain node scores corresponding to the Y nodes.

[0199] Specifically, Figure 6 Each edge in the diffusion topology graph 60b has a score, so the service server can Figure 4 The node score of each node is determined by all the edge scores in the diffusion topology graph 60b. One feasible way is to use the PageRank algorithm to calculate the node score of each node, where PageRank is a web page ranking algorithm that mainly determines the importance of a page by analyzing the link structure between web pages. The ranking of each page depends on the number and quality of other web pages linked to it. PageRank was originally used to improve search engine results, and is now also used in graph network analysis and other fields. Therefore, the business server determines the node score corresponding to each node through the PageRank algorithm shown in Table 3.

[0200] Table 3

[0201]

[0202] Among them, the meanings of some identifiers in Table 3 are the same as those in Table 1 and Table 2 above, so they are not explained here. Among them, PR(v′) represents the PageRank value of the node in the diffusion topology graph, which is called the node score in this application.

[0203] Step S208: Determine the Y node scores as identification scores corresponding to the Y object identifications.

[0204] Specifically, it can be understood that the node is used to represent the object identity, so the node score represents the identity score of the object identity, so Figure 7 as well as Figure 7 The combination ofFigure 3 12 identity scores for the sample.

[0205] Step S209 : merging two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group.

[0206] Specifically, for each abnormal diffusion group, the business server calculates the average score of its public object identifier. If the average is greater than the merge score threshold, the two abnormal diffusion groups in the abnormal diffusion group are merged. This process can be implemented by the algorithm shown in Table 4 below.

[0207] Table 4

[0208]

[0209] C(T i′ ) represents the abnormal diffusion group T i ′’s public object identifier, T combine represents an abnormal merged diffusion group. It can be understood that for multiple abnormal diffusion groups, there may be one or more abnormal merged diffusion groups; S(T i ′,T j ′) represents the common fraction.

[0210] After processing all abnormal diffusion groups, the service server traverses the set, each element represents the group number that needs to be merged, and merges the abnormal diffusion groups. The merging score threshold in the embodiment of the present application is an adjustable parameter.

[0211] After the merging step is completed, the business server removes all diffused object identifiers and retains the abnormal object identifiers from the abnormal group. The reason for implementing this step is that the goal of this embodiment of the application is to identify newly registered object identifiers, and the diffused object identifiers may not be newly registered. These object identifiers are only used for diffusion and do not need to be retained in the final result. This process can be implemented as shown in Table 5 below.

[0212] Table 5

[0213]

[0214] Among them, T in Table 5 final Indicates an abnormally merged population.

[0215] As can be seen from the above, this application can not only effectively identify batch registration object identifiers, but also merge potentially related abnormal groups. Merging small groups into large groups facilitates batch review, improves review efficiency, and helps identify and reduce major illegal risks.

[0216] The method for identifying batch registration object identifiers based on registration information and transaction characteristics proposed in the embodiments of the present application can produce the following beneficial effects:

[0217] 1. Comprehensive risk analysis: By combining multiple information dimensions (such as object name, business scope, opening bank, region, etc.), it provides a more comprehensive risk assessment and can more accurately identify and distinguish truly abnormal behavior from normal behavior.

[0218] 2. Improved accuracy: Improved the accuracy of identifying batch registration resource management identifiers, reducing false positives and missed negatives.

[0219] 3. Reveal hidden relationships: Through the methods of node diffusion (i.e., object identifier diffusion) and abnormal group merging, it is possible to identify and reveal the potential connections between seemingly independent batch registration object identifiers, helping to discover larger-scale illegal networks.

[0220] 4. Improve efficiency: The method of merging small groups into large groups is conducive to batch review, reducing analysis time and resource consumption, and improving review efficiency.

[0221] 5. Strong adaptability: The technologies and methods adopted can flexibly adapt to changes in illegal behaviors and maintain effective detection of illegal activities.

[0222] 6. Better resource allocation: By effectively identifying and prioritizing large-scale illegal risks, detection resources can be more rationally allocated to ensure that they focus on the most abnormal behaviors.

[0223] 7. Improve strategy effectiveness: Identifying and analyzing large groups can help develop more effective intervention measures, thereby more effectively blocking illegal activities.

[0224] Further, see Figure 1 , Figure 1 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 7 The data processing method can be performed by a business server (for example, Figure 8 The service server 100 shown in FIG. 1 may also be executed by a terminal device (for example, the above Figure 8 The terminal device 200a shown in the figure can also be executed by the business server and the terminal device interactively. For ease of understanding, the embodiment of the present application takes the method executed by the business server as an example for explanation, that is, the business server executes the data processing method as a computer device. Figure 4 As shown, the data processing method may at least include the following steps S301 to S309.

[0225] Step S301: obtain first detailed data corresponding to multiple resource management identifiers, merge the first detailed data with the same object identifier among the multiple first detailed data to obtain multiple second detailed data; wherein the object identifiers corresponding to the multiple second detailed data are different from each other.

[0226] Specifically, identifying batch registration objects (such as merchants) plays a vital role in suspicious transaction detection. Criminals often use automated technology to batch register a large number of payment resource management identifiers (such as business accounts) based on false or fraudulent information, and activate resource management identifiers in batches to carry out illegal activities. After completing the illegal activities, they will quickly withdraw cash and transfer resources. Effective and timely identification of these batch entities can significantly improve the accuracy and efficiency of anti-anomaly measures, thereby protecting the integrity and security of the financial system, and is an important part of improving the effectiveness of suspicious transaction detection.

[0227] Based on the above reasons, the present application embodiment proposes a method for identifying batch registration object identifiers based on registration information and transaction characteristics. The method is mainly implemented by the following steps:

[0228] 1. Select recently registered resource management identifiers as analysis objects.

[0229] Among them, recent is an adjustable time period, such as 7 days, one month, etc., that is, the resource management identifiers registered within the last 7 days are extracted, or the resource management identifiers registered within the last 30 days are extracted.

[0230] Get the first detailed data corresponding to multiple resource management identifiers. The embodiment of this application does not limit the content of the first detailed data, which can be set according to the actual application scenario, including but not limited to the business scope of the object (such as business scope), opening bank (that is, the bank bound to the resource management identifier), region, legal person, ultimate beneficiary, object identifier, object name, etc. Please refer to Figure 8 , Figure 8 This is a data processing scenario provided by the embodiment of the present application. Figure 8 . Figure 8 The example plurality of resource management identifiers includes resource management identifier 1, resource management identifier 2, resource management identifier 3, resource management identifier 4, resource management identifier 5, and resource management identifier 6. Figure 8 The example first detailed data includes object name, object ID, industry, region, and business scope. It can be understood that the first detailed data is divided based on the resource management ID.

[0231] like Figure 2As shown, the first detailed data of resource identification 1 is as follows: the object name is Name 1, the object identification is Identification 1, the industry is retail, the region is Region 1, and the business scope is food and daily necessities; the first detailed data of resource identification 2 is as follows: the object name is Name 1, the object identification is Identification 1, the industry is retail, the region is Region 2, and the business scope is food and daily necessities; the first detailed data of resource identification 2 is as follows: the object name is Name 1, the object identification is Identification 1, the industry is technology, the region is Region 1, and the business scope is software development; the first detailed data of resource identification 4 is as follows: the object name is Name 2, the object identification is Identification 2, the industry is e-commerce, the region is Region 3, and the business scope is electronic products and software services; the first detailed data of resource identification 5 is as follows: the object name is Name 2, the object identification is Identification 2, the industry is e-commerce, the region is Region 4, and the business scope is electronic products and software services; the first detailed data of resource identification 6 is as follows: the object name is Name 3, the object identification is Identification 3, the industry is technology, the region is Region 1, and the business scope is software development.

[0232] The embodiment of the present application does not limit the number of first detailed data (i.e., resource management identifiers) and can be set according to the actual application scenario. If there is only one resource management identifier (i.e., only one first detailed data), the following multi-dimensional identification processing is performed.

[0233] The business server merges the first detailed data with the same object identifier among the multiple first detailed data to obtain multiple second detailed data divided by the object identifier. Figure 9 , the number of object identifiers is 3, namely identifier 1, identifier 2 and identifier 3, so 3 second detailed data can be obtained; the second detailed data of identifier 1 is as follows: the resource management identifier includes resource identifier 1, resource identifier 2 and resource identifier 3, the object name is name 1, the industry includes retail and technology, the region includes region 1 and region 2, and the business scope includes food, daily necessities and software development; the second detailed data of identifier 2 is as follows: the resource management identifier includes resource identifier 4 and resource identifier 5, the object name is name 2, the industry includes e-commerce, the region includes region 3 and region 4, and the business scope includes electronic products and software services; the second detailed data of identifier 3 is as follows: the resource management identifier includes resource identifier 6, the object name is name 3, the industry includes technology, the region includes region 1, and the business scope includes software development.

[0234] The embodiment of the present application does not limit the number of second detailed data (ie, object identifiers) and can be set according to the actual application scenario. If there is only one object identifier (ie, only one second detailed data), the business server performs multi-dimensional identification processing.

[0235] It is understandable that the suspicious transaction detection in the embodiment of the present application performs transaction detection based on the object dimension, so detailed data is integrated according to the object identifier, thereby avoiding the detection of a large number of resource management identifiers.

[0236] Step S302 : dividing the plurality of object identifiers according to the object names respectively included in the plurality of second detailed data to obtain A groups; A is a positive integer; the object identifiers respectively included in the A groups are different from each other.

[0237] Specifically, a text recognition model is obtained, and the object names respectively included in the multiple second detailed data are input into the text recognition model; through the text recognition model, feature extraction is performed on the multiple object names respectively to obtain multiple text vectors; wherein, a text vector is obtained by extracting features of an object name through the text recognition model; a vector clustering model is obtained, and the multiple text vectors are input into the vector clustering model respectively, and the multiple text vectors are clustered by the vector clustering model to obtain A text vector cluster clusters; according to the A text vector cluster clusters, the multiple object identifiers are divided and processed to obtain A groups; wherein, the text vector corresponding to the object identifier in a group belongs to a text vector cluster cluster.

[0238] The embodiments of the present application do not limit the model type of the text recognition model and can be set according to the actual application scenario. For example, ALBERT can be used as the text recognition model. ALBERT (A Lite BERT) is an optimized natural language processing model based on Bidirectional Encoder Representations from Transformers (BERT). It mainly improves training speed and efficiency by sharing parameters and reducing model complexity, while reducing memory consumption, making it more efficient when processing large-scale data sets, especially in text embedding and semantic understanding.

[0239] Text vectors can be understood as embeddings. In machine learning and natural language processing, embedding refers to the process of converting words, phrases, or other data items into fixed-length real-number vectors. This conversion helps machines understand and process language data because it captures semantic relationships and contextual meaning between words.

[0240] The embodiments of this application do not limit the model type of the vector clustering model; it can be set according to the actual application scenario. For example, DBSCAN can be used as a text clustering model. DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a popular spatial clustering algorithm that is particularly suitable for noisy datasets. It forms clusters based on the density of data points, can identify clusters of various shapes, and does not require a pre-specified number of clusters. This makes DBSCAN particularly useful when processing complex or non-standard datasets.

[0241] In addition, the embodiment of the present application can also use a graph neural network (GNN) to divide and process multiple object identifiers. The specific process is as follows:

[0242] 1. Data representation: Graph neural networks represent data through graphs (nodes / vertices represent entities, such as merchants; edges represent the relationships between entities), which is suitable for capturing complex relationships and interactions.

[0243] 2. Information aggregation: GNN updates the representation of each node by aggregating information about the node and its neighbors, which allows the algorithm to capture the relationships and structural features between nodes.

[0244] 3. Learning tasks: These can be node classification (determining whether each object identifier has unusual properties), graph classification (determining the level of unusualness of the entire object identifier network), or link prediction (predicting which resource management identifiers may be related).

[0245] Graph neural networks have the following advantages:

[0246] 1. Richer data representation: GNN can directly process relational data between object identifiers and can better represent complex network structures than vector embedding methods.

[0247] 2. Dynamic learning capability: GNN can continuously optimize node representation through the learning process and adapt to new data patterns, rather than just static clustering.

[0248] 3. Deep relationship mining: GNN can capture deeper network patterns, such as the internal structure of a group.

[0249] 4. Generalization ability: After learning, GNN also has good generalization ability for new nodes or structures and can process unseen data.

[0250] Step S303 : performing identification processing on the A groups in multiple information dimensions to obtain group attributes corresponding to the A groups.

[0251] Specifically, the A groups include a group B c , c is a positive integer, and c is less than or equal to A; group B c is identified in multiple information dimensions to obtain identification results corresponding to the multiple information dimensions respectively; if there is an abnormal identification result in the multiple identification results, an abnormal group attribute is determined as the group attribute corresponding to group B c ; if the multiple identification results are all normal identification results, a normal group attribute is determined as the group attribute corresponding to group B c .

[0252] The multiple information dimensions include a name information dimension, and the multiple identification results include a name identification result; the specific process of identifying group B c in the multiple information dimensions to obtain identification results corresponding to the multiple information dimensions respectively can include: in the multiple second detailed data, an object name corresponding to an object identifier in group B c is obtained, and the obtained object name is determined as an object name set; a reference name with an abnormal attribute is obtained, and the reference name is matched with the object names in the object name set; if there is an object name in the object name set that matches the reference name successfully, an abnormal identification result is determined as the name identification result; if there is no object name in the object name set that matches the reference name successfully, a normal identification result is determined as the name identification result.

[0253] The multiple information dimensions include a registration information dimension, and the multiple identification results include a registration identification result; the specific process of identifying group B c in the multiple information dimensions to obtain identification results corresponding to the multiple information dimensions respectively can include: in the multiple second detailed data, registration information corresponding to an object identifier in group B c is obtained, and the obtained registration information is determined as a registration information set; in the registration information set, a maximum number of the same registration information is obtained, and a ratio between the maximum number and a total number of the registration information in the registration information set is determined as an abnormal information ratio; the abnormal information ratio is compared with an abnormal information ratio threshold value; if the abnormal information ratio is greater than the abnormal information ratio threshold value, an abnormal identification result is determined as the registration identification result; if the abnormal information ratio is equal to or less than the abnormal information ratio threshold value, a normal identification result is determined as the registration identification result.

[0254] The multiple information dimensions include a business information dimension, and the multiple identification results include a business identification result; the specific process of identifying group B c in the multiple information dimensions to obtain identification results corresponding to the multiple information dimensions respectively can include: in the multiple second detailed data, business information corresponding to an object identifier in group Bc The business scope corresponding to the object identifier in the business scope is determined as a business scope set; similarity processing is performed on every two business scopes in the business scope set to obtain D similarities corresponding to the business scope set; D is a positive integer; the similarity mean corresponding to the D similarities is determined, and the similarity mean is compared with a similarity threshold; if the similarity mean is equal to or greater than the similarity threshold, the abnormal recognition result is determined as the business recognition result; if the similarity mean is less than the similarity threshold, the normal recognition result is determined as the business recognition result.

[0255] After obtaining the group through the above cluster analysis, the business server needs to determine the group attributes of the group through multi-dimensional analysis, that is, whether it is an abnormal group.

[0256] The embodiments of this application do not limit the number of multiple information dimensions, and can be set according to actual application scenarios, including at least two. The embodiments of this application also do not limit the information dimensions, and can be set according to actual application scenarios. For ease of understanding and description, this application uses the name information dimension, registration information dimension, and business scope dimension as examples.

[0257] In terms of name information, the business server can identify object identifiers that match suspicious naming patterns using regular expressions or pattern matching. The object names of abnormal object identifiers typically follow specific naming patterns, which the business server can collect based on prior knowledge. If at least one object name corresponding to a group meets any abnormal pattern, the business server can determine that the group is abnormal. For example, if "Name 1" is a suspicious naming pattern, the group corresponding to Name 1 is considered abnormal.

[0258] In the registration information dimension, there can be multiple types of registration information, such as Figure 9 For example, if there are multiple types of registration information in an industry or region, the abnormal identification result is determined as the identification result when the abnormal information ratio corresponding to each type of registration information is greater than the abnormal information ratio threshold. The abnormal information ratio threshold can be an adjustable parameter.

[0259] In terms of business scope, business servers can use Jaccard similarity. Jaccard similarity measures the similarity between two sets by calculating the ratio of common elements to the total elements in the two sets. This method is commonly used in text analysis, data mining, and machine learning to assess the similarity of different datasets.

[0260] Taking the above three information dimensions into consideration, if a group meets at least one information dimension, the business server can determine that the group is an abnormal group, that is, the group attribute of the group is an abnormal group attribute.

[0261] Step S304: among the A groups, the group with the abnormal group attribute is determined as an abnormal group.

[0262] Specifically, the method proposed in the embodiment of the present application comprehensively considers multiple information dimensions and introduces the concepts of group merging and relationship network analysis to improve the accuracy and effectiveness of identifying illegal activities.

[0263] Step S305 , obtaining multiple abnormal groups; the multiple abnormal groups all refer to groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other.

[0264] Step S306 , acquiring relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier.

[0265] Step S307 , performing diffusion processing on multiple abnormal groups respectively through relational data to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group belong to associated object identifiers.

[0266] Step S308 , performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups.

[0267] Step S309 : merging two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group.

[0268] For the specific implementation process of steps S305 to S309, please refer to the above Figure 9 Steps S101 to S105 in the corresponding embodiment are not described in detail here.

[0269] From the above, it can be seen that the embodiment of the present application can identify abnormal groups with related relationships through group diffusion, and can transform small groups into large groups by merging abnormal groups with related relationships, so the audit efficiency of abnormal object identification can be improved. By determining larger-scale abnormal groups, large-scale abnormal groups can be processed first, so audit resources can be allocated more reasonably, thereby improving resource utilization.

[0270] Further, see Figure 10 , Figure 10 : is a structural diagram of a data processing device provided in an embodiment of the present application. The above-mentioned data processing device 1 can be a computer program (including program code) running on a computer device, for example, the data processing device 1 is an application software; the data processing device 1 can be used to execute the corresponding steps of the method provided in the embodiment of the present application. Figure 10 As shown, the data processing device 1 may include:

[0271] The acquisition module 11 is used to acquire multiple abnormal groups; the multiple abnormal groups are groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other;

[0272] The acquisition module 11 is further configured to acquire relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; and the associated object identifier is an object identifier that has an associated relationship with the source object identifier.

[0273] The processing module 12 is configured to perform diffusion processing on multiple abnormal groups respectively through the relational data to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through the relational data, and the diffusion object identifiers included in an abnormal diffusion group are associated object identifiers;

[0274] The processing module 12 is further configured to perform object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups;

[0275] The processing module 12 is further configured to merge two abnormal diffusion groups in the abnormal diffusion group group according to the same diffusion object identifier contained in the abnormal diffusion group group to obtain an abnormal merged group.

[0276] In a possible implementation, the plurality of abnormal groups include abnormal group E f , f is a positive integer, and f is less than or equal to the total number of multiple abnormal groups; abnormal group E f Including the exception object identifier G h , h is a positive integer, and h is less than or equal to the abnormal population E f The total number of abnormal object identifiers in the ; the total number of source object identifiers and the total number of associated object identifiers are both Z, where Z is a positive integer; there is an association relationship between a source object identifier and an associated object identifier;

[0277] The processing module 12 performs diffusion processing on the multiple abnormal groups respectively through the relational data to obtain multiple abnormal diffusion groups, which are used to perform the following operations:

[0278] Identify the abnormal object G h Match with Z source object identifiers. If there is an exception object identifier G in the Z source object identifiers, h If the source object ID is the same, then among the Z associated object IDs, obtain the one with the exception object ID G h The identifier of the associated object with which the association relationship exists;

[0279] will be identified with the exception object G h The associated object ID with an associated relationship is determined as the abnormal object ID G h The diffusion object identifier;

[0280] Identify the abnormal object G h The diffusion object identifier is added to the abnormal group E f , will add the abnormal object identifier G h The abnormal group E identified by the diffusion object f , identified as abnormal group E f The corresponding abnormal diffusion group.

[0281] In a possible implementation, the multiple abnormal diffusion groups include a first abnormal diffusion group and a second abnormal diffusion group;

[0282] The processing module 12 performs object identification processing on every two abnormal diffusion groups in the multiple abnormal diffusion groups to obtain an abnormal diffusion group group with the same diffusion object identification, which is used to perform the following operations:

[0283] comparing the diffusion object identifiers in the first abnormal diffusion group with the diffusion object identifiers in the second abnormal diffusion group;

[0284] If the first abnormal diffusion group and the second abnormal diffusion group include the same diffusion object identifier, the first abnormal diffusion group and the second abnormal diffusion group are determined as an abnormal diffusion group group;

[0285] The same diffusion object identifiers included in the first abnormal diffusion group and the second abnormal diffusion group are determined as the same diffusion object identifiers of the abnormal diffusion group group.

[0286] In a possible implementation, the plurality of abnormal diffusion groups include Y object identifiers; Y is a positive integer greater than 1;

[0287] The data processing device 1 further includes:

[0288] A determination module 13 is used to determine identification scores corresponding to the Y object identifications;

[0289] In a possible implementation, the processing module performs merging processing on two abnormal diffusion groups in the abnormal diffusion group set according to the same diffusion object identifier included in the abnormal diffusion group set, to obtain an abnormal merged group, for performing the following operations:

[0290] The same diffusion object identifier included in the abnormal diffusion group set is determined as a common object identifier of the abnormal diffusion group set.

[0291] Among the Y identifier scores, a common identifier score of the common object identifier is obtained.

[0292] The merging processing is performed on the two abnormal diffusion groups in the abnormal diffusion group set according to the common identifier score, to obtain the abnormal merged group.

[0293] In a possible implementation, the processing module 12 performs merging processing on two abnormal diffusion groups in the abnormal diffusion group set according to the common identifier score, to obtain an abnormal merged group, for performing the following operations:

[0294] The common identifier score is subjected to mean value processing, to obtain a common score corresponding to the abnormal diffusion group set.

[0295] The common score is compared with a merging score threshold value, and if the common score is equal to or greater than the merging score threshold value, the merging processing is performed on the two abnormal diffusion groups in the abnormal diffusion group set, to obtain an abnormal merged diffusion group.

[0296] The diffusion object identifier in the abnormal merged diffusion group is deleted, and the abnormal merged diffusion group after the deletion of the diffusion object identifier is determined as the abnormal merged group.

[0297] In a possible implementation, the total number of the abnormal diffusion group set is at least two, and the at least two abnormal diffusion group sets include a first abnormal diffusion group set and a second abnormal diffusion group set different from the first abnormal diffusion group set; the common score includes a first common score corresponding to the first abnormal diffusion group set and a second common score corresponding to the second abnormal diffusion group set.

[0298] If the common score is equal to or greater than the merging score threshold value, the processing module 12 performs the merging processing on the two abnormal diffusion groups in the abnormal diffusion group set, to obtain an abnormal merged diffusion group, for performing the following operations:

[0299] If the first common score is equal to or greater than the merging score threshold value and the second common score is equal to or greater than the merging score threshold value, the two abnormal diffusion groups in the first abnormal diffusion group set and the two abnormal diffusion groups in the second abnormal diffusion group set are compared.

[0300] If there is a same abnormal diffusion group in the first abnormal diffusion group and the second abnormal diffusion group, the two abnormal diffusion groups in the first abnormal diffusion group and the two abnormal diffusion groups in the second abnormal diffusion group are merged to obtain an abnormal merged diffusion group.

[0301] In a possible implementation, the determination module 13 determines the identification scores corresponding to the Y object identifications, and performs the following operations:

[0302] Construct an original topological map based on multiple abnormal groups;

[0303] According to the multiple abnormal diffusion groups, the original topology map is diffused to obtain a diffusion topology map; the diffusion topology map includes Y nodes; one node is used to represent one object identifier among the Y object identifiers;

[0304] Perform node score processing on the diffusion topology graph to obtain the node scores corresponding to Y nodes;

[0305] The Y node scores are determined as the identification scores corresponding to the Y object identifications respectively.

[0306] In a possible implementation, the plurality of abnormal groups include Z abnormal object identifiers; the Z abnormal object identifiers belong to Y object identifiers; Z is a positive integer greater than 1;

[0307] The determination module 13 constructs an original topology map based on the multiple abnormal groups, and is used to perform the following operations:

[0308] Generate Z nodes for representing Z abnormal object identifiers; wherein one node is used to represent one abnormal object identifier;

[0309] Determine an edge score between every two nodes in the Z nodes according to an association relationship between every two abnormal object identifiers in the Z abnormal object identifiers;

[0310] Construct the original topology graph based on Z nodes and the edge scores between every two nodes.

[0311] In a possible implementation, the Z abnormal object identifiers include a first abnormal object identifier and a second abnormal object identifier different from the first abnormal object identifier;

[0312] The determination module 13 determines the edge score between every two nodes in the Z nodes according to the association relationship between every two abnormal object identifiers in the Z abnormal object identifiers, and is used to perform the following operations:

[0313] Obtaining first registration information corresponding to the first abnormal object identifier and second registration information corresponding to the second abnormal object identifier;

[0314] Performing similarity processing on the first registration information and the second registration information to obtain a similarity value between the first registration information and the second registration information;

[0315] The similarity value between the first registration information and the second registration information is determined as the edge score between the first node and the second node; the first node is used to represent the first abnormal object identifier; the second node is used to represent the second abnormal object identifier; the first node and the second node both belong to Z nodes.

[0316] In a possible implementation, the relationship data further includes an association weight for characterizing the degree of association between the source object identifier and the associated object identifier; the plurality of abnormal diffusion groups include the abnormal object identifier I j , and the exception object identifier I j Diffusion object identifier K j ; j is a positive integer, and j is less than or equal to the total number of diffusion object identifiers in multiple abnormal diffusion groups; diffusion object identifier K j Belongs to the associated object identifier; abnormal object identifier I j Belongs to the source object identifier;

[0317] The determination module 13 performs diffusion processing on the original topology map according to the multiple abnormal diffusion groups to obtain a diffusion topology map for performing the following operations:

[0318] Generate the identifier K for characterizing the diffusion object j The third node of

[0319] In relational data, obtain the identifier I used to characterize the abnormal object j And the diffusion object identifier K j The association weight of the degree of association between them;

[0320] Identify the exception object as I j And the diffusion object identifier K j The association weight between them is determined as the edge score between the third node and the fourth node; the fourth node refers to the identifier I used to characterize the abnormal object in the original topology graph. j Node;

[0321] The third node and the edge score are added to the original topology graph, and the original topology graph with the third node and the edge score added is determined as a diffusion topology graph.

[0322] In a possible implementation, the data processing device 1 further includes:

[0323] The acquisition module 11 is further configured to acquire first detailed data corresponding to a plurality of resource management identifiers, and merge first detailed data with the same object identifier among the plurality of first detailed data to obtain a plurality of second detailed data; wherein the object identifiers corresponding to the plurality of second detailed data are different from each other;

[0324] The processing module 12 is further configured to divide the plurality of object identifiers into A groups according to the object names respectively included in the plurality of second detailed data, where A is a positive integer and the object identifiers respectively included in the A groups are different from each other;

[0325] The processing module 12 is further configured to perform identification processing on the A groups in multiple information dimensions to obtain group attributes corresponding to the A groups;

[0326] The determination module 13 is configured to determine, among the A groups, a group with an abnormal group attribute as an abnormal group.

[0327] In a possible implementation, the processing module 12 divides the plurality of object identifiers according to the object names respectively included in the plurality of second detailed data to obtain A groups for performing the following operations:

[0328] obtaining a text recognition model, and inputting the object names respectively included in the plurality of second detailed data into the text recognition model;

[0329] Through the text recognition model, feature extraction is performed on multiple object names respectively to obtain multiple text vectors; wherein, a text vector is obtained by extracting features from an object name through the text recognition model;

[0330] Obtain a vector clustering model, input multiple text vectors into the vector clustering model respectively, and cluster the multiple text vectors through the vector clustering model to obtain A text vector clusters;

[0331] According to A text vector clusters, multiple object identifiers are divided and processed to obtain A groups; wherein, the text vector corresponding to the object identifier in a group belongs to a text vector cluster.

[0332] In one possible implementation, group A includes group B c , c is a positive integer, and c is less than or equal to A;

[0333] The processing module 12 performs identification processing on the A groups in multiple information dimensions to obtain group attributes corresponding to the A groups, which are used to perform the following operations:

[0334] In multiple information dimensions, group B cPerform recognition processing to obtain recognition results corresponding to multiple information dimensions;

[0335] If there are abnormal recognition results among multiple recognition results, the abnormal group attribute is determined to be group B c Corresponding group attributes;

[0336] If multiple recognition results are all normal recognition results, the normal group attribute is determined to be group B c The corresponding group attributes.

[0337] In one possible implementation, the plurality of information dimensions include a name information dimension, and the plurality of recognition results include a name recognition result;

[0338] The processing module 12 analyzes the group B in multiple information dimensions. c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0339] Among the plurality of second detailed data, obtain group B c The object name corresponding to the object identifier in the object name is determined as an object name set;

[0340] Get the reference name with the exception attribute, match the reference name with the object name in the object name collection;

[0341] If there is an object name in the object name set that successfully matches the reference name, the abnormal recognition result is determined as the name recognition result;

[0342] If there is no object name in the object name set that successfully matches the reference name, the normal recognition result is determined as the name recognition result.

[0343] In one possible implementation, the plurality of information dimensions include a registration information dimension, and the plurality of recognition results include a registration recognition result;

[0344] The processing module 12 analyzes the group B in multiple information dimensions. c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0345] Among the plurality of second detailed data, obtain group B c The registration information corresponding to the object identifier in the object identifier is determined as a registration information set;

[0346] In the registration information set, obtaining a maximum number of identical registration information, and determining a ratio between the maximum number and the total number of registration information in the registration information set as an abnormal information ratio;

[0347] Compare the abnormal information ratio with the abnormal information ratio threshold. If the abnormal information ratio is greater than the abnormal information ratio threshold, the abnormal recognition result is determined as the registered recognition result.

[0348] If the abnormal information ratio is equal to or less than the abnormal information ratio threshold, the normal recognition result is determined as the registered recognition result.

[0349] In one possible implementation, the multiple information dimensions include a business information dimension, and the multiple identification results include a business identification result;

[0350] The processing module 12 analyzes the group B in multiple information dimensions. c Perform recognition processing to obtain recognition results corresponding to multiple information dimensions, which are used to perform the following operations:

[0351] Among the plurality of second detailed data, obtain group B c The business scope corresponding to the object identifier in the object is determined as a business scope set;

[0352] Perform similarity processing on every two business scopes in the business scope set to obtain D similarities corresponding to the business scope set; D is a positive integer;

[0353] Determine the mean similarity value corresponding to the D similarities, and compare the mean similarity value with the similarity threshold;

[0354] If the similarity mean is equal to or greater than the similarity threshold, the anomaly identification result is determined as the business identification result;

[0355] If the similarity mean is less than the similarity threshold, the normal recognition result is determined as the business recognition result.

[0356] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0357] From the above, it can be seen that the embodiment of the present application can identify abnormal groups with related relationships through group diffusion, and can transform small groups into large groups by merging abnormal groups with related relationships, so the audit efficiency of abnormal object identification can be improved. By determining larger-scale abnormal groups, large-scale abnormal groups can be processed first, so audit resources can be allocated more reasonably, thereby improving resource utilization.

[0358] Further, see Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. ​ As shown, the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. In some embodiments, the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As ​ As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0359] exist ​ In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0360] Acquire multiple abnormal groups; the multiple abnormal groups all refer to groups including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other;

[0361] Acquire relationship data including a source object identifier and an associated object identifier; the source object identifier is an abnormal object identifier included in multiple abnormal groups; the associated object identifier is an object identifier that has an associated relationship with the source object identifier;

[0362] Through relational data, diffusion processing is performed on multiple abnormal groups respectively to obtain multiple abnormal diffusion groups; an abnormal diffusion group is obtained by diffusion processing an abnormal group through relational data, and the diffusion object identifiers included in an abnormal diffusion group belong to the associated object identifiers;

[0363] Performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups;

[0364] According to the same diffusion object identifier contained in the abnormal diffusion group, two abnormal diffusion groups in the abnormal diffusion group are merged to obtain an abnormal merged group.

[0365] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method or device in the above embodiments, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here.

[0366] The present application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the data processing method or apparatus described in the preceding embodiments, which are not described in detail here. Furthermore, the description of the beneficial effects of the same method is not described in detail here.

[0367] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0368] The present application also provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, enabling the computer device to perform the data processing methods or apparatuses described in the preceding embodiments, which are not further detailed here. Furthermore, the beneficial effects of the same methods are not further detailed here.

[0369] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0370] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0371] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A data processing method, characterized in that: include: Acquire multiple abnormal groups; each of the multiple abnormal groups refers to a group including abnormal object identifiers; the abnormal object identifiers included in the multiple abnormal groups are different from each other; Acquire relationship data including a source object identifier and an associated object identifier; the source object identifier belongs to an abnormal object identifier included in the multiple abnormal groups; The associated object identifier refers to an object identifier that has an associated relationship with the source object identifier; Performing diffusion processing on the multiple abnormal groups respectively through the relationship data to obtain multiple abnormal diffusion groups; An abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through the relationship data, and the diffusion object identifiers included in the abnormal diffusion group belong to the associated object identifiers; performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups; According to the same diffusion object identifier contained in the abnormal diffusion group, two abnormal diffusion groups in the abnormal diffusion group are merged to obtain an abnormal merged group.

2. The method according to claim 1, characterized in that The plurality of abnormal groups include abnormal group E f , f is a positive integer, and f is less than or equal to the total number of the plurality of abnormal groups; the abnormal group E f Including the exception object identifier G h , h is a positive integer, and h is less than or equal to the abnormal population E f The total number of abnormal object identifiers in the; the total number of the source object identifiers and the total number of the associated object identifiers are both Z, where Z is a positive integer; there is an association relationship between a source object identifier and an associated object identifier; The step of performing diffusion processing on the plurality of abnormal groups respectively through the relationship data to obtain a plurality of abnormal diffusion groups includes: The abnormal object identifier G h Match with Z source object identifiers, if there is a match with the abnormal object identifier G in the Z source object identifiers h If the source object identifier is the same as the abnormal object identifier G, then among the Z associated object identifiers, obtain the h The identifier of the associated object with which the association relationship exists; The exception object identifier G h The associated object identifier with an associated relationship is determined as the abnormal object identifier G h The diffusion object identifier; The abnormal object identifier G h The diffusion object identifier is added to the anomaly group E f , the abnormal object identifier G will be added h The abnormal group E identified by the diffusion object f , identified as the abnormal group E f The corresponding abnormal diffusion group.

3. The method according to claim 1, characterized in that The plurality of abnormal diffusion groups include a first abnormal diffusion group and a second abnormal diffusion group; The performing object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification includes: comparing the diffusion object identifiers in the first abnormal diffusion group with the diffusion object identifiers in the second abnormal diffusion group; If the first abnormal diffusion group and the second abnormal diffusion group include the same diffusion object identifier, determining the first abnormal diffusion group and the second abnormal diffusion group as an abnormal diffusion group group; The same diffusion object identifiers included in the first abnormal diffusion group and the second abnormal diffusion group are determined as the same diffusion object identifiers of the abnormal diffusion group group.

4. The method according to claim 1, wherein The plurality of abnormal diffusion groups include Y object identifiers; Y is a positive integer greater than 1; The method further comprises: Determine identification scores corresponding to the Y object identifications respectively; Then, according to the same diffusion object identifier contained in the abnormal diffusion group group, two abnormal diffusion groups in the abnormal diffusion group group are merged to obtain an abnormal merged group, including: Determining the same diffusion object identifiers contained in the abnormal diffusion group as the common object identifier of the abnormal diffusion group; Obtaining the public identification score of the public object identification from the Y identification scores; According to the common identification score, two abnormal diffusion groups in the abnormal diffusion group are merged to obtain an abnormal merged group.

5. The method according to claim 4, characterized in that The step of merging two abnormal diffusion groups in the abnormal diffusion group according to the public identification score to obtain an abnormal merged group includes: Performing mean processing on the public identification scores to obtain the public scores corresponding to the abnormal diffusion group; Comparing the common score with a merge score threshold, and if the common score is equal to or greater than the merge score threshold, merging two abnormal diffusion groups in the abnormal diffusion group to obtain an abnormal merged diffusion group; The diffusion object identifier in the abnormal merged diffusion group is deleted, and the abnormal merged diffusion group after the diffusion object identifier is deleted is determined as the abnormal merged group.

6. The method according to claim 5, characterized in that The total number of the abnormal diffusion group groups is at least two, and the at least two abnormal diffusion group groups include a first abnormal diffusion group group and a second abnormal diffusion group group different from the first abnormal diffusion group group; the common score includes a first common score corresponding to the first abnormal diffusion group group and a second common score corresponding to the second abnormal diffusion group group; If the common score is equal to or greater than the merge score threshold, merging the two abnormal diffusion groups in the abnormal diffusion group to obtain an abnormal merged diffusion group includes: If the first common score is equal to or greater than the merge score threshold, and the second common score is equal to or greater than the merge score threshold, comparing the two abnormal diffusion groups in the first abnormal diffusion group and the two abnormal diffusion groups in the second abnormal diffusion group; If there is a same abnormal diffusion group in the first abnormal diffusion group and the second abnormal diffusion group, the two abnormal diffusion groups in the first abnormal diffusion group and the two abnormal diffusion groups in the second abnormal diffusion group are merged to obtain an abnormal merged diffusion group.

7. The method according to claim 4, characterized in that Determining the identification scores corresponding to the Y object identifications respectively includes: constructing an original topological map according to the multiple abnormal groups; Performing diffusion processing on the original topology map according to the multiple abnormal diffusion groups to obtain a diffusion topology map; the diffusion topology map includes Y nodes; one node is used to represent one object identifier among the Y object identifiers; Performing node score processing on the diffusion topology graph to obtain node scores corresponding to the Y nodes respectively; The Y node scores are determined as identification scores corresponding to the Y object identifications respectively.

8. The method according to claim 7, characterized in that The plurality of abnormal groups include Z abnormal object identifiers; the Z abnormal object identifiers belong to the Y object identifiers; Z is a positive integer greater than 1; The constructing of an original topological map according to the plurality of abnormal groups includes: Generating Z nodes for representing the Z abnormal object identifiers; wherein one node is used to represent one abnormal object identifier; Determine an edge score between every two nodes in the Z nodes according to an association relationship between every two abnormal object identifiers in the Z abnormal object identifiers; An original topological graph is constructed according to the Z nodes and the edge scores between every two nodes.

9. The method according to claim 8, characterized in that The Z abnormal object identifiers include a first abnormal object identifier and a second abnormal object identifier different from the first abnormal object identifier; Determining the edge score between every two nodes in the Z nodes according to the association relationship between every two abnormal object identifiers in the Z abnormal object identifiers includes: Obtaining first registration information corresponding to the first abnormal object identifier and second registration information corresponding to the second abnormal object identifier; performing similarity processing on the first registration information and the second registration information to obtain a similarity value between the first registration information and the second registration information; The similarity value between the first registration information and the second registration information is determined as an edge score between a first node and a second node; the first node is used to represent the first abnormal object identifier; the second node is used to represent the second abnormal object identifier; the first node and the second node both belong to the Z nodes.

10. The method according to claim 7, characterized in that The relationship data also includes an association weight for characterizing the degree of association between the source object identifier and the associated object identifier; the multiple abnormal diffusion groups include abnormal object identifiers I j , and the abnormal object identifier I j Diffusion object identifier K j ; j is a positive integer, and j is less than or equal to the total number of diffusion object identifiers in the plurality of abnormal diffusion groups; the diffusion object identifier K j Belongs to the associated object identifier; the abnormal object identifier I j an identifier belonging to the source object; The step of performing diffusion processing on the original topology map according to the plurality of abnormal diffusion groups to obtain a diffusion topology map includes: Generate a K for characterizing the diffusion object j The third node of In the relational data, obtain the identifier I used to characterize the abnormal object j And the diffusion object identifier K j The association weight of the degree of association between them; The abnormal object is identified as I j And the diffusion object identifier K j The association weight between them is determined as the edge score between the third node and the fourth node; the fourth node refers to the abnormal object identifier I in the original topology graph. j Node; The third node and the edge score are added to the original topology graph, and the original topology graph with the third node and the edge score added is determined as a diffusion topology graph.

11. The method according to claim 1, wherein The method further comprises: Acquire first detailed data corresponding to a plurality of resource management identifiers, merge first detailed data with the same object identifier among the plurality of first detailed data to obtain a plurality of second detailed data; wherein the object identifiers corresponding to the plurality of second detailed data are different from each other; Divide the plurality of object identifiers according to the object names respectively included in the plurality of second detailed data to obtain A groups; A is a positive integer; the object identifiers respectively included in the A groups are different from each other; Identify the A groups separately in multiple information dimensions to obtain group attributes corresponding to the A groups; Among the A groups, the group whose group attribute is an abnormal group attribute is determined as an abnormal group.

12. The method according to claim 11, characterized in that The plurality of object identifiers are divided according to the object names respectively included in the plurality of second detailed data to obtain A groups, including: obtaining a text recognition model, and inputting the object names respectively included in the plurality of second detailed data into the text recognition model; Using the text recognition model, feature extraction is performed on multiple object names to obtain multiple text vectors; wherein a text vector is obtained by extracting features from one object name using the text recognition model; Obtaining a vector clustering model, inputting the plurality of text vectors into the vector clustering model respectively, and performing clustering processing on the plurality of text vectors using the vector clustering model to obtain A text vector clusters; According to the A text vector clusters, multiple object identifiers are divided and processed to obtain A groups; wherein the text vectors corresponding to the object identifiers in a group belong to a text vector cluster.

13. The method according to claim 11, characterized in that The group A includes group B c , c is a positive integer, and c is less than or equal to A; The identification processing is performed on the A groups in multiple information dimensions to obtain group attributes corresponding to the A groups, including: The group B is analyzed in multiple information dimensions. c Performing recognition processing to obtain recognition results corresponding to the multiple information dimensions respectively; If there is an abnormal recognition result among the multiple recognition results, the abnormal group attribute is determined to be the group B c Corresponding group attributes; If multiple recognition results are all normal recognition results, the normal group attribute is determined to be the group B c The corresponding group attributes.

14. The method according to claim 13, characterized in that The multiple information dimensions include a name information dimension, and the multiple recognition results include a name recognition result; The group B is analyzed in multiple information dimensions. c Performing recognition processing to obtain recognition results corresponding to the multiple information dimensions, including: In the plurality of second detailed data, the group B is obtained c The object name corresponding to the object identifier in the object name is determined as an object name set; Obtaining a reference name with an abnormal attribute, and matching the reference name with an object name in the object name set; If there is an object name in the object name set that successfully matches the reference name, determining the abnormal recognition result as the name recognition result; If there is no object name in the object name set that successfully matches the reference name, a normal recognition result is determined as the name recognition result.

15. The method according to claim 13, characterized in that The multiple information dimensions include a registration information dimension, and the multiple recognition results include a registration recognition result; The group B is analyzed in multiple information dimensions. c Performing recognition processing to obtain recognition results corresponding to the multiple information dimensions, including: In the plurality of second detailed data, the group B is obtained c The registration information corresponding to the object identifier in the object identifier is determined as a registration information set; In the registration information set, obtaining a maximum number of identical registration information, and determining a ratio between the maximum number and a total number of registration information in the registration information set as an abnormal information ratio; Comparing the abnormal information ratio with an abnormal information ratio threshold, and if the abnormal information ratio is greater than the abnormal information ratio threshold, determining the abnormal identification result as the registered identification result; If the abnormal information ratio is equal to or less than the abnormal information ratio threshold, the normal recognition result is determined as the registered recognition result.

16. The method according to claim 13, characterized in that The multiple information dimensions include a business information dimension, and the multiple identification results include a business identification result; The group B is analyzed in multiple information dimensions. c Performing recognition processing to obtain recognition results corresponding to the multiple information dimensions, including: In the plurality of second detailed data, the group B is obtained c The business scope corresponding to the object identifier in the object is determined as a business scope set; Perform similarity processing on every two business scopes in the business scope set to obtain D similarities corresponding to the business scope set; D is a positive integer; Determine a similarity mean corresponding to the D similarities, and compare the similarity mean with a similarity threshold; If the similarity mean is equal to or greater than the similarity threshold, determining the abnormality identification result as the business identification result; If the similarity mean is less than the similarity threshold, the normal recognition result is determined as the service recognition result.

17. A data processing device, characterized in that: include: An acquisition module, configured to acquire a plurality of abnormal groups; each of the plurality of abnormal groups refers to a group including abnormal object identifiers; and the abnormal object identifiers respectively included in the plurality of abnormal groups are different from each other; The acquisition module is further configured to acquire relationship data including a source object identifier and an associated object identifier; the source object identifier belongs to an abnormal object identifier included in the multiple abnormal groups; The associated object identifier refers to an object identifier that has an associated relationship with the source object identifier; a processing module, configured to perform diffusion processing on the plurality of abnormal groups respectively according to the relationship data to obtain a plurality of abnormal diffusion groups; An abnormal diffusion group is obtained by performing diffusion processing on an abnormal group through the relationship data, and the diffusion object identifiers included in the abnormal diffusion group belong to the associated object identifiers; The processing module is further configured to perform object identification processing on every two abnormal diffusion groups in the plurality of abnormal diffusion groups to obtain an abnormal diffusion group group having the same diffusion object identification; the abnormal diffusion group group includes two different abnormal diffusion groups; The processing module is further configured to merge two abnormal diffusion groups in the abnormal diffusion group according to the same diffusion object identifier contained in the abnormal diffusion group to obtain an abnormal merged group.

18. A computer device, characterized in that: include: processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 16.

20. A computer program product, characterized in that The computer program product comprises a computer program stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 16.