Object recognition method, apparatus, electronic device, and storage medium

CN115859187BActive Publication Date: 2026-09-25TENPAY PAID TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111109153.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2026-09-25
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

[0003]目前,对于风险用户的识别,通常是借助于其他用户报损、用户自身的交易行为(比如与用户交易相关联的商户等)等,虽然该方式能够识别出一些风险用户,但是该方式的时效性较差、且识别覆盖率具有很大的局限性

Benefits of technology

本申请实施例提供的方案,在对待识别对象进行识别时,同时考虑了待识别对象自身的对象相关数据和该对象与其他对象之间的关联关系,由于一个对象的对象相关数据反映了该对象的特征,而不同对象类型的对象的特征通常是不同的,因此,可以基于待识别对象的对象相关数据来初步评估该对象的对象类型。而一个对象与其他对象之间的关联关系会对该对象产生影响,因此,本申请实施例的方法,进一步考虑对象之间的关联关系、以及各对象自身的标签(即待识别对象的第一标签、第一样本对象的标注标签和第二标签),可以在基于待识别对象的对象相关数据预测出的该对象的第一标签的基础上,融入对象之间的相互影响,从而得到更加准确的识别结果。此外,由于本申请的该方法,无需依赖对象的投诉、报损,可以实现对象的提前预防识别,更好的满足了时效性的要求,尤其是风险识别领域对于时效性的要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115859187B_ABST
    Figure CN115859187B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an object identification method and device, electronic equipment and storage medium, and relate to the fields of financial payment, payment security, big data, cloud technology, block chain, vehicle-mounted terminal and artificial intelligence. The method comprises: obtaining object-related data of to-be-identified objects; predicting a first label of each object through an object identification model based on the object-related data of each to-be-identified object, obtaining a reference data set comprising object-related data of a plurality of first sample objects with labeled labels and second labels, determining a first association relationship between each object in the to-be-identified objects and the first sample objects according to the object-related data of the to-be-identified objects and the first sample objects; and obtaining an identification result of the to-be-identified objects according to the first label of the to-be-identified objects, the labeled labels and the second labels of the first sample objects, and the first association relationship. Based on the method, the object type of unknown objects can be identified in a timely, accurate and effective manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of mobile payment, payment security, big data, vehicle terminals and artificial intelligence. Specifically, this application relates to an object recognition method, device, electronic device and storage medium. Background Technology

[0002] With the rapid development of science and technology, online payments and transfers have become commonplace in people's lives. While science and technology bring convenience, the forms and methods of online fraud are also constantly evolving. Effectively preventing and avoiding various forms of commercial fraud, and identifying users engaging in fraudulent activities, has always been a crucial research issue for relevant technical personnel.

[0003] Currently, the identification of risky users usually relies on other users' loss reports and the user's own transaction behavior (such as merchants associated with the user's transactions). Although this method can identify some risky users, it has poor timeliness and a very limited coverage. Summary of the Invention

[0004] To address at least one of the problems existing in the prior art, embodiments of this application provide an object recognition method, apparatus, electronic device, and storage medium, which can better meet the requirements of timeliness and coverage in object recognition.

[0005] To achieve the above objectives, the solutions provided in this application are as follows: On one hand, embodiments of this application provide an object recognition method, the method comprising: Obtain object-related data for at least one object to be identified; For each object to be identified, the object's first label is predicted by the object recognition model based on the object's object-related data. The first label of an object represents the object type to which the object belongs among multiple object types. Obtain a reference dataset, which includes object-related data and second labels for multiple first sample objects with labeled tags. The labeled tag of a first sample object represents the true object type to which the object belongs among multiple object types, and the second label of an object represents the probability of the object belonging to each object type among multiple object types. Based on the object-related data of each object to be identified and each first sample object, determine at least one first association relationship between each object in the object to be identified and multiple first sample objects; The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship; For each object to be identified, the identification result is determined based on the object's second label.

[0006] On the other hand, embodiments of this application provide an object recognition device, which includes: The first prediction module is used to acquire object-related data of at least one object to be identified; for each object to be identified, the first label of the object is predicted by the object recognition model based on the object-related data of the object, and the first label of an object represents the object type to which the object belongs among multiple object types; The reference dataset acquisition module is used to acquire a reference dataset, which includes object-related data and second labels for multiple first sample objects with labeled tags. The labeled tag of a first sample object represents the true object type to which the object belongs among multiple object types, and the second label of an object represents the probability of the object belonging to each object type among multiple object types. The second prediction module is used to determine a first association relationship between at least one object to be identified and each object in multiple first sample objects based on object-related data of each object to be identified and each first sample object, and to determine a second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship. The recognition result determination module is used to determine the recognition result of each object to be recognized based on the second label of each object to be recognized.

[0007] Optionally, the second prediction module can be specifically used for: The first label of each object to be identified is used as the annotation label and the initial second label of the object to be identified. Based on the annotation label and the second label of each object to be identified and the first sample object, and based on the first association relationship, at least one label propagation is performed between the object to be identified and the first sample object to obtain the updated label of each object to be identified and the first sample object. For each object to be identified, the updated labels of each object with the first association relationship are merged according to the first association relationship to obtain the second label of the object.

[0008] Optionally, the second prediction module may perform the following operations during each label propagation: For each object in the object to be identified and the first sample object, the second label of the object is updated based on the second label of each object that has a relationship with the object, according to the first association relationship; for each object, the updated label of the object is obtained by fusing the updated second label of the object and the labeled label of the object, and the updated label of the object is used as the second label of the object in the next label propagation.

[0009] Optionally, the object-related data includes at least one specified type of object-related data, and the first association includes an association of one type corresponding to each specified type of object-related data; correspondingly, the second prediction module can be used for: Obtain the weight corresponding to each type of association; determine the second label of each object to be identified based on the first label of each object to be identified, the labeled label and second label of each first sample object, each type of association, and the weight corresponding to each type of association. Optionally, the second prediction module can be used to: for each object among at least one object to be identified and multiple first sample objects, determine the influence of the object based on the object-related data of the object; determine the second label of each object to be identified based on the first label of each object to be identified, the labeled label and second label of each first sample object, the influence of each object to be identified and the first sample objects, and the first association.

[0010] Optionally, the object-related data includes at least one specified type of object-related data, the first association includes an association of a type corresponding to each specified type of object-related data, and the influence of each of the at least one object to be identified and a plurality of first sample objects, including the influence of each object corresponding to each type of association.

[0011] Optionally, the second prediction module can be used to: determine the proportion of each object type in at least one object to be identified and multiple first sample objects based on the first label of each object to be identified and the annotation label of each first sample object; use the proportion of each object type as a weight to weight the first label of the corresponding object type in at least one object to be identified and to weight the annotation labels of the corresponding object type in multiple first sample objects; and determine the second label of each object to be identified based on the weighted first label of each object to be identified, the weighted annotation label and second label of each first sample object, and the first association relationship.

[0012] Optionally, the object recognition model is obtained by the model training module by performing the following operations: Obtain a first training dataset, which includes object-related data of multiple second sample objects with labeled labels and object-related data of multiple unlabeled third sample objects. The multiple second sample objects include multiple objects of each of multiple object types. Based on object-related data of multiple second sample objects, the initial classification model is trained until the first training termination condition is met, resulting in the first classification model. For each third sample object, based on the object-related data of that object, the object type of the object is predicted by the first classification model, and the label of the object is determined according to the object type. Based on the object-related data of multiple second sample objects and the object-related data of multiple third sample objects with labels, the first classification model is trained again until the second training termination condition is met, resulting in the object recognition model.

[0013] Optionally, the reference dataset is obtained by the reference dataset acquisition module in the following ways: Obtain a second training dataset, which includes object-related data of multiple first sample objects with labeled tags; determine the second association relationship between objects in the second training dataset based on the object-related data of each first sample object; use the labeled tag of each first sample object as the initial third tag of the object, and repeat the following operations until the updated third tags of multiple first sample objects meet the preset conditions, and determine the third tag of each first sample object that meets the preset conditions as the second tag of the object; based on the second association relationship and the labeled tags and third tags of each first sample object, obtain the updated fourth tag of each first sample object by performing tag propagation among multiple first sample objects; for each first sample object, obtain the new third tag of the object by fusing the fourth tags of each first sample object that has an association relationship with the object, according to the second association relationship.

[0014] Optionally, the reference dataset acquisition module can also be used for: After each label propagation, new data is acquired, including object-related data of at least one sample object with labeled tags; each sample object in the new data is taken as the new first sample object, and the second training dataset is updated based on the new data; based on the object-related data of each first sample object in the updated second training dataset, the second association relationship between each object in the updated second training dataset is determined, and the updated second association relationship is obtained. When the reference dataset acquisition module obtains the updated fourth label for each first sample object, it can be used for: The annotation label of each newly added first sample object is used as the third label of that object. Based on the updated second association relationship, as well as the updated annotation labels and third labels of each first sample object, the updated fourth label of each first sample object is obtained by propagating the labels among the updated multiple first sample objects.

[0015] Optionally, the labels for each sample object in the newly added data are obtained in the following way: Obtain object-related data for at least one unlabeled object, wherein at least one sample object includes at least one unlabeled object; for each of the at least one unlabeled objects, based on the object-related data of that object, predict the first label of that object through an object recognition model, and use the first label of that object as the label of that object.

[0016] Optionally, for each label propagation, the reference dataset acquisition module is also used for: Based on object-related data of multiple first sample objects, similar object pairs are determined among the multiple first sample objects; wherein, satisfying preset conditions includes the value setting conditions of the loss function; The loss function includes a first loss function and a second loss function. For each label propagation, the value of the first loss function characterizes the difference between the labeled label of each first sample object and the new third label, while the value of the second loss function characterizes the difference between the new third labels of each pair of similar objects.

[0017] In another aspect, embodiments of this application provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the method provided in embodiments of this application.

[0018] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in embodiments of this application.

[0019] In another aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method provided in embodiments of this application.

[0020] In another aspect, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any of the optional embodiments of this application described above.

[0021] The beneficial effects of the technical solutions provided in this application are: The solution provided in this application, when identifying an object, considers both the object's own object-related data and its relationships with other objects. Since object-related data reflects the object's characteristics, and the characteristics of different object types are usually different, the object type can be initially assessed based on the object-related data. Furthermore, the relationships between an object and other objects influence that object. Therefore, the method in this application further considers the relationships between objects and the labels of each object (i.e., the first label of the object to be identified, the annotation label of the first sample object, and the second label). It can incorporate the mutual influence between objects based on the first label predicted from the object-related data of the object to be identified, thereby obtaining a more accurate identification result. In addition, since this method does not rely on object complaints or damage reports, it can achieve early preventative identification of objects, better meeting the timeliness requirements, especially in the field of risk identification. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0023] Figure 1 A flowchart illustrating an object recognition method provided in an embodiment of this application; Figures 2a to 2d A schematic diagram of several object types provided in the examples of this application; Figure 3 This is a schematic diagram of the structure of an object recognition system provided in an embodiment of this application; Figure 4 A flowchart illustrating an object recognition method provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the principle of a training method for an object recognition model provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the principle of tag propagation provided in the example of this application; Figures 7a to 7c Schematic diagrams illustrating several different examples of label propagation provided for the purposes of this application; Figure 8 This is a schematic diagram of the structure of an object recognition device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device to which this application is applicable. Detailed Implementation

[0024] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0025] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term, including all or any unit and all combinations of one or more associated listed items, for example, “A and / or B” indicates implementation as “A,” or implementation as “A,” or implementation as “A and B.”

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. To better understand the relevant technology, some technical terms used in this application will first be introduced: This application addresses the problems existing in current methods for identifying target types (such as risky targets, i.e., users / objects engaging in fraudulent activities, referring to users who profit through illegal / socially unethical means), and proposes a method for identifying targets (i.e., targets with fraud risk, referring to the transaction risk of black market operators illegally obtaining user assets through inducement, false information, etc.). Currently, the identification of risky users often relies on other users' loss reports and the user's own transaction behavior. User risk tags (marks of users engaging in fraudulent activities) are fragmented, and when identifying fraud risk, only a single user risk tag is used to correlate risk identification with other users or merchants. In past practices, users have only served as a medium for the transmission of single risks, and the maintenance of user tags is costly, time-consuming, and labor-intensive. Existing methods for identifying risky targets have at least the following problems: 1) Poor timeliness: Throughout the entire lifecycle of black market activities (black industry / illegal industry / malicious industry, referring to industries that profit through illegal / socially immoral means), black market activities often involve large-scale fraudulent activities occurring simultaneously. Relying on identification methods based on other users' reports of losses means that when a risky user is flagged, other merchants may have already completed the entire fraud process, resulting in a large number of reported losses. This lack of early prevention significantly impacts control over black market funds.

[0027] 2) Insufficient Coverage: Since most fraudulent activities are currently based on internet technology, the cost of registering an account is almost zero. In order to carry out fraudulent transactions and fund transfers more quickly and efficiently, black market operators often possess a large number of accounts. However, the solution of identifying risky users by relying on customer complaints and associating with black market merchants (merchants with fraudulent activities) has significant limitations and cannot comprehensively cover black market accounts.

[0028] 3) Weak correlation. Existing user risk tagging systems are often independent of each other based on different business scenarios. Although the sources of clues vary in the process of user risk identification, extensive practice has shown that different risky users may play different roles in the same fraud case, and there are also subtle connections between different risky users such as social information and transaction behavior. However, existing identification methods cannot achieve correlation identification in different business scenarios.

[0029] To address at least one of the problems existing in the prior art and to better meet the needs of risk identification, this application provides a new object identification method. Based on this method, a risk user relationship network can be built, which not only helps to construct a user risk system, but also clarifies the life cycle of black market activities, providing a new path for the early identification of fraud risks.

[0030] Optionally, the object recognition method provided in this application embodiment can be applied to big data processing, such as by implementing it based on cloud technology. The data computation involved in this application embodiment can be performed using cloud computing. For example, the computation of steps such as training the object recognition model and determining the object's label based on label propagation can be performed using cloud computing.

[0031] Big data refers to data sets that cannot be captured, managed, and processed within a certain timeframe using conventional software tools. It represents massive, rapidly growing, and diverse information assets that require new processing models to achieve stronger decision-making, insightful discovery, and process optimization capabilities. With the advent of the cloud era, big data has attracted increasing attention. Big data requires specialized technologies to effectively process large amounts of data within a tolerable timeframe. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the internet, and scalable storage systems. Among these, cloud technology, based on the cloud computing business model, encompasses network technologies, information technologies, integration technologies, management platform technologies, and application technologies. It can form resource pools, providing on-demand, flexible, and convenient access. Cloud computing technology will become a crucial support.

[0032] Optionally, the solutions provided in this application can also be implemented based on Artificial Intelligence (AI) technology. For example, a trained risk identification model can predict the first risk label of an object, or machine learning can be used to obtain a reference dataset based on a loss function. Artificial intelligence technology is a comprehensive discipline involving a wide range of fields, including both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0033] Optionally, the data involved in this application embodiment (such as object-related data) can be stored using cloud storage or blockchain-based storage, which can effectively protect data security. Here, blockchain refers to a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0034] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0035] Figure 1 This illustration shows a flowchart of an object identification method provided in an embodiment of this application. The method can be executed by any electronic device, such as a server (e.g., a cloud server, a physical server, or a server cluster). The method can be implemented as an application or as a plugin or functional module of an existing application. For example, it can be a new functional module in a transaction-based application (e.g., mobile payment). The application's server can execute the method of this embodiment to identify the tags of the object to be identified, determining whether the object is a target type object (e.g., whether it is a non-risk object) and, if so, its risk type (object type, i.e., what kind of fraudulent behavior the object exhibits). The method can also be executed by a terminal device, which can identify the tags of the object to be identified and obtain the identification result. The terminal device includes a user terminal, including but not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. Optionally, in practical applications, to better ensure the security of object information, the method can be executed by a server.

[0036] like Figure 1 As shown in the figure, the object recognition method provided in this application embodiment may include the following steps S110-S140.

[0037] Step S110: Obtain object-related data for at least one object to be identified.

[0038] In this application embodiment, the objects may include, but are not limited to, users, merchants, etc. An object can be represented by its object identifier. The form of the object identifier is not limited in this application embodiment, as long as it is information that can uniquely represent an object. For example, it may include, but is not limited to, the object's contact information, the object's account identifier, etc. The object's account identifier may be the object's social media account, such as the object's account in the application (e.g., a user's registered account name, nickname, etc. in the application). For ease of description, in some of the embodiments described below, an object's account may be used to represent the object.

[0039] In this embodiment, object-related data for an object includes the object's interaction data. This object-related data can be the object's interaction behavior data (also called social behavior data), which refers to data related to the object's social interactions. Specifically, it can include data related to the object's interaction behavior with other objects. In practical applications, the specific social behavior data used can be configured according to requirements. Furthermore, object-related data can be the object's social behavior data obtained with the object's authorization.

[0040] Optionally, an object's social behavior data may include its social / interaction information and transaction information. Social information reflects the object's social activity level, such as the number of friends, the number of other objects following the object, or the number of people who forwarded or shared a post by the object, or the number of people on a forum. The criteria for determining friends are not limited in this application; for example, two objects that follow each other can be friends. An object's transaction information refers to information related to transactions between the object and other objects. Transaction information may include, but is not limited to, payment behavior information and transfer information (including payments / transfers from the object to other objects, and payments / transfers from other objects to the object). Specific transaction information may include, but is not limited to, transaction time, initiator and recipient (e.g., A transfers money to B, where A is the initiator and B is the recipient), transaction amount, and transaction type (whether it's a transfer, a red envelope, or other forms).

[0041] Step S120: For each object to be identified, based on the object-related data of the object, the first label of the object is predicted by the object recognition model. The first label of an object represents the object type to which the object belongs among multiple object types.

[0042] Among these, object type, also known as risk type, refers to the type of fraudulent behavior an object may be engaging in. The first label, also known as the first risk label, characterizes the risk type of the object predicted based on its relevant data.

[0043] Object recognition models (also known as risk recognition models) are neural network models pre-trained on a training dataset. The input to this model is object-related data, or pre-processed object-related data. The output is the object type corresponding to the object-related data. For example, the object-related data can be pre-processed into a fixed format, such as converting it into a vector of a specified format, before being input into the model, which then predicts the object type.

[0044] In this embodiment, the object recognition model can be a classification model, which can be a multi-classification model. Each object type in the multiple object types corresponds to a category in the classification model. Through this model, the category corresponding to the social behavior data can be predicted. The object type represented by this category is the object type of the object to which the social behavior data belongs. In practical applications, this embodiment does not limit the data format of the model output. For example, it can be a category identifier or a one-dimensional vector. The number of elements (i.e., numbers) in the vector is equal to the total number of types among the multiple object types mentioned above. Each element corresponds to a type, and the element value of each element can be 0 or 1. For example, if only one element has a value of 1 and the others are all 0, the type corresponding to the element with a value of 1 is the predicted object type, which is the first label mentioned above.

[0045] Furthermore, in practical implementation, the aforementioned multiple object types can include multiple target types and one non-target type. Each target type corresponds to a type of fraudulent behavior, i.e., a risk type, while the non-target type corresponds to a user without fraudulent behavior, i.e., a non-risk user. In other words, no risk can also be considered a risk type. If the model predicts that the risk type is no risk, then the initial identification result for that object considers it not a risky object. For example, if there are two object types, A and B (i.e., two target types), then the object identification model can be a three-class classification model. This model can predict whether an object is type A, type B, or a risk-free type (i.e., a non-target type).

[0046] This application does not limit the specific training method for the object recognition model. The training termination conditions described above can also be configured according to application requirements.

[0047] In an optional embodiment of this application, the object recognition model may be trained in the following manner: Obtain a first training dataset, which includes object-related data of multiple second sample objects with labeled labels and object-related data of multiple unlabeled third sample objects. The multiple second sample objects include multiple objects of each of multiple object types. Based on the object-related data of multiple second sample objects, the initial classification model is trained until the first training termination condition is met, and the first classification model is obtained. For each third sample object, based on the object-related data of that object, the object type of the object is predicted by the first classification model, and the label of the object is determined according to the object type. Based on object-related data from multiple second sample objects and object-related data from multiple third sample objects with labeled tags, the first classification model is trained again until the second training termination condition is met, thus obtaining the object recognition model.

[0048] Because the interactive behavior characteristics (social behavior characteristics) of objects differ in different scenarios, in order to ensure that different types of objects do not interfere with each other during the model learning process and lead to misjudgment, in this optional solution of the application, when training the object recognition model based on the training dataset, multiple training data of different object types are used for model training. That is, for each object type, the training dataset contains object-related data of multiple sample objects of that type. Through training, the model can learn the social behavior characteristics of objects of different object types from the object-related data of sample objects of different object types.

[0049] Furthermore, since obtaining labeled sample data usually requires manual intervention and the amount of sample data is typically limited, this alternative solution utilizes a semi-supervised learning approach for model training. The training dataset includes both labeled and unlabeled sample data. To ensure model training accuracy, labeled sample data is used for iterative training in the first stage of training, enabling the trained model to meet certain performance requirements, i.e., satisfying the first training termination condition. This condition can be configured according to actual needs; for example, if the model's prediction accuracy exceeds a set value, the model can predict the object type corresponding to the unlabeled sample data. The object-related data of the third sample object can be input into the first classification model that satisfies the first training termination condition to obtain the first label of each third sample object. This label is then used as the annotation label (pseudo-label) of the third sample object. The model can then continue training based on labeled and pseudo-labeled sample data. When the model achieves the expected results, training can end, resulting in an object recognition model that meets application requirements. This model can preliminarily predict the first label of the object to be identified.

[0050] Step S130: Obtain a reference dataset, which includes object-related data and second labels for multiple first sample objects with labeled tags.

[0051] In this context, the label of a first sample object represents the true object type to which the object belongs among multiple object types, and the second label of an object represents the probability of the object belonging to each of the multiple object types.

[0052] To facilitate understanding, as an example, suppose there are 5 types of object types. An object's label can be represented as [1, 0, 0, 0, 0], and the second label can be represented as [p1, p2, p3, p4, p5]. Here, p1 to p5 represent the probability that the object is each of the 5 object types, and the sum of the 5 probabilities equals 1. The label indicates that the object's true object type is the object type corresponding to the element with a value of 1 among the 5 object types.

[0053] The reference dataset can be understood as a real sample dataset, which contains relevant data on objects with multiple known risk types, including object-related data, annotation labels, and secondary labels.

[0054] In this embodiment of the application, for each of the first sample objects, its label and second label can be understood as the real label of the object. The second label can be understood as the probability distribution of the object belonging to each of the multiple object types when the real object type of the sample object is the object type corresponding to the label.

[0055] In practical applications, fraudulent activities often involve multiple stages and may involve multiple different risky users (i.e., users / objects at risk). Throughout the lifecycle of a fraudulent activity, different risky users may act on different stages of the same fraudulent activity, and there are subtle connections between different risky users through social information, transaction behavior, etc. Therefore, a certain type of risky user is likely to be associated with risky users of the same or different types, and there can be transmission and mutual influence between users of different risk types. Therefore, in this embodiment, a label and a second label are used to reflect a user's own object type from two different levels, as well as the probability that the user belongs to each object type when considering the association between the user and other users. In other words, the second label is a risk label that takes into account the mutual influence between users. The specific method of obtaining the reference dataset is not limited in this embodiment.

[0056] Step S140: Based on the object-related data of each object to be identified and each first sample object, determine the first association relationship between at least one object to be identified and each object in the multiple first sample objects.

[0057] Step S150: Determine the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship.

[0058] Step S160: For each object to be identified, determine the identification result of the object based on the second label of the object to be identified.

[0059] The first association relationship between the at least one object to be identified and each of the multiple first sample objects includes the association relationship between the objects to be identified, and the association relationship between the object to be identified and the first sample objects. This association relationship can also be referred to as a social association relationship or an interactive association relationship.

[0060] Since object-related data contains interaction data between that object and other objects, the social relationships between two objects can be determined based on their object-related data. This application does not limit the granularity of the relationship classification. Optionally, the relationship between objects can include whether or not there is a relationship between them. It can also be further subdivided into different types of relationships. For example, object-related data can include multiple different types of object-related data, and the presence or absence of a particular type of relationship between objects can be determined based on each type of object-related data.

[0061] Optionally, object-related data for an object can include various types of data such as the object's transfer information, red envelope (sending or receiving red envelopes) information, and entity information corresponding to the object. Entity information refers to the entity information used by the object when engaging in social behavior, such as the object's contact information and transaction account (e.g., bank card number, virtual resource account). Based on the transfer information of each object, it can be determined whether there is a corresponding relationship between objects of this type of data; based on the red envelope information of each object, it can be determined whether there is a corresponding relationship between objects of this type. In other words, one type of behavioral data can correspond to one type of relationship. Of course, in practical applications, the relationship can also be unclassified. It can be determined whether there is a relationship between objects based on various types of object-related data. For example, if any type of object-related data of two objects indicates that the two objects are related, then the relationship between the objects can be determined.

[0062] In practical applications, since the social relationships between objects can affect the attribute information of objects, in the field of risk identification, if an object A is a risky object, such as an object with fraudulent behavior, and another ordinary object B (an object without risk) is associated with object A (for example, the two have had a payment transaction), then object B may also become a potentially risky object. That is, risk can spread due to the interaction information between objects. Considering this, the solution provided in this application further considers the relationship between objects when determining the identification result of the object to be identified, thereby improving the accuracy and comprehensiveness of object identification.

[0063] The object identification method provided in this application, when identifying an object whose risk is unknown, considers both the object's own social behavior data and its social relationships with other objects. Since social behavior data reflects the social characteristics of the object and other objects, and the social characteristics of risky objects are usually different from those of non-risky objects, and the social characteristics of objects belonging to different risk types are also usually different, the risk type of the object can be initially assessed based on its social behavior data. Furthermore, since the social relationships between an object and other objects affect that object, especially since risky objects affect objects with which they are related, further consideration is given to the social relationships between objects, as well as the risk labels of each object (i.e., the first risk label of the object to be identified, the label of the first sample object, and the second risk label). Based on the first risk label predicted from the social behavior data of the object to be identified, the mutual influence between objects can be incorporated to determine a more accurate second risk label for the object to be identified, thereby obtaining the object's risk assessment result based on this label.

[0064] Furthermore, the method provided in this application embodiment can automatically identify objects based on a reference dataset and object-related data of the object to be identified, without relying on the loss reports of other objects. Therefore, objects can be assessed whenever needed, better meeting the timeliness requirements in practical applications. It can predict risky objects in advance, allowing for preventative measures based on the identification results. For example, if an object is identified as risky, other objects transacting with that object can receive risk warnings to prevent falling into fraud traps. Risky objects can also be subject to appropriate control, or identified risky objects can be further verified manually for proactive prevention and crackdown. Moreover, when conducting risk assessment, the method in this application embodiment can leverage the relationships between objects to achieve a more comprehensive risk assessment, effectively expanding the coverage of risk object assessment.

[0065] After obtaining the second label of the object to be identified, the identification result of the object can be determined based on the label. This identification result can include whether the object is a risk object, i.e., whether it belongs to the target type; if the object is a risk object, its object type(s). Alternatively, the second label can be directly used as the identification result of the object to be identified, and the probability of the object belonging to each object type can be obtained through the label. Optionally, the object type corresponding to the probability in the second label that is greater than or equal to a set threshold can be determined as the object type of the object to be identified, or the object type corresponding to the highest probability can be determined as the object type of the object to be identified. If the object type with the highest probability is not risky, then the object can be considered to be currently a non-risky object, i.e., not a target type. Of course, further evaluation can be performed on objects that are not risky.

[0066] In an optional embodiment of this application, determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship may include: The first label of each object to be identified is used as the annotation label and the initial second label of the object to be identified. Based on the annotation label and the second label of each object to be identified and the first sample object, and based on the first association relationship, at least one label propagation is performed between the object to be identified and the first sample object to obtain the updated label of each object to be identified and the first sample object. For each object to be identified, the updated labels of all objects that have the first association relationship with the object are merged to obtain the second label of the object.

[0067] In this alternative approach, a second label for the object to be identified can be obtained through label propagation. Since related objects influence each other, if one object is a risky object, its risk type (label) may also be propagated to other related objects. In other words, related objects are more likely to be risky objects themselves. Therefore, given that each object has its own label (the first label of the object to be identified, and the annotation labels and second labels of the sample objects), at least one label propagation can be performed based on the relationships between objects. Afterward, for the object to be identified, its second label can be obtained by fusing the labels of all related objects (including the sample objects and the object to be identified).

[0068] Label propagation algorithms are graph-based semi-supervised learning methods that leverage the information transferability of knowledge graphs to propagate label information along behavioral paths. The basic idea is to use the label information of labeled nodes to predict the label information of unlabeled nodes, with node labels being passed to other nodes based on the similarity between nodes. This optional solution in the embodiment of this application optimizes existing label propagation algorithms. For an object to be identified, its first label is predicted based on its related data. Then, based on the relationships between objects, risk labels are propagated between objects; that is, a risk label of one object can be propagated to other objects with which it has a relationship. The number of times label propagation is performed can be configured according to application requirements.

[0069] Each tag propagation includes the following operations: For each object in the object to be identified and the first sample object, the second label of the object is updated based on the second label of each object that has an association relationship with the object, according to the first association relationship; For each object, the updated label of the object is obtained by fusing the updated second label of the object and the annotation label of the object, and the updated label of the object is used as the second label of the object in the next label propagation.

[0070] Assuming the label propagation occurs once, for each of the at least one target object and multiple first sample objects, its second label can be updated based on the second labels of the objects with which it is associated. For example, the second labels of the objects with which it is associated can be merged (e.g., added and then standardized) to obtain the updated label. This updated label is then merged with the label of its object type (first risk label / labeling label) to obtain the merged label of that object, which is the updated label after this label propagation. Subsequently, for each target object, its second label is obtained by merging the merged risk labels of the objects with which it is associated.

[0071] If the label propagation is performed more than once, the above operation can be performed again based on the second label of each object (including the object to be identified and the first sample object) obtained in the previous propagation, and the second label of the object to be identified obtained in the last propagation can be used as the final second label.

[0072] In an optional embodiment of this application, the object-related data includes at least one specified type of object-related data, and the first association relationship includes an association relationship of one type corresponding to each specified type of object-related data; Accordingly, the second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, including: Obtain the weight corresponding to each type of association; The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, each type of association, and the weight corresponding to each type of association.

[0073] In this optional approach, the relationships corresponding to each type of object-related data can be determined according to its type. This allows for a more granular measurement of whether an object is associated with other objects in various social behaviors, thus more accurately and comprehensively representing an object's social relationships. The specific types included in the specified categories can be configured according to requirements; this application embodiment does not limit this. For example, object-related data can include multiple data types, and the specified type can be one or more of these types. This application embodiment also does not limit the specific method of classifying object-related data types; the classification rules for each data type can be set according to actual needs and application scenarios.

[0074] In practical applications, since different types of relationships have different degrees of influence, each type of relationship has its own corresponding weight in order to more accurately assess the relationships between objects. This allows relationships with different influence capabilities to play different roles in assessing risk objects, further improving the accuracy of object identification.

[0075] In an optional embodiment of this application, the method may further include: For each of at least one object to be identified and multiple first sample objects, determine the object's influence based on the object's object-related data; Accordingly, the second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, including: The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, the influence of each object to be identified and the first sample object, and the first association relationship.

[0076] In this context, an object's influence refers to the extent to which it can affect other objects, representing its social competence to some extent. In practical applications, the influence of different objects typically varies. For example, object-related data includes transfer information; a user who transfers to more than 30 accounts clearly has a significantly different influence than a user who transfers to only 2 accounts. Furthermore, the likelihood that an object's label with different levels of influence will affect other objects differs. Therefore, to more accurately assess the second label of the object to be identified, this alternative solution further considers the influence of each object.

[0077] Optionally, when determining the second label of an object to be identified based on label propagation, the label can be weighted using the influence of each object during each label propagation process. For example, if a single label propagation is performed, for each object in the object to be identified and the first sample objects, the second label (which is the initial second label, i.e., the first label, for the object to be identified) can be weighted using the influence of that object, and then a label propagation is performed based on the weighted label. If multiple label propagations are performed, the second label of the object obtained in the previous propagation can be weighted before each label propagation.

[0078] Optionally, the object-related data of an object includes at least one specified type of object-related data, the first association includes an association corresponding to a type of object-related data for each specified type, and the influence of each of the at least one object to be identified and the plurality of first sample objects includes the influence of each object corresponding to each type of association.

[0079] In other words, when classifying object-related data, the influence of each type of object-related data can be determined according to its type, thereby measuring the influence of an object in various social behaviors in a more granular way and representing the influence of an object more accurately and comprehensively.

[0080] Optionally, for each object, the final influence of the object can be obtained by fusing the influence corresponding to each type. For example, the influence corresponding to each type can be multiplied together.

[0081] In an optional embodiment of this application, the method may further include: Based on the first risk of each object to be identified and the label of each first sample object, determine the proportion of each object type in at least one object to be identified and multiple first sample objects; Accordingly, the second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, including: The proportion of objects of each object type is used as the weight to weight the first label of the corresponding object type in at least one object to be identified, and to weight the annotation labels of the corresponding object type in multiple first sample objects; The second label of each object to be identified is determined based on the weighted first label of each object to be identified, the weighted annotation label and second label of each first sample object, and the first association relationship.

[0082] For the aforementioned object to be identified and the second sample object, each object has its own corresponding object type, namely, the first label of the object to be identified and the annotation label of the second sample object. Since the number of objects under different object types usually varies, for a certain object type, if the number of objects belonging to that object type is larger, then the possibility of the label of that object type being propagated to the object to be identified will be greater. Therefore, in this optional embodiment of the present application, when determining the second label of the object to be identified, the proportion of the number of objects of each object type is further considered, and the object labels of the corresponding object types (the first label of the object to be identified and the annotation label of the second sample object) are weighted according to this proportion, so that the influence of the object label is positively correlated with the number of objects of the corresponding object type, which is more in line with the actual situation and can more accurately predict the second label of the object to be identified.

[0083] Optionally, in the label-based propagation processing method, the object labels of the objects to be identified and the first sample objects of the corresponding object type can be weighted according to the number of objects of each object type each time the label propagation is performed.

[0084] In an optional embodiment of this application, the reference dataset may be obtained in the following ways: Obtain a second training dataset, which includes object-related data of multiple first sample objects with labeled tags; Based on the object-related data of each first sample object, determine the second association relationship between objects in the second training dataset; Use the annotation label of each first sample object as the initial third label of that object, and repeat the following operations until the updated third labels of multiple first sample objects meet the preset conditions. Then, determine the third label of each first sample object that meets the preset conditions as the second label of that object: Based on the second association relationship and the third label of each first sample object, the updated fourth label of each first sample object is obtained by propagating the label among multiple first sample objects; for each first sample object, according to the second association relationship, the new third label of the object is obtained by fusing the fourth labels of each first sample object that has an association relationship with the object.

[0085] As described above, labels can propagate between different objects. If objects have engaged in social interactions, especially certain types of social interactions related to fraud, such as money transfers and payments, then the risk label of an object is very likely to be propagated to the objects that interacted with it. To better learn the propagation effect of labels between different objects and to predict the second label of the object to be identified, this alternative solution of the application, based on a large number of sample objects with labeled labels, considers the mutual influence between objects (i.e., the association between objects and the labeled labels of sample objects), and adopts a label propagation method between objects to update the labels of objects until, when preset conditions are met, the final updated label of each object is obtained based on the label propagation result. This label is used as the second label of the sample object. Since this label is an updated label that incorporates the label propagation effect between different objects under the premise of knowing the labeled labels of the objects, it is possible to further determine the second label of the object to be identified based on the labeled labels and second labels of these sample objects, after the first label of the object to be identified (which can be understood as the preliminary labeled label of the object to be identified) has already been predicted, and based on the association between the object to be identified and these sample objects, to propagate the labels between objects and further determine the second label of the object to be identified. For the specific operations of each tag propagation, please refer to the corresponding description above, which will not be repeated here.

[0086] In an optional embodiment of this application, after each tag propagation, the method further includes: Acquire new data, which includes object-related data for at least one sample object with labeled tags; Each sample object in the new data is used as the first new sample object, and the second training dataset is updated based on the new data. Based on the object-related data of each first sample object in the updated second training dataset, the second association relationship between each object in the updated second training dataset is determined, and the updated second association relationship is obtained. Accordingly, based on the second association relationship and the third label of each first sample object, the updated fourth label of each first sample object is obtained by propagating the label among multiple first sample objects, including: The annotation label of each newly added first sample object is used as the third label of that object. Based on the updated second association relationship and the updated third labels of each first sample object, the updated fourth label of each first sample object is obtained by propagating the label among the updated multiple first sample objects.

[0087] To improve the generalization ability of learning, when learning the effect of label propagation between sample objects, the training dataset can be updated by adding new sample data after each label propagation. This increases the number of sample data and incorporates more relationships between objects, making the risk labels of the learned sample objects more general.

[0088] In an optional embodiment of this application, the annotation labels of each sample object in the newly added data are obtained in the following way: Obtain object-related data for at least one unlabeled object, wherein at least one sample object includes at least one unlabeled object; For each object in at least one unlabeled object, based on the object's object-related data, the object's first label is predicted by the object recognition model, and the object's first label is used as the object's annotation label.

[0089] In practical applications, new data can be object-related data of manually labeled sample objects, or social behavior data of risky objects reported by the objects. Considering the labor costs and the volume of new data, in this alternative solution of the application, the label of the new data can be the first label predicted by a trained object recognition model, and this label can be used as the annotation label.

[0090] In an optional embodiment of this application, the method may further include: Based on the object-related data of multiple first sample objects, similar object pairs among the multiple first sample objects are identified; Among them, meeting the preset conditions includes the value of the loss function meeting the set conditions; The loss function includes a first loss function and a second loss function. For each label propagation, the value of the first loss function characterizes the difference between the labeled label of each first sample object and the new third label, while the value of the second loss function characterizes the difference between the new third risks of each similar object pair.

[0091] In this optional scheme, the first loss function can constrain the difference between the updated label of a sample object and its original label to be as close as possible, while the second loss function can constrain the updated labels of similar sample objects to be as similar as possible. This scheme allows label propagation learning to have good accuracy and generalization ability, better meeting application requirements. Optionally, when determining similar object pairs, the similarity between two objects can be determined based on specific types of object-related data in the object-related data. For example, if the similarity between specific types of object-related data of two objects is greater than a set value, then the two objects can be considered a similar object pair. The specific feature type(s) are not limited in this embodiment and can be configured according to actual needs; for example, it could be object transfer data.

[0092] The object identification method provided in this application proposes to construct a user risk system (user identification system) by building and disseminating user (i.e., object) tags, which can then be applied to identify fraud risks in advance, that is, to identify users with risks and the types of risks associated with those users.

[0093] The method provided in this application can be applied to the mobile payment field, where existing technologies often separate the identification of commercial fraud and social fraud risks. However, numerous case studies have revealed that black market accounts (i.e., risky users / merchants, referred to as risky users) play an indispensable role in both commercial and social fraud chains. Their main tasks include, but are not limited to, social media traffic generation, account nurturing, transaction guidance, and fund transfer (i.e., various target types and risk levels). Based on the method provided in this application, risky users can be identified from different scenarios, and then a tag propagation algorithm can be used to spread the risk among these users, constructing a user risk system. This user risk system can then be applied to fraud risk identification, providing a new path for uncovering suspicious black market activities.

[0094] To better understand and illustrate the solution provided in this application, a specific optional implementation method of this application will be described below in the context of a mobile payment scenario.

[0095] To facilitate understanding, we will first introduce the various stages involved in the illegal industry. Throughout the entire process of illegal fraud, it often relies on black market accounts (also known as illegal accounts / risky accounts, i.e., accounts of risky users / merchants, representing risky users) to achieve multiple stages such as attracting traffic, nurturing accounts, guiding transactions, and transferring funds (each stage corresponds to a different target type). The specific manifestations at different stages have the following different characteristics: 1) Traffic generation: such as Figure 2aAs shown, lead generation is a primary method used by illicit industries to find targets for fraud. Risky accounts typically leverage large internet platforms to publish a variety of highly attractive, enticing messages, disseminating these messages to ordinary users. Once users inquire about details, they begin to use pre-designed scams and persuasive tactics to carry out the fraud. These accounts are often dedicated to phishing; once a scam is successful, the account is deleted. Therefore, their social information (i.e., data related to the target account) differs significantly from that of legitimate social media accounts.

[0096] 2) Account nurturing: such as Figure 2b As shown, account farming often occurs in the early stages of a risky merchant's registration. To create the illusion of a healthy business, to reserve funds for later transfers, or to evade risk control regulations, illegal businesses often make multiple payments to the merchant in advance. These transactions are usually completed by a single account, either a few large or many small transactions, and the transaction records are untraceable. In some scenarios, these transactions may also be completed by multiple accounts, i.e., multiple people farming accounts.

[0097] 3) Guiding transactions: such as Figure 2c As shown, the behavior of guiding transactions often occurs in certain specific scenarios. While guiding users to make payments, the black market will also make payments to risky merchants, hiding among normal users. However, the transaction frequency and amount are higher than those of ordinary users. That is, risky accounts will guide ordinary users to make transactions (become scammed) by participating in transactions (guiding transactions).

[0098] 4) Fund transfer: including money laundering (an act of legalizing illicit gains), such as... Figure 2d As shown, because illegal businesses often operate multiple merchants simultaneously, withdrawn funds will flow into other risky merchants or other risky accounts at the same time. When a risky merchant is penalized, the illegal business may, to ensure its funds are not frozen, recover the funds reserved during the account-raising process through refunds, such as... Figure 2d As shown, a risky merchant refunds funds back to the corresponding account (the risky account shown in the figure). These accounts can then transfer funds to other accounts / merchants (the ellipses and arrows in the figure indicate that the accounts / merchants can further transfer funds), thereby transferring the illegal proceeds.

[0099] The method provided in the embodiments of this application will be described below in conjunction with the fraud scenarios involving multiple stages listed above.

[0100] Figure 3 This paper shows a schematic diagram of the structure of an object recognition system applicable to an embodiment of this application. Figure 4 A schematic diagram illustrating the implementation process of the object recognition method in this scenario is shown. Figure 3As shown, the system may include a server 10 and multiple terminal devices (only terminal devices 21 and 22 are shown in the figure). The terminal devices can communicate with the server 10 via a network. The sample object library 11 on the server 10 side stores a large amount of object-related data of the first sample objects with labels, which is the object-related data of the sample users. In other words, the sample object library 11 stores a reference dataset. Terminal devices 21 and 22 can be the terminal devices of the object to be identified, A and B, respectively. Optionally, the server 10 can be an application server with mobile payment and user interaction functions. The users of the terminal devices, i.e., the objects, can interact through the application, such as sending messages to each other, adding friends, etc., and can also conduct transactions and make mobile payments through the application. With user authorization, the server 10 can obtain the user's relevant information and achieve risk identification of the user by executing the method provided in the embodiments of this application.

[0101] like Figure 4 As shown, the optional implementation process of this method may include the following steps S1 to S5.

[0102] Step S1: Train the object recognition model based on the training dataset.

[0103] like Figures 2a to 2d As shown, risky accounts (representing malicious users, i.e., risky users) are present throughout the entire lifecycle of the black market. These risky accounts exhibit different characteristics in different scenarios. To ensure that different types of malicious users do not interfere with each other during model learning, leading to misjudgments, this embodiment can train the model separately according to the different types of risky accounts (i.e., different object types). The model training can be performed by server 10 or other electronic devices. Server 10 predicts the risk type of objects by calling the trained object recognition model. This embodiment uses the training device 30 as an example to illustrate the model training steps.

[0104] This solution utilizes semi-supervised learning for model training, and the specific implementation process is as follows: 1. Model Grouping: This involves classifying risky accounts into different risk types. First, based on the lifecycle of the illegal industry, different types of risky users (i.e., risky accounts) are grouped. For example, risky accounts responsible for fund transfers need to achieve a closed loop in fund inflows and outflows, thus sharing similar characteristics with accounts used for account nurturing. However, they exhibit behavioral differences at different time windows; nurturing accounts typically appear in the early stages. Therefore, time windows can be used to differentiate between the two types of risky users for model training. Similarly, risky accounts used for lead generation and those used for payment guidance are also trained separately. Of course, the training dataset also includes non-risky accounts, i.e., users who are not part of the target category.

[0105] This step can be performed manually or by electronic devices according to predefined classification rules. Through this step, accounts can be grouped and labeled according to their risk type based on different characteristics. A classification model can then be trained using object-related data from these labeled accounts to obtain an object recognition model.

[0106] 2. Sample acquisition: i.e., the second training dataset ( Figure 3 Construction of the training dataset 12 shown This step uses risky accounts (those already labeled with risk types) and normal accounts (those without risk, i.e., sample objects without risk) as the targets for model learning. The object-related data of these accounts (i.e., the interaction information between these accounts and other accounts, such as social information and payment behavior information) are used as feature variables for model recognition.

[0107] For example, payment behavior information refers to interaction information related to payments / transactions, which may include payments from one account to another, or payments from another account to that account. Social information, on the other hand, is interaction information other than payment behavior information, such as the account's friend information / friendship level, activity level, etc.

[0108] In real-world scenarios, risky accounts typically lure users into transactions through chat, posting virtual information, and other means. The object-related data of risky accounts differs significantly from that of normal social media accounts, and different types of risky accounts also exhibit different characteristics in their object-related data. Therefore, the object-related data of already labeled risky and normal accounts can be used as sample data to train the model.

[0109] The sample data may also include social behavior data of multiple accounts with unknown risk types (corresponding to the third sample object mentioned above).

[0110] 3. Model training: The above sample data is used to train the model. When the training meets certain conditions, the model (i.e. the first classification model mentioned above) is used to label accounts with unknown risk types. This yields the labeled accounts with unknown risk types, i.e., pseudo-labels.

[0111] During training, the model's input is either object-related data of the account or preprocessed object-related data, and the model's output is the predicted risk type of the account, which is the first label.

[0112] 4. Model Validation: Train the pseudo-labels together with the labeled samples. When the model achieves the expected results, stop training to obtain the object recognition model.

[0113] Figure 5 The diagram illustrates the principle of an optional model training method provided in an embodiment of this application. Figure 5 As shown in the figure, the labeled samples are the sample data with labels, namely the object-related data of risky accounts with labels and the object-related data of normal accounts (whose labels indicate no risk). The unlabeled samples represent the object-related data of risky accounts with the above unknown risk types. The machine learning model is the object recognition model to be trained. As can be seen from the figure, the labeled samples include sample data of multiple risk types (Category 1, Category 2, ... shown in the figure).

[0114] During model training, labeled samples are first used for repeated training until the first training termination condition is met (e.g., one or more preset training metrics meet certain conditions), resulting in a first classification model. Then, this model is used to predict labels for unlabeled samples. Specifically, object-related data of the unlabeled samples can be input into the model to obtain the predicted first label. This label is then used as the pseudo-label for the unlabeled samples, resulting in pseudo-labeled samples. Afterward, the model continues iterative training based on the labeled sample data and these pseudo-labeled sample data until the model's performance reaches the expected level, such as the model's loss function converging, resulting in a well-trained object recognition model.

[0115] Step S2: Construct a reference dataset based on label propagation.

[0116] Similarly, this step can be performed by server 10 or other electronic devices, providing the constructed reference dataset to server 10 for use. In this embodiment, the construction of the reference dataset is also completed by training device 30.

[0117] Semi-supervised learning (i.e., risk identification models) helps address the timeliness issue of user risk detection. However, in the process of user risk labeling, to ensure the accuracy of model training, different types of risky users are labeled separately, which limits the expansion of the risk user system. Furthermore, the behavioral characteristics of cybercriminals constantly change as they operate using illegitimate accounts. Therefore, relying solely on models for user risk identification is not conducive to the long-term operation of the user risk system. Based on this, this step can leverage the information transferability of knowledge graphs to disseminate user risk labels.

[0118] The preceding text described the different roles that risky accounts play in the entire lifecycle of malicious activities in the black market. Based on differences in users' social interactions, payment behaviors, and other characteristics, semi-supervised learning can be used to label different types of users. For labeled users, i.e., users with tags, the risk tags can be disseminated based on relationships between users, such as entity affiliations and fund flows (such as transfers and sending red envelopes).

[0119] like Figure 6 In the diagram shown, each node represents a user. The diagram illustrates three known risk types of users: users with the first target type (such as risky users engaged in account farming), users with the second target type (such as risky users engaged in traffic generation), and users with the third target type (such as risky users engaged in guiding transactions). There are also some users with unknown risk types (unknown users). Users may be related to each other (the relationship can be determined based on users' social behavior data). Risk tags can be transferred between users with related relationships. As shown in the diagram, the risk tags of users with known risk types will be passed on to unknown users with related risk types. Tag transfer also occurs between users with known risk types and those with related relationships.

[0120] Figures 7a to 7c Several examples of risk label propagation are illustrated schematically, in which, Figure 7a As an example of one-way risk label propagation, if a user of target type A (such as a user engaging in account farming) and an unknown user have had a fund transfer (such as a bank transfer transaction), the risk label of the risk user (target type A label, such as an account farming label) will be passed on to the unknown user. Figure 7b As an example of the circular propagation of multiple types of risk labels, if a user of target type A and a user of target type B (such as a user at risk of fund transfer) have transferred funds, their risk labels will be passed on to each other. At the same time, they may also pass on labels to unknown users. In this case, a closed loop of label propagation may occur, i.e., endless propagation. At this point, a loss function is used to break out of the loop. Figure 7cAs an example of multi-source risk label propagation, the risk label of an unknown user may be obtained through more than one path. Risk users of different risk types (users of target type A and target type B shown in the figure) may all be associated with the same unknown user, and the label information of these risk users will also be passed on to the unknown user.

[0121] It is evident that tags associated with users can spread and influence each other. Therefore, these factors need to be considered to more comprehensively and accurately assess a user's risk outcomes.

[0122] Tag propagation can be conducted through multiple iterations, based on the relationships between users. These relationships can be categorized into several types. For example, object relationships can be divided into three types: red envelope relationships, transfer relationships, and entity relationships. Red envelope and transfer relationships are based on the flow of funds. If two users (i.e., accounts) have sent or received red envelopes, they are considered to have a red envelope relationship. If two users (i.e., accounts) have transferred funds (including payment transfers or other transfer methods), they are considered to have a transfer relationship. Entity relationships are defined as two users being associated with the same entity (e.g., both using the same contact method).

[0123] It is understandable that the above description of the relationships is just an example. In actual applications, different partitioning methods can be configured according to different application scenarios.

[0124] The implementation process of the label propagation algorithm is as follows: initialization: (Loss function at initialization) When Loss decreases: Tag propagation: by the first The propagation results of the wheel The relationship R with the user is obtained. The propagation results of the wheel

[0125] Results Summary: Based on the first The propagation results of the wheel Summarized

[0126] Loss calculation: based on the result set calculate

[0127] Output: When The result at the minimum

[0128] in, This represents the annotation labels of each first sample object during the initialization phase. This represents the updated labels of each first sample object obtained after n rounds of label propagation. The user association relationship R is the second association relationship mentioned earlier. The result summary refers to obtaining the fused risk label corresponding to each sample object by fusing the updated labels of all objects that have an association relationship with that object. In the next round of label propagation, the process is based on the fused label corresponding to each sample object and the relationship between the sample objects, until the loss function meets the set conditions, such as reaching the minimum, that is, when the value of the loss function no longer decreases, the iteration is completed, and the fused label of each sample object corresponding to the minimum value of the loss function is taken as the second label of each sample object.

[0129] The following section details the specific implementation of the labeling algorithm, along with an explanation of the meanings of the parameters mentioned above: 1. The loss function used in the label propagation algorithm to determine whether multiple iterations have ended can be expressed as follows:

[0130]

[0131] in, For the first loss function, For the second loss function, and These are the preset weights for the loss function.

[0132] The specific meanings of each parameter in the loss function are as follows: 1) Set I represents the set of all labeled users, which is the number of the first sample objects, and S represents the set of all similar users in set I, which is the set of similar object pairs.

[0133] It is the label for the i-th user / account; This is the predicted label for the i-th user obtained through the label propagation algorithm (i.e., the fused label mentioned above). Assume there are four types of object types, i.e., risk types. and Both can be one-dimensional vectors, and vectors have four elements. The element corresponding to the user's tag has a value of 1, while the other three values ​​are 0. These are four probability values, representing the probability that a user belongs to each risk type after the current sub-tag propagates.

[0134] 2) This represents the importance of the i-th risk label, i.e., the importance of the i-th tagged user. A user's importance can be determined based on their relevant data, and the specific calculation method is not limited. For example, in the process of fund transfer, when a risky user transfers a larger amount of funds, the risk information can be considered more valid, and the user's importance is considered greater.

[0135] This represents the similarity between two users a and b (any pair of similar objects). Optionally, it can be represented by the degree of overlap in fund-related accounts. This is calculated as the intersection of the accounts used for fund transfers (the number of fund transfers between the two users) / the union of the accounts used for fund transfers (the total number of fund transfers between the two users and all other users). Therefore, priority is given to user pairs with a high degree of overlap in their fund transfer accounts. In other words, the higher the overlap in the accounts used by two users, the more likely they are to have the same risk type.

[0136] 3) Let represent the cosine distance between the predicted user vector (i.e., predicted label) of account i and the labeled label of account i in the nth round of label propagation; This represents the cosine distance between the predicted user vectors of accounts a and b in the nth round of label propagation. This represents the predicted user vector for user i. , These represent the predicted user vectors for user a and user b in the nth round (i.e., the user's second label in the next propagation).

[0137] 2. The expression for tag propagation can be represented as:

[0138]

[0139] The meanings of each parameter in this expression are as follows: 1) Set R represents the set of relationships between users, for example... There are three types of association relationships, where r represents one of the association types. 2) This represents the influence factor (i.e., the weight of each type of association) for the association type *r*. Since different association types have varying degrees of influence—for example, fewer users have entity associations, while red envelopes and transfers have significantly different limits on fund amounts—the influence factor is used to adjust the combined weights of different association types. The value of the influence factor for each association type can be set according to needs or experience. For instance, the factor for entity associations is larger, and the factor for transfer associations can be larger than the factor for red envelope associations.

[0140] 3) The influence matrix represents the influence of an object on each type of association (r). A user who transfers money to more than 30 accounts clearly has a significantly different influence than a user who transfers to only 2 accounts. The influence matrix characterizes a user's influence weight; for example, by standardizing the number of accounts a user is associated with, the user's influence weight can be obtained.

[0141] Suppose there are N users in set I. It can be represented as a vector with N element values. For example, the vector has N rows and 1 column. The element value of each row represents the influence of a user in the corresponding type of association, that is, the influence of the user in the corresponding type of social behavior.

[0142] 4) This indicates the path through which the label is propagated.

[0143] Assuming the user relationship network, i.e., set I, has N nodes, then the matrix... have If account i transferred funds to 10 accounts, then... In the table, the values ​​of the 10 transfer account columns corresponding to account i are all 0.1, and the values ​​of the other columns are all 0. An element with a value of 0 indicates that the account is not associated with account i, while an element with a non-zero value indicates that the account is associated with account i. The value of the element represents the degree of association, which is the numerical value used to represent the association when calculating.

[0144] If the association type r is an entity association, assuming account i has entity associations with 5 accounts, the corresponding value is 0.2, and the others are all 0.

[0145] 5) This represents the result of the nth round of label propagation. The result of the (n+1)th round is obtained through the propagation of the result of the nth round and the addition of user labels, which is the addition of new data.

[0146] For example, if the number of users in set I is N during a label propagation, and the number of newly added sample objects is M after obtaining the propagation result, then the number of users in set I during the next round of label propagation will be N+M.

[0147] 6) The weight represents the risk type (i.e., the proportion of sample objects of each risk type in set I). Since the number of accounts varies under different risk types, standardization is required using weights. y represents the labeled user matrix, which is the label of each sample object in set I.

[0148] In other words, a normalized weight can be calculated for different risk types based on the number of labeled users for each risk type. For example, if there are 4 risk types, and the number of labeled users for each risk type is a1, a2, a3, and a4, then the weight of the i-th risk type can be expressed as: a i / (a1+a2+a3+a4).

[0149] Y is the label matrix of all users in set I. Assuming there are N labeled users and 4 risk types in the first round of label propagation, this matrix can be an N x 4 matrix, where each row represents a user's label. In each row, one element has a value of 1, and the other three have values ​​of 0. The risk type corresponding to the element with a value of 1 is the true object type of that sample. Assuming there are N+m labeled users in the second round of label propagation, Y can be an N+m x 4 matrix.

[0150] Based on the above tag propagation formula, the tags of each user in set I can be continuously updated through multiple iterations.

[0151] Assuming the label propagation goes through n rounds, the propagation result is obtained. For an account x in set I, after n rounds of label propagation to all its associated accounts A, its corresponding result vector (predicted label) can be represented as follows:

[0152] in, This represents a standardization function, such as the softmax function. This represents a user associated with user x. As can be seen from this expression, the second risk label for user x can be obtained by fusing and normalizing the updated labels of all users associated with user x. All associated accounts of a user are considered associated users. The user whose row contains non-zero values.

[0153] Specifically, in the iterative process, each iteration yields a corresponding result. Assume there are N tagged users and 4 risk types, vector It can be an N x 4 (or 4 x N) matrix, where the four values ​​in the i-th row (which can be simply called the user vector) represent the probabilities of the i-th user belonging to the four risk types. This yields... Then, for the i-th user, the user vectors of all associated users are summed and then standardized to obtain the prediction vector for the i-th user, which is the vector used to calculate the loss function for this iteration. If user i has 3 associated users, then the user vectors of these 3 users are superimposed and then standardized.

[0154] Through continuous iterative updates, the user vectors obtained when the loss no longer decreases are used as the final risk labels (i.e., second labels) for these labeled users. These second labels are then used to predict the identification results of the target object in the dataset. Assuming the final iteration has 5,000 labeled users, the final result will be user vectors for 5,000 users. The annotations, secondary labels, and object-related data of these 5,000 users can then be used as a reference dataset.

[0155] Step S3: Server 10 obtains object-related data of the user to be identified, i.e., user-related data.

[0156] Step S4: Server 10 calls the object recognition model to predict the first label of the user to be identified.

[0157] Specifically, the object-related data of each user to be identified is input into the object user identification model. The model predicts the initial risk label, or first label, of each user to be identified, which is the risk type of the user that the model initially determines.

[0158] Step S5: Server 10 determines the second label of the user to be identified based on the reference dataset.

[0159] Server 10 predicts the final risk label, or second label, for each user based on the reference dataset and object-related data of the users to be identified, and determines the identification result of the users based on the final risk labels. This step may include: a. Determine multiple types of associations between each user to be identified and other users (including other users to be identified and sample objects), including but not limited to the aforementioned entity associations, red envelope associations, transfer associations, etc.

[0160] b. Based on the first risk label of each user to be identified obtained according to the following label propagation formula and step S32, a second label of each user to be identified is obtained through at least one label propagation:

[0161]

[0162] As an example, suppose the number of users to be identified is M, and the number of sample users is N. During the identification phase, the number of nodes (i.e., the number of users) in the user relationship network is M+N.

[0163] At this point, regarding the parameters mentioned above in the label propagation formula, This represents the influence factor for the association type r. The influence factor corresponding to each type of association can be preset according to actual needs or experimental values, and can be combined with the influence factors from the previous iteration stages. same.

[0164] For influence matrix For each of the M+N users, the influence factor (i.e., influence or influence weight) corresponding to each type of association between that user and other users can be determined based on that user's association with other users. Similarly, the propagation path of each user in tag propagation can be determined based on the association between each user and other users. .

[0165] For example, taking relation type r as an example, for M+N users, an influence matrix can be obtained. The matrix contains M+N values, representing the influence weight of each of the M+N users. for A 3D matrix.

[0166] This represents the weight of the risk type, and its value is the same as in the iteration phase. In the application phase, Y represents the initial risk labels for N+M users. For users to be identified, the initial risk label is the first label predicted by the object recognition model; for sample users, it is the labeled label of that sample user.

[0167] During this application phase, in the first round of tag propagation... The second label of each sample user, that is, the second label of N sample users (i.e., the label of the last iteration). ) and the first labels of M users to be identified.

[0168] Based on the above label propagation formula, calculate the value at this time. , It is The matrix, where k represents the number of risk types, such as 4 types. If only one label propagation is performed, according to... ,pass The final result vector for each user to be identified can be calculated. This is the second label for each user, and the vector includes k probability values. The risk type corresponding to the highest probability value or the probability value exceeding a threshold can be determined as the risk type of the user to be identified. If multiple label propagation operations are performed, the result vectors of each user (including users to be identified and sample users) obtained from the first label propagation are used as the basis for the second propagation. The initial value is used to update the label again based on the label propagation formula. This operation is repeated until the propagation number reaches the set number (i.e., the maximum number of propagations set in advance). The result vector of the user to be identified obtained in the last propagation is used as the second label of the user to be identified.

[0169] Understandably, in practice, to avoid infinite loops, when calculating the result vector for each tag propagation, the result vector for each user should be calculated one by one, and the order of calculation is not limited. However, once the result vector for a user has been calculated, it will not be calculated again because the result vectors of other users related to it change again.

[0170] Furthermore, in practical applications, various types of new social behavior data of risky users can be continuously collected, which means the training dataset 12 can be continuously updated and expanded. The risk identification model can be retrained periodically or when the amount of updated data reaches a certain level to further improve model performance. Similarly, the data in the sample object library 11 can also be updated to expand the amount of sample user data.

[0171] The method provided in this application, for the first time, deconstructs the illegal industry based on its lifecycle, identifies and labels different types of risky accounts using models, and innovatively employs a tag propagation algorithm based on user relationships to propagate user risk tags, thus improving the risk user system. Based on this method, not only are object profiles of different risk types depicted, but the long-term maintenance of risk user tags is also ensured, allowing for better application in the strategic crackdown on risky users and providing a new approach to early identification of risky users. Compared with existing methods, the solution provided in this application has at least the following advantages: 1) It can improve the timeliness of discovering risky users.

[0172] For each possible stage of fraudulent activities by illicit industries, risk identification of users can be achieved at any stage through similarity analysis, or correlation analysis, of risky users, without relying solely on delayed information such as customer complaints. This allows for the pre-identification and strategic crackdown on fraudulent transactions by leveraging users with different risk types in various scenarios. This approach is better suited to different fraud scenarios and combat methods, improving the timeliness and accuracy of fraud identification strategies.

[0173] 2) Increased the coverage of user risk labels.

[0174] A tag propagation algorithm based on user relationship graphs leverages the information connections between users to spread risk tags, expanding the coverage of risky users. Through the construction and propagation of user risk tags, a user risk system can characterize the risk attributes of all users who have made transactions (such as mobile payments), and it also has numerous applications in the early identification of fraud risks. For example: 1. For high-risk users who pose a risk of attracting unwanted traffic, their social network can be identified, alerting other users to potential risks associated with transacting with them. For example, when a user engages in a large transaction with a newly added friend, real-time strategies can be implemented to prevent the user from falling into a fraud trap.

[0175] 2. For users at risk of account nurturing, merchants with fraudulent potential can be identified in advance by analyzing the user's previous payment behavior. Merchants with frequent transactions can be identified in advance, and merchants who may engage in fraudulent transactions later can be identified during the account nurturing stage and penalized accordingly.

[0176] 3. For users at risk of fund transfer or money laundering, the system can monitor their fund flows and promptly prevent illegal fund transfers. For example, when such users make large-scale fund withdrawals, real-time monitoring can be implemented to prevent fund transfers.

[0177] 4. In the process of establishing a user risk system, users without any attributes may be discovered, including many alt accounts and zombie accounts. These may be tools used by illegal industries for later malicious activities, and can provide new source data for identifying fraud risks.

[0178] For example, if the label propagation algorithm predicts that the probability / weight of each risk attribute of an account is 0, then the result vector for that account is... If all attribute dimensions are 0, the account can be considered a secondary / zombie account. The social information and payment behavior information of such accounts can be used as new samples for the risk identification model. Through training, the model can not only predict various types of risky accounts, but also identify secondary / zombie accounts.

[0179] Based on the same principle as the method provided in the embodiments of this application, the embodiments of this application also provide an object recognition device, such as... Figure 8 As shown, the object recognition device 100 may include a first prediction module 110, a reference dataset acquisition module 120, a second prediction module 130, and a recognition result determination module 140.

[0180] The first prediction module 110 is used to acquire object-related data of at least one object to be identified; for each object to be identified, the first label of the object is predicted by the object recognition model based on the object-related data of the object, and the first label of an object represents the object type to which the object belongs among multiple object types. The reference dataset acquisition module 120 is used to acquire a reference dataset, which includes object-related data and second labels of multiple first sample objects with labeled labels. The labeled label of a first sample object represents the real object type to which the object belongs among multiple object types, and the second label of an object represents the probability of the object belonging to each object type among multiple object types. The second prediction module 430 is used to determine a first association relationship between at least one object to be identified and each object in multiple first sample objects based on object-related data of each object to be identified and each first sample object, and to determine a second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship. The recognition result determination module 140 is used to determine the recognition result of each object to be recognized based on the second label of each object to be recognized.

[0181] Optionally, the second prediction module can be specifically used for: The first label of each object to be identified is used as the annotation label and the initial second label of the object to be identified. Based on the annotation label and the second label of each object to be identified and the first sample object, and based on the first association relationship, at least one label propagation is performed between the object to be identified and the first sample object to obtain the updated label of each object to be identified and the first sample object. For each object to be identified, the updated labels of each object with the first association relationship are merged according to the first association relationship to obtain the second label of the object.

[0182] Optionally, the second prediction module may perform the following operations during each label propagation: For each object in the object to be identified and the first sample object, the second label of the object is updated based on the second label of each object that has a relationship with the object, according to the first association relationship; for each object, the updated label of the object is obtained by fusing the updated second label of the object and the labeled label of the object, and the updated label of the object is used as the second label of the object in the next label propagation.

[0183] Optionally, the object-related data includes at least one specified type of object-related data, and the first association includes an association of one type corresponding to each specified type of object-related data; correspondingly, the second prediction module can be used for: Obtain the weight corresponding to each type of association; determine the second label of each object to be identified based on the first label of each object to be identified, the labeled label and second label of each first sample object, each type of association, and the weight corresponding to each type of association. Optionally, the second prediction module can be used to: for each object among at least one object to be identified and multiple first sample objects, determine the influence of the object based on the object-related data of the object; determine the second label of each object to be identified based on the first label of each object to be identified, the labeled label and second label of each first sample object, the influence of each object to be identified and the first sample objects, and the first association.

[0184] Optionally, the object-related data includes at least one specified type of object-related data, the first association includes an association of a type corresponding to each specified type of object-related data, and the influence of each of the at least one object to be identified and the plurality of first sample objects, including the influence of each object corresponding to each type of association.

[0185] Optionally, the second prediction module can be used to: determine the proportion of each object type in at least one object to be identified and multiple first sample objects based on the first label of each object to be identified and the annotation label of each first sample object; use the proportion of each object type as a weight to weight the first label of the corresponding object type in at least one object to be identified and to weight the annotation labels of the corresponding object type in multiple first sample objects; and determine the second label of each object to be identified based on the weighted first label of each object to be identified, the weighted annotation label and second label of each first sample object, and the first association relationship.

[0186] Optionally, the object recognition model is obtained by the model training module by performing the following operations: Obtain a first training dataset, which includes object-related data of multiple second sample objects with labeled labels and object-related data of multiple unlabeled third sample objects. The multiple second sample objects include multiple objects of each of multiple object types. Based on object-related data of multiple second sample objects, the initial classification model is trained until the first training termination condition is met, resulting in the first classification model. For each third sample object, based on the object-related data of that object, the object type of the object is predicted by the first classification model, and the label of the object is determined according to the object type. Based on the object-related data of multiple second sample objects and the object-related data of multiple third sample objects with labels, the first classification model is trained again until the second training termination condition is met, resulting in the object recognition model.

[0187] Optionally, the reference dataset is obtained by the reference dataset acquisition module in the following ways: Obtain a second training dataset, which includes object-related data of multiple first sample objects with labeled tags; determine the second association relationship between objects in the second training dataset based on the object-related data of each first sample object; use the labeled tag of each first sample object as the initial third tag of the object, and repeat the following operations until the updated third tags of multiple first sample objects meet the preset conditions, and determine the third tag of each first sample object that meets the preset conditions as the second tag of the object; based on the second association relationship and the labeled tags and third tags of each first sample object, obtain the updated fourth tag of each first sample object by performing tag propagation among multiple first sample objects; for each first sample object, obtain the new third tag of the object by fusing the fourth tags of each first sample object that has an association relationship with the object, according to the second association relationship.

[0188] Optionally, the reference dataset acquisition module can also be used for: After each label propagation, new data is acquired, including object-related data of at least one sample object with labeled tags; each sample object in the new data is taken as the new first sample object, and the second training dataset is updated based on the new data; based on the object-related data of each first sample object in the updated second training dataset, the second association relationship between each object in the updated second training dataset is determined, and the updated second association relationship is obtained. When the reference dataset acquisition module obtains the updated fourth label for each first sample object, it can be used for: The annotation label of each newly added first sample object is used as the third label of that object. Based on the updated second association relationship, as well as the updated annotation labels and third labels of each first sample object, the updated fourth label of each first sample object is obtained by propagating the labels among the updated multiple first sample objects.

[0189] Optionally, the labels for each sample object in the newly added data are obtained in the following way: Obtain object-related data for at least one unlabeled object, wherein at least one sample object includes at least one unlabeled object; for each of the at least one unlabeled objects, based on the object-related data of that object, predict the first label of that object through an object recognition model, and use the first label of that object as the label of that object.

[0190] Optionally, for each label propagation, the reference dataset acquisition module is also used for: Based on object-related data of multiple first sample objects, similar object pairs are determined among the multiple first sample objects; wherein, satisfying preset conditions includes the value setting conditions of the loss function; The loss function includes a first loss function and a second loss function. For each label propagation, the value of the first loss function characterizes the difference between the labeled label of each first sample object and the new third label, while the value of the second loss function characterizes the difference between the new third labels of each similar object pair.

[0191] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0192] Based on the same principles as the methods and apparatus provided in the embodiments of this application, the embodiments of this application also provide an electronic device, which may include a memory and a processor. The memory stores a computer program, and the processor, when running the computer program, is used to execute the methods provided in any optional embodiment of this application, or to execute the actions performed by the apparatus provided in any optional embodiment of this application.

[0193] As an optional embodiment, Figure 9 The diagram shows a structural schematic of an electronic device according to an embodiment of this application, such as... Figure 9As shown, the electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0194] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0195] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0196] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0197] The memory 4003 stores computer programs that execute embodiments of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0198] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.

[0199] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0200] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any of the optional embodiments of this application described above.

[0201] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0202] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.

Claims

1. An object recognition method, characterized in that, The method is performed by a computer device and includes: Obtain object-related data for at least one object to be identified, wherein the object-related data for each object to be identified includes the social behavior data of the object to be identified, and the object to be identified is a user or a merchant; For each object to be identified, the first label of the object is predicted by the object recognition model based on the object-related data of the object. The first label of an object represents the object type to which the object belongs among multiple object types. The multiple object types include multiple target types and one non-target type. Each target type corresponds to a risk type, and the risk type is a fraud behavior type. Obtain a reference dataset, which includes object-related data and second labels for multiple first sample objects with labeled tags. The labeled tag of a first sample object represents the true object type to which the object belongs among the multiple object types, and the second label of an object represents the probability that the object belongs to each of the multiple object types. Based on the object-related data of each of the objects to be identified and each of the first sample objects, a first association relationship is determined between the at least one object to be identified and each of the objects in the plurality of first sample objects; The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship; For each object to be identified, the identification result of the object type to which the object belongs is determined based on the second label of the object to be identified.

2. The method according to claim 1, characterized in that, The step of determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship includes: The first label of each object to be identified is used as the annotation label and the initial second label of the object to be identified. Based on the annotation label and the second label of each object to be identified and the first sample object, and based on the first association relationship, at least one label propagation is performed between the object to be identified and the first sample object to obtain the updated label of each object to be identified and the first sample object. For each object to be identified, the updated tags of each object that has the first association relationship with the object are merged to obtain the second tag of the object.

3. The method according to claim 2, characterized in that, Each tag propagation involves the following operations: For each object in the object to be identified and the first sample object, the second label of the object is updated based on the second label of each object that has an association relationship with the object, according to the first association relationship; For each object, the updated label of the object is obtained by fusing the updated second label of the object and the annotation label of the object, and the updated label of the object is used as the second label of the object in the next label propagation.

4. The method according to any one of claims 1 to 3, characterized in that, The object-related data includes at least one specified type of object-related data, the first association relationship includes an association relationship of one type corresponding to each specified type of object-related data, the object-related data includes multiple types of data, and the specified type is one or more types preset from the multiple types; The step of determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship includes: Obtain the weight corresponding to each type of association; The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, each type of association, and the weight corresponding to each type of association.

5. The method according to any one of claims 1 to 3, characterized in that, Also includes: For each of the at least one object to be identified and the plurality of first sample objects, the influence of the object is determined based on the object-related data of the object, and the influence of each object is used to characterize the object's ability to influence other objects; The step of determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship includes: The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, the influence of each object to be identified and the first sample object, and the first association relationship.

6. The method according to claim 5, characterized in that, The object-related data includes at least one specified type of object-related data, which includes multiple types of data. The specified type is one or more types preset from the multiple types. The first association relationship includes an association relationship of a type corresponding to each specified type of object-related data. The influence of each of the at least one object to be identified and each of the multiple first sample objects includes the influence of each object corresponding to each type of association relationship. The influence of each object corresponding to each type of association relationship is used to characterize: the ability of the object to influence objects that are associated with the object of the type.

7. The method according to any one of claims 1 to 3, characterized in that, Also includes: Based on the first label of each object to be identified and the annotation label of each first sample object, determine the proportion of the number of each object type in the at least one object to be identified and the plurality of first sample objects; The step of determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship includes: The proportion of objects of each object type is used as a weight to weight the first label of the corresponding object type in the at least one object to be identified, and to weight the annotation labels of the corresponding object types in the multiple first sample objects; The second label of each object to be identified is determined based on the weighted first label of each object to be identified, the weighted annotation label and second label of each first sample object, and the first association relationship.

8. The method according to claim 1, characterized in that, The object recognition model was trained in the following way: Obtain a first training dataset, which includes object-related data of multiple second sample objects with labeled labels and object-related data of multiple unlabeled third sample objects. The multiple second sample objects include multiple objects whose real object type is each of the multiple object types. Based on the object-related data of the multiple second sample objects, the initial classification model is trained until the first training termination condition is met, and the first classification model is obtained. For each of the third sample objects, based on the object-related data of the object, the object type of the object is predicted by the first classification model, and the label of the object is determined according to the object type. Based on the object-related data of the multiple second sample objects and the object-related data of the multiple third sample objects with labeled tags, the first classification model is trained again until the second training termination condition is met, and the object recognition model is obtained.

9. The method according to claim 1, characterized in that, The reference dataset was obtained in the following way: Obtain a second training dataset, which includes object-related data of multiple first sample objects with labeled tags; Based on the object-related data of each of the first sample objects, determine the second association relationship between the objects in the second training dataset; The annotation label of each first sample object is used as the initial third label of that object. The following operations are repeated until the updated third labels of the multiple first sample objects meet a preset condition. The third label of each first sample object that meets the preset condition is determined as the second label of that object: Based on the second association relationship and the annotation labels and third labels of each of the first sample objects, the updated fourth label of each first sample object is obtained by propagating the labels among the multiple first sample objects; for each first sample object, a new third label of the object is obtained by fusing the fourth labels of each first sample object that has an association relationship with the object, according to the second association relationship.

10. The method according to claim 9, characterized in that, After each tag propagation, the method further includes: Acquire new data, which includes object-related data of at least one sample object with labeled tags; Each sample object in the newly added data is used as the first newly added sample object, and the second training dataset is updated based on the newly added data. Based on the object-related data of each first sample object in the updated second training dataset, the second association relationship between each object in the updated second training dataset is determined, and the updated second association relationship is obtained. The step of obtaining an updated fourth label for each first sample object based on the second association relationship and the annotation labels and third labels of each first sample object, through label propagation among the multiple first sample objects, includes: The annotation label of each newly added first sample object is used as the third label of that object. Based on the updated second association relationship, as well as the updated annotation labels and third labels of each first sample object, the updated fourth label of each first sample object is obtained by propagating the labels among the updated multiple first sample objects.

11. The method according to claim 10, characterized in that, The labels for each sample object in the newly added data were obtained in the following way: Obtain object-related data for at least one unlabeled object, wherein the at least one sample object includes the at least one unlabeled object; For each of the at least one unlabeled objects, based on the object-related data of that object, the first label of that object is predicted by the object recognition model, and the first label of that object is used as the label of that object.

12. The method according to claim 9, characterized in that, The method further includes: Based on the object-related data of the plurality of first sample objects, similar object pairs are determined among the plurality of first sample objects; Among them, the conditions for satisfying the preset conditions include the value setting conditions of the loss function; The loss function includes a first loss function and a second loss function. For each label propagation, the value of the first loss function characterizes the difference between the labeled label of each first sample object and the new third label, and the value of the second loss function characterizes the difference between the new third labels of each pair of similar objects.

13. An object recognition device, characterized in that, The device is applied to computer equipment and includes: The first prediction module is used to acquire object-related data of at least one object to be identified, wherein the object-related data of each object to be identified includes the social behavior data of the object to be identified; for each object to be identified, a first label of the object is predicted by an object recognition model based on the object-related data of the object, wherein the first label of an object represents the object type to which the object belongs among multiple object types, and the object type is a risk type; wherein the object is a user or a merchant; the multiple object types include multiple target types and one non-target type, each target type corresponds to a risk type, and the risk type is a fraudulent behavior type; The reference dataset acquisition module is used to acquire a reference dataset, which includes object-related data and second labels of multiple first sample objects with labeled tags. The labeled tag of a first sample object represents the real object type to which the object belongs among the multiple object types, and the second label of an object represents the probability that the object belongs to each object type among the multiple object types. The second prediction module is used to determine a first association relationship between the at least one object to be identified and each object in the plurality of first sample objects based on object-related data of each object to be identified and each first sample object, and to determine a second label of each object to be identified based on a first label of each object to be identified, a label and a second label of each first sample object, and the first association relationship; The identification result determination module is used to determine the identification result of the object type to which each object belongs based on the second label of each object to be identified.

14. The apparatus according to claim 13, characterized in that, The second prediction module, when determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, is specifically used for: The first label of each object to be identified is used as the annotation label and the initial second label of the object to be identified. Based on the annotation label and the second label of each object to be identified and the first sample object, and based on the first association relationship, at least one label propagation is performed between the object to be identified and the first sample object to obtain the updated label of each object to be identified and the first sample object. For each object to be identified, the updated tags of each object that has the first association relationship with the object are merged to obtain the second tag of the object.

15. The apparatus according to claim 14, characterized in that, The second prediction module performs the following operations during each label propagation: For each object in the object to be identified and the first sample object, the second label of the object is updated based on the second label of each object that has an association relationship with the object, according to the first association relationship; For each object, the updated label of the object is obtained by fusing the updated second label of the object and the annotation label of the object, and the updated label of the object is used as the second label of the object in the next label propagation.

16. The apparatus according to any one of claims 13-15, characterized in that, The object-related data includes at least one specified type of object-related data, the first association relationship includes an association relationship of one type corresponding to each specified type of object-related data, the object-related data includes multiple types of data, and the specified type is one or more types preset from the multiple types; The second prediction module, when determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, is specifically used for: Obtain the weight corresponding to each type of association; The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, each type of association, and the weight corresponding to each type of association.

17. The apparatus according to any one of claims 13-15, characterized in that, The second prediction module is also used for: For each of the at least one object to be identified and the plurality of first sample objects, the influence of the object is determined based on the object-related data of the object, and the influence of each object is used to characterize the object's ability to influence other objects; The second prediction module, when determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, is specifically used for: The second label of each object to be identified is determined based on the first label of each object to be identified, the annotation label and second label of each first sample object, the influence of each object to be identified and the first sample object, and the first association relationship.

18. The apparatus according to claim 17, characterized in that, The object-related data includes at least one specified type of object-related data, which includes multiple types of data. The specified type is one or more types preset from the multiple types. The first association relationship includes an association relationship of a type corresponding to each specified type of object-related data. The influence of each of the at least one object to be identified and each of the multiple first sample objects includes the influence of each object corresponding to each type of association relationship. The influence of each object corresponding to each type of association relationship is used to characterize: the ability of the object to influence objects that are associated with the object of the type.

19. The apparatus according to any one of claims 13-15, characterized in that, The second prediction module is also used for: Based on the first label of each object to be identified and the annotation label of each first sample object, determine the proportion of the number of each object type in the at least one object to be identified and the plurality of first sample objects; The second prediction module, when determining the second label of each object to be identified based on the first label of each object to be identified, the annotation label and second label of each first sample object, and the first association relationship, is specifically used for: The proportion of objects of each object type is used as a weight to weight the first label of the corresponding object type in the at least one object to be identified, and to weight the annotation labels of the corresponding object types in the multiple first sample objects; The second label of each object to be identified is determined based on the weighted first label of each object to be identified, the weighted annotation label and second label of each first sample object, and the first association relationship.

20. The apparatus according to claim 13, characterized in that, The object recognition model is trained by the model training module in the following way: Obtain a first training dataset, which includes object-related data of multiple second sample objects with labeled labels and object-related data of multiple unlabeled third sample objects. The multiple second sample objects include multiple objects whose real object type is each of the multiple object types. Based on the object-related data of the multiple second sample objects, the initial classification model is trained until the first training termination condition is met, and the first classification model is obtained. For each of the third sample objects, based on the object-related data of the object, the object type of the object is predicted by the first classification model, and the label of the object is determined according to the object type. Based on the object-related data of the multiple second sample objects and the object-related data of the multiple third sample objects with labeled tags, the first classification model is trained again until the second training termination condition is met, and the object recognition model is obtained.

21. The apparatus according to claim 13, characterized in that, The reference dataset is obtained by the reference dataset acquisition module in the following ways: Obtain a second training dataset, which includes object-related data of multiple first sample objects with labeled tags; Based on the object-related data of each of the first sample objects, determine the second association relationship between the objects in the second training dataset; The annotation label of each first sample object is used as the initial third label of that object. The following operations are repeated until the updated third labels of the multiple first sample objects meet a preset condition. The third label of each first sample object that meets the preset condition is determined as the second label of that object: Based on the second association relationship and the annotation labels and third labels of each of the first sample objects, the updated fourth label of each first sample object is obtained by propagating the labels among the multiple first sample objects; for each first sample object, a new third label of the object is obtained by fusing the fourth labels of each first sample object that has an association relationship with the object, according to the second association relationship.

22. The apparatus according to claim 21, characterized in that, After each label propagation, the reference dataset acquisition module is further used for: Acquire new data, which includes object-related data of at least one sample object with labeled tags; Each sample object in the newly added data is used as the first newly added sample object, and the second training dataset is updated based on the newly added data. Based on the object-related data of each first sample object in the updated second training dataset, the second association relationship between each object in the updated second training dataset is determined, and the updated second association relationship is obtained. When the reference dataset acquisition module obtains an updated fourth label for each first sample object based on the second association relationship and the labeled and third labels of each first sample object, by performing label propagation among the multiple first sample objects, it is specifically used for: The annotation label of each newly added first sample object is used as the third label of that object. Based on the updated second association relationship, as well as the updated annotation labels and third labels of each first sample object, the updated fourth label of each first sample object is obtained by propagating the labels among the updated multiple first sample objects.

23. The apparatus according to claim 22, characterized in that, The labels for each sample object in the newly added data were obtained in the following way: Obtain object-related data for at least one unlabeled object, wherein the at least one sample object includes the at least one unlabeled object; For each of the at least one unlabeled objects, based on the object-related data of that object, the first label of that object is predicted by the object recognition model, and the first label of that object is used as the label of that object.

24. The apparatus according to claim 21, characterized in that, The reference dataset acquisition module is also used for: Based on the object-related data of the plurality of first sample objects, similar object pairs are determined among the plurality of first sample objects; Among them, the conditions for satisfying the preset conditions include the value setting conditions of the loss function; The loss function includes a first loss function and a second loss function. For each label propagation, the value of the first loss function characterizes the difference between the labeled label of each first sample object and the new third label, and the value of the second loss function characterizes the difference between the new third labels of each pair of similar objects.

25. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-12.

26. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-12.

27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Method and device for predicting object category

    CN108629358A

  • Object recognition method and device, computer system and readable storage medium

    CN113094595A