A method, apparatus, electronic device, and storage medium for identifying an object state

By analyzing the naming pattern and similarity of merchant names, dividing object sets, and using resource status information of object sets for state recognition, the problem of low merchant identification accuracy in the prior art is solved, and fast and accurate merchant status recognition and crackdown effects are achieved.

CN114723509BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110014363.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-06
Publication Date
2025-07-08
Estimated Expiration
2041-01-06

AI Technical Summary

Technical Problem

In the prior art, the accuracy of identifying bad merchants based on risk control strategy platforms is low, resulting in poor crackdown effects and it is difficult to effectively identify and manage the problems of uneven merchant quality.

Method used

By analyzing the naming mode of merchant names, using part-of-speech annotation and similarity calculation, dividing object sets, mining similarity between objects, and quickly and accurately identifying state based on resource state information of object sets.

Benefits of technology

It realizes rapid and accurate identification of the target object status, improves the efficiency and accuracy of merchant status identification, and can promptly crack down on bad merchants and protects user funds safety and ecological health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723509B_ABST
    Figure CN114723509B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and in particular, to a method, device, electronic device, and storage medium for identifying the state of an object, so as to improve the efficiency and accuracy of object state identification. Among them, the method includes: determining a naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object; determining an object set to which the target object belongs according to the naming pattern, where the object set is obtained by dividing according to the similarity between each object; determining the target resource state of the target object according to the resource state information of the object set, where the resource state information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set. The object set in this application is obtained by dividing according to similarity. Therefore, when identifying the target resource state based on the resource state information of the object set to which the target object belongs, fast and accurate state identification can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of the mobile payment field, people's shopping patterns have become increasingly diverse. There have emerged shopping payment methods such as scanning codes or face recognition using payment applications. Merchants need to be registered in the payment application first before customers can use the payment application for payment transactions. However, usually the quality of merchants varies. The number of normal merchants generally accounts for the vast majority, but there will also be some bad merchants hidden among them. These bad merchants with violations that disrupt the ecological health, harm users' fund security and experience should be cracked down on.

[0003] In the related art, a risk control strategy platform can be used to identify and crack down on bad merchants. Specifically, data rules at the merchant dimension are mainly used for strategy strikes, and strategies are hit for interception and strikes. However, the recognition accuracy of this method is relatively low, and the cracking effect is not good. Similarly, for other objects, it is the same principle. Therefore, how to improve the recognition efficiency and accuracy of object status is an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of the present application provide a method, device, electronic device and storage medium for recognizing the status of an object, so as to improve the recognition efficiency and accuracy of the object status.

[0005] A method for recognizing the status of an object provided by an embodiment of the present application includes:

[0006] Determine the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object;

[0007] Determine the object set to which the target object belongs according to the naming pattern, where the object set is obtained by dividing according to the similarity between each object;

[0008] Determine the target resource status of the target object according to the resource status information of the object set, where the resource status information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

[0009] A device for recognizing the status of an object provided by an embodiment of the present application includes:

[0010] A name processing unit, configured to determine the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object;

[0011] A set determination unit, configured to determine the object set to which the target object belongs according to the naming pattern, where the object set is obtained by dividing according to the similarity between each object;

[0012] A status recognition unit, configured to determine the target resource status of the target object according to the resource status information of the object set, where the resource status information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

[0013] Optionally, if there are multiple target objects, for any one of the target objects, the set determination unit is specifically configured to:

[0014] If there is no candidate object set with the same naming pattern as the any one target object, obtain the similarity between the any one target object and each of the other target objects;

[0015] Filter out at least one other target object whose similarity with the any one target object reaches a preset similarity threshold;

[0016] Form a new object set by the any one target object and the at least one other target object filtered out, and use the new object set as the object set to which the target object belongs.

[0017] Optionally, the naming patterns of the candidate objects in the same candidate object set are the same, and the similarity between any two candidate objects reaches the similarity threshold.

[0018] Optionally, the apparatus further includes:

[0019] A set partitioning unit, configured to obtain the similarity between any two objects before the set determination unit determines the object set to which the target object belongs according to the naming pattern; for each object with the same naming pattern, query each connected component in the connected graph constructed based on the similarity between the objects; respectively form an object set by the objects corresponding to the vertices in each connected component, where the connected graph is constructed by taking each object as a vertex and connecting the vertices corresponding to any two objects whose similarity reaches the similarity threshold; or,

[0020] Obtain the similarity between any two objects; for each object with the same naming pattern, perform clustering on the objects based on the similarity between the objects to obtain at least one object set.

[0021] Optionally, the set partitioning unit is specifically configured to:

[0022] Obtain the text similarity between the object names of the any two objects; and

[0023] Respectively obtain the information similarity between the registration information corresponding to the any two objects;

[0024] Weight the text similarity with the information similarity corresponding to the registration information to obtain the similarity between any two objects.

[0025] Optionally, the resource status information of the object set is determined according to the historical resource processing methods associated with other objects in the object set except the target object; the apparatus further includes:

[0026] A status update unit, configured to, after the status recognition unit determines the target resource status of the target object according to the resource status information of the object set, obtain the resource processing method associated with the target object when performing resource processing based on the target object;

[0027] Update the resource status information of the object set to which the target object belongs according to the resource processing method of the target object.

[0028] Optionally, the name processing unit is specifically configured to:

[0029] Input the object name of the target object into a first status recognition model, perform word segmentation processing on the object name based on the first status recognition model, and perform part-of-speech tagging on each word obtained through word segmentation processing;

[0030] Determine the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word;

[0031] The status recognition unit is specifically configured to:

[0032] Obtain the target resource status of the target object determined according to the resource status information of the object set output by the first status recognition model;

[0033] Wherein, the first status recognition model is trained based on a training sample data set, and the training sample data set at least includes a plurality of training samples, and each training sample includes the object name of a candidate object.

[0034] Optionally, the apparatus further includes:

[0035] A model training unit, configured to train the first status recognition model in the following manner:

[0036] Select at least two training samples from the training sample data set;

[0037] Input the object names of the candidate objects in each training sample into a second status recognition model respectively, and perform word segmentation processing on the object names in each training sample based on the second status recognition model;

[0038] Perform part-of-speech tagging on each of the segmented words, and determine the naming pattern corresponding to each candidate object according to the part-of-speech tagging results of the segmented words;

[0039] According to the similarity between each candidate object, divide the candidate objects under the same naming pattern to obtain at least one candidate object set;

[0040] Based on the historical resource processing methods of the candidate objects in each candidate object set respectively, determine the resource status information corresponding to each candidate object set;

[0041] Output the resource status information of each candidate object set via the second status recognition model;

[0042] According to the resource status information output by the second status recognition model, perform at least one parameter adjustment on the second status recognition model to obtain the first status recognition model.

[0043] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores program codes. When the program codes are executed by the processor, the processor executes the steps of any one of the above object status recognition methods.

[0044] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any one of the above object status recognition methods.

[0045] An embodiment of the present application provides a computer-readable storage medium, which includes program codes. When the program product runs on an electronic device, the program codes are used to cause the electronic device to execute the steps of any one of the above object status recognition methods.

[0046] The beneficial effects of the present application are as follows:

[0047] The embodiments of the present application provide a method, an apparatus, an electronic device, and a storage medium for identifying the state of an object. Since the embodiments of the present application design a naming pattern, mine and analyze the naming characteristics of the object name, determine the similarity between objects based on this characteristic, and then divide the object set, so as to realize the mining of object sets with high object name similarity. Therefore, when identifying the state of a target object, only the naming pattern of the target object needs to be used to analyze the object set to which the target object belongs, and then the state of the target object can be analyzed based on the resource state information of the target object set. Since the resource state information of the target object set is obtained by analyzing the resource processing methods associated with the objects in the object set, and these objects belong to the same object set as the target object, it means that the resource processing methods between these objects are also very close. Therefore, analyzing based on the resource state information of the object set can achieve fast and accurate state identification.

[0048] Other features and advantages of the present application will be described in the following specification, and part of them will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0050] Figure 1 is an optional schematic diagram of an application scenario in the embodiments of the present application;

[0051] Figure 2 is a schematic flowchart of a method for identifying the state of an object in the embodiments of the present application;

[0052] Figure 3 is a schematic diagram of an undirected graph G in the embodiments of the present application;

[0053] Figure 4 is a schematic diagram of merchant gang aggregation in the embodiments of the present application;

[0054] Figure 5 is a schematic diagram of a historical gang in the embodiments of the present application;

[0055] Figure 6 is a specific flowchart of combating real-time online merchant transactions for newly registered merchants in the embodiments of the present application;

[0056] Figure 7A schematic flowchart of an offline model training and online model incremental calculation in an embodiment of the present application;

[0057] Figure 8 A schematic flowchart of the training process of a state recognition model in an embodiment of the present application;

[0058] Figure 9 A flowchart of the implementation of a quasi-real-time deployment solution in an embodiment of the present application;

[0059] Figure 10 A schematic diagram of the composition structure of an object state recognition device in an embodiment of the present application;

[0060] Figure 11 A schematic diagram of the hardware composition structure of a first electronic device applying the embodiment of the present application;

[0061] Figure 12 A schematic diagram of the hardware composition structure of a second electronic device applying the embodiment of the present application. Detailed implementation manners

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0063] Some concepts involved in the embodiments of the present application are introduced below.

[0064] Connectivity of vertices and connected graphs: In an undirected graph G, if there is a path from vertex v i to vertex v j then vertices v i and v j are said to be connected. Among them, an undirected graph refers to a graph where the edges have no direction. In an undirected graph G, if any two different vertices in V(G) are connected (i.e., there is a path), then G is called a connected graph.

[0065] Connected components: In graph theory, a maximal connected subgraph of an undirected graph G is called a connected component. In a connected component, any two vertices are connected to each other by a path. Any connected graph has only one connected component, that is, the connected graph itself. A non-connected undirected graph is composed of multiple connected components, and there is no path connecting the connected components to each other.

[0066] BFS (Breadth First Search, breadth - first search): Also known as width - first search, it is one of the simplest graph search algorithms and is also the prototype of many important graph algorithms. It belongs to a blind search method, aiming to systematically expand and check all nodes in the graph to find the result. In other words, it does not consider the possible location of the result, but thoroughly searches the entire graph until the result is found.

[0067] DFS (Depth First Search, depth - first search): The purpose is to reach the leaf nodes of the searched structure (i.e., those HTML (HyperText Markup Language) files that do not contain any hyperlinks). In an HTML file, when a hyperlink is selected, the linked HTML file will perform a depth - first search, that is, a single chain must be completely searched before searching the results of the remaining hyperlinks. The depth - first search follows the hyperlinks on the HTML file until it can no longer go deeper, then returns to an HTML file and continues to select other hyperlinks in that HTML file. When there are no more hyperlinks to select, it means the search has ended.

[0068] Resource status information and resource processing methods: Taking the merchant as the object, the resource status information mainly refers to the transaction status of the merchant set, whether it is good or suspicious. The same principle applies when the object is an individual. In this article, the merchant is mainly used as an example for detailed introduction. The resource processing methods mainly refer to the relevant operation methods, information, etc. when the object conducts transactions. For example, whether there are illegal operations such as brushing orders, and it can further include transaction details information.

[0069] The embodiments of this application relate to artificial intelligence (Artificial Intelligence, AI) and machine learning technologies, and are designed based on computer vision technology and machine learning (Machine Learning, ML) in artificial intelligence.

[0070] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence.

[0071] Artificial intelligence is also about studying the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technologies mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in multiple fields, such as common smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, robots, intelligent healthcare, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.

[0072] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0073] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Compared with data mining that finds mutual characteristics from big data, machine learning pays more attention to the design of algorithms, enabling computers to automatically "learn" rules from data and use the rules to predict unknown data.

[0074] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. In the embodiments of this application, when performing state recognition on a target object, a state recognition model based on machine learning or deep learning is adopted. Based on this model, the target object is grouped to determine the object set to which the object belongs, and then the target resource state of the target object is determined based on the resource state information of the object set. The resource state of the object identified by the method in the embodiments of this application is more accurate.

[0075] The method for training a state recognition model proposed in the embodiment of the present application can be divided into two parts, including a training part and an application part; wherein, the training part involves the technical field of machine learning. In the training part, the state recognition model is trained by the technology of machine learning, so that the training samples containing candidate objects given in the embodiment of the present application are used to train the state recognition model. After the training samples pass through the state recognition model, the output results of the state recognition model are obtained, and the model parameters are continuously adjusted through the optimization algorithm in combination with the output results; the application part is used to use the state recognition model trained in the training part to perform state recognition on the target object, and obtain the target resource state of the target object, so as to better manage the object and implement strategic strikes on bad objects, etc. In addition, it should be noted that the state recognition model in the embodiment of the present application can be either online training or offline training, which is not specifically limited here.

[0076] The following is a brief introduction to the design concept of the embodiment of the present application:

[0077] With the development of Internet technology, people's shopping patterns are becoming increasingly diversified. Shopping platforms in contemporary society are not limited to cash payment shopping, but also use payment applications to scan codes or scan faces for shopping. Of course, merchants need to first settle in payment applications before customers can use the payment application to make payment transactions.

[0078] Among the merchants settled in payment applications, the quality of merchants is usually uneven. The number of normal merchants generally accounts for the vast majority, but there are also some bad merchants hidden among them. High-quality merchants that provide most of the transaction contributions should be supported, while bad merchants that violate regulations, disrupt the health of the ecosystem, and harm the security of users' funds and experience should be cracked down. For example, the black industry registers merchants in batches in a short period of time, engages in gambling, fraud, etc., and the transaction amount grows explosively, and the transaction time window is short. However, the names of these merchants are not exactly the same, but they are highly similar. In order to better manage merchants, it is necessary to assign corresponding merchant operation strategies to each merchant to encourage or crack down on them. When identifying and cracking down on these bad merchants based on the risk control strategy platform, the data rules of the merchant dimension are mainly used for strategic crackdowns, which have low recognition accuracy and poor crackdown effects.

[0079] In view of this, the embodiments of the present application provide a method, an apparatus, an electronic device, and a storage medium for identifying the state of an object. Since the embodiments of the present application design a naming pattern, mine and analyze the naming characteristics of the object name, and determine the similarity between objects based on this characteristic, and then divide the object set, so as to realize the mining of the object set with high object name similarity. Therefore, when identifying the state of the target object, only the naming pattern of the target object needs to be used to analyze the object set to which the target object belongs, and then the state of the target object is analyzed based on the resource state information of the target object set. Since the resource state information of the target object set is obtained by analyzing the resource processing methods associated with the objects in the object set, and these objects belong to the same object set as the target object, it means that the resource processing methods between these objects are also very close. Therefore, analyzing based on the resource state information of the object set can achieve fast and accurate state identification.

[0080] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0081] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of the present application. The application scenario diagram includes two terminal devices 110 and a server 120. The terminal devices 110 and the server 120 can communicate through a communication network.

[0082] In an alternative embodiment, the communication network is a wired network or a wireless network. The terminal devices 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.

[0083] In the embodiments of the present application, the terminal device 110 is an electronic device used by a user. The electronic device can be a personal computer, a mobile phone, a tablet computer, a notebook, an e-book reader, a smart home, etc., which are computer devices with certain computing capabilities and running instant messaging software and websites or social software and websites. Each terminal device 110 communicates with the server 120 through a wireless network. The server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0084] Among them, the status recognition model can be deployed on the server 120 for training. A large number of training samples can be stored in the server 120 for training the status recognition model. Optionally, after the status recognition model is trained based on the training method in the embodiments of the present application, the trained status recognition model can be directly deployed on the server 120 or the terminal device 110. Generally, the status recognition model is directly deployed on the server 120. In the embodiments of the present application, the status recognition model is mainly used to recognize the status of the target object, including but not limited to merchant gang recognition, and can also be extended to other scenarios with naming similarity such as personal gang recognition scenarios. An online training method or an offline training technical solution can be adopted for offline analysis.

[0085] In addition, in the model deployment part of the embodiments of the present application, it includes but is not limited to the online deployment method of incremental calculation and flushing cache features. It is also possible to perform offline calculation of group information and group information according to the application scenario, and directly flush the model results into the cache features for online deployment, which will not be specifically limited here.

[0086] In a possible application scenario, the training samples in the present application can be stored using cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed storage file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0087] In a possible application scenario, in order to facilitate reducing communication latency, servers 120 can be deployed in each region, or for load balancing, different servers 120 can respectively serve the regions corresponding to each terminal device 110. Multiple servers 120 can achieve data sharing through blockchain. Multiple servers 120 are equivalent to a data sharing system composed of multiple servers 120. For example, the terminal device 110 is located at location a and is communicatively connected to the server 120, and the terminal device 110 is located at location b and is communicatively connected to other servers 120.

[0088] For each server 120 in the data sharing system, it has a node identifier corresponding to the server 120. Each server 120 in the data sharing system can store the node identifiers of other servers 120 in the data sharing system, so that subsequently, according to the node identifiers of other servers 120, the generated blocks can be broadcast to other servers 120 in the data sharing system. A node identifier list as shown in the following table can be maintained in each server 120, and the server 120 name and the node identifier are stored in the node identifier list correspondingly. Among them, the node identifier can be an IP (Internet Protocol) address and any other information that can be used to identify the node. Only the IP address is used as an example in Table 1 for illustration.

[0089] Table 1

[0090]

[0091]

[0092] Next, in combination with the above-described application scenario, the method for identifying the object state provided by the exemplary embodiment of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0093] Refer to Figure 2 As shown, it is a flowchart of the implementation of a method for identifying the object state provided by an embodiment of the present application. The specific implementation process of the method is as follows:

[0094] S21: Determine the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object;

[0095] S22: Determine the object set to which the target object belongs according to the naming pattern, where the object set is obtained by dividing according to the similarity between each object;

[0096] In the embodiment of the present application, first, it is necessary to perform word segmentation processing on the object name of the target object to obtain each included word segment. Among them, the word segment can be a word or a character, and no specific limitation is made here. In the embodiment of the present application, mainly words are used as examples for illustration.

[0097] In addition, the target object can refer to a merchant or an individual, etc. In the embodiment of the present application, mainly a newly registered merchant is used as an example for illustration. When a merchant registers, the object set to which the merchant belongs is determined based on the above process. In this way, when the merchant conducts a transaction, the target resource state of the merchant can be analyzed according to the resource state information of the object set, and it can be judged whether to crack down on it.

[0098] In the embodiments of the present application, by observing the data of merchant names, it can be known that there are several naming patterns (also known as part-of-speech patterns) in merchant registration names. For example, "Bingbingbang Department Store in XX District", that is, generally merchants will name themselves in the way of place name + personal name + business entity. The goal of part-of-speech tagging is to mark each word with a single label, which represents its usage and syntactic role, such as nouns, verbs, adjectives, etc.

[0099] For example, the target object is merchant 1, and the merchant name of this merchant is Bingbingbang Department Store in XX District. After word segmentation, three words are obtained, namely XX District # Bingbingbang # Department Store. Next, part-of-speech tagging is performed on these three words. The part-of-speech tagging result is NR#NR#NN, where NR represents the subclass of personal names (or place names) in the large category of nouns, and NN represents the subclass of work-related nouns in the large category of nouns.

[0100] In step S22, each object set is obtained by dividing according to the similarity between objects. However, when determining the target object, the similarity between the target object and each object in the object set is unknown. Therefore, it is necessary to determine the object set to which the target object belongs according to the naming pattern. Finally, the target resource state of the target object is analyzed based on the resource state information of this object set, as shown in step S23.

[0101] S23: Determine the target resource state of the target object according to the resource state information of the object set, where the resource state information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

[0102] In the embodiments of the present application, the resource state information of a certain object set is determined according to the resource processing methods associated with some or all of the objects in this object set. Among them, the resource state information of the object set is mainly used to indicate whether there is abnormal operation, suspicious transaction, or normal state, etc.

[0103] Generally, it will be more accurate to analyze and obtain the resource state information of the object set based on the resource processing methods associated with all objects. Of course, it is also possible to analyze and obtain the resource state information of the object set based on the resource processing methods of some objects, and no specific limitation is made here.

[0104] In the above embodiments, by designing a naming pattern, the naming characteristics of the analysis object names are mined, analyzed from the dimension of the object names, and the similarity between objects is determined based on this characteristic, and then the object set is divided to realize the mining of the object set with high object name similarity. Therefore, when identifying the state of the target object, only the naming pattern of the target object needs to be used to analyze the object set to which the target object belongs, and then the state of the target object is analyzed based on the resource state information of the target object set. Since the resource state information of the object set to which the target object belongs is obtained by analyzing the resource processing methods associated with the objects in the object set, and these objects belong to the same object set as the target object, it means that the resource processing methods between these objects are also very close. Therefore, analyzing based on the resource state information of the object set can achieve fast and accurate state recognition.

[0105] In an alternative embodiment, the naming patterns of the objects in the same object set are the same, and the similarity between any two objects reaches a preset similarity threshold. Therefore, before determining the object set to which the target object belongs, it is first necessary to divide the registered objects into object sets, which can be specifically divided into the following two division methods:

[0106] Division method one:

[0107] Obtain the similarity between any two objects; for each object with the same naming pattern, query each connected component in the connected graph constructed based on the similarity between the objects; respectively form an object set with the objects corresponding to the vertices in each connected component, where the connected graph is constructed by taking each object as a vertex and connecting the vertices corresponding to any two objects whose similarity reaches the similarity threshold.

[0108] In graph theory, a graph G is defined as consisting of a finite non-empty vertex set V(G) and a finite edge set E(G), denoted as:

[0109] G=(V(G),E(G))

[0110] Among them, each element of E(G) is an unordered pair of vertices in V(G), called an edge of G.

[0111] In an undirected graph G, if there is a path from vertex v i to vertex v j then vertex v i and vertex v jis connected. If any two different vertices in V(G) are connected (i.e., there is a path), then G is called a connected graph. A maximal connected subgraph of an undirected graph G is called a connected component, and any two vertices in the connected component are connected to each other through a path. There is only one connected component in any connected graph, that is, the connected graph itself, and a non-connected undirected graph is composed of multiple connected components.

[0112] For example Figure 3 As shown, it is a schematic diagram of an undirected graph G listed in the embodiments of the present application. Among them, graph G is composed of 3 maximal connected subgraphs. Any two vertices in the 3 connected components are connected to each other. There is no path connecting the connected components to each other.

[0113] There are also many methods for calculating the connected components of a graph, such as using breadth-first search (BFS) or depth-first search (DFS). Taking the depth-first search algorithm as an example, the algorithm logic for finding the connected components of an undirected graph is as follows:

[0114]

[0115]

[0116] In the above algorithm, the content after " / / " after each line of code represents the annotation (explanation) of this line of code. For example, end means end. Among them, for...do... represents a loop statement in the C language. Among them, the DFS(V, k) program is as follows:

[0117]

[0118] Based on the DFS algorithm, starting from a vertex v of graph G, visit any of its adjacent vertices w1, and then starting from w1, visit the adjacent vertices that have not been visited until all adjacent vertices have been visited. Then, take a step back to the previously visited vertex and check whether there are still adjacent vertices that have not been visited. And so on, repeat the above process until all vertices in the connected graph have been visited. Through DFS, all the connected components of the undirected graph can be obtained.

[0119] It should be noted that the above-listed algorithm codes are only examples and are not limited to this. No specific limitations are made here.

[0120] Taking the scenario of the present application as an example, calculate the connected components of merchants with the same naming pattern. Take the merchants as vertices, take the similarity between merchants as edges, and at the same time set a similarity threshold. When the similarity is greater than or equal to the similarity threshold, the similarity between merchants can be used as an edge to construct a graph. Regarding the setting of the similarity threshold, through experiments, a similarity threshold setting listed in the present application is shown in Table 2, where VV represents a verb:

[0121] Table 2

[0122] Number of parts of speech Threshold Example Less than or equal to 3 0.68 For example: NR#NR#NN 4 0.75 For example: NR#NR#VV#NN 5 0.8 For example: NR#NN#NN#NN#NN 6 0.83 For example: NR#NN#NN#NN#NN#NN 7 0.85 For example: NN#NN#NN#NN#NN#NN#NN Greater than or equal to 8 0.87 For example: NR#NR#NR#VV#NN#NN#NN#NN

[0123] That is to say, the similarity threshold in the embodiments of the present application is related to the number of parts of speech. When the number of parts of speech is different, the corresponding similarity threshold may be different, as shown in the above table. In the embodiments of the present application, mainly taking the number of parts of speech as 3 as an example for illustration, the corresponding similarity threshold is 0.68. Through the above threshold setting, the connected components under the same naming pattern are used as the standard for aggregating merchant groups (merchant sets), and the merchant group aggregation as shown in Figure 4 can be obtained, which includes a total of three merchant sets.

[0124] Among them, 6 miscellaneous goods stores in Jiang, D County on the upper left form a merchant set. In the upper right, 8 merchants including Fu's Department Store in County A, Hao's Department Store in County A, Hao's Department Store in County A, Fu's Department Store in County A, Pan's Department Store in County A, Pan's Department Store in County A, Di's Department Store in County A, and Di's Department Store in County A form a merchant set. Figure 4 In the lower middle, 6 miscellaneous goods stores in Jiang, D County and 6 general merchandise stores in Wu, D County form a merchant set. It should be noted that Figure 4 there are some merchants with the same merchant name but different geographical locations, which can be regarded as different merchants.

[0125] From Figure 4 it can be seen that through the calculation of connected components, merchants with the naming pattern of [D County] + XX + [Department Store] are grouped into one category, and merchants with the naming pattern of [D County] + XX + [General Merchandise Store] are grouped into one category. Similarly, the same principle applies to other merchants. It is precisely such merchants with this naming pattern that become a powerful input for suspicious gang analysis.

[0126] Partition method two

[0127] Obtain the similarity between any two objects; for each object with the same naming pattern, cluster the objects based on the similarity between the objects to obtain at least one object set.

[0128] In the embodiments of the present application, when partitioning each object into a set by clustering, an unsupervised model or a deep learning model can be used for partitioning, etc. It is mainly clustered based on the similarity between the objects, and no specific limitation is made here.

[0129] In the embodiments of the present application, merchants under the same naming pattern are subjected to similarity calculation to measure the similarity and correlation between merchants. Through similarity, pairwise similar merchants can be associated together. Considering that some static information such as legal person information and ID card information registered and used by generally similar suspicious groups has some aggregation characteristics. Therefore, in the embodiments of the present application, when calculating the similarity between merchants, in addition to considering the similarity of merchant names, some aggregation characteristics of merchant registration information are also considered, and these aggregation information are incorporated into the similarity calculation, thereby defining the weighted similarity calculation.

[0130] In an alternative embodiment, the similarity between any two objects is obtained by weighting the text similarity between the object names of these two objects and the information similarity between the corresponding registration information.

[0131] Taking merchants as an example, the registration information of merchants includes legal person information, bank card information, mobile phone numbers, etc. When calculating the information similarity between the registration information of two merchants, it is necessary to calculate the legal person information similarity, bank card information similarity, and mobile phone number similarity of the two merchants respectively.

[0132] In the embodiments of the present application, for two merchants, the corresponding text similarity or information similarity is calculated based on the Dice coefficient, where the Dice coefficient is a set similarity metric function, usually used to calculate the similarity between two samples. The formula is as follows:

[0133]

[0134] Taking the calculation of text similarity as an example, when calculating the text similarity of the merchant names of merchant A and merchant B based on the above formula 1, A represents the merchant name of merchant A, B represents the merchant name of merchant B, and both A and B can be expressed in vector form.

[0135] After calculating the text similarity and information similarity between merchant A and merchant B based on the above formula 1, the similarity between merchant A and merchant B can be obtained by weighting. When weighting the text similarity and information similarity, the similarity calculation formula is as follows:

[0136] Similarity t =αΥ t +(1-α)Similarity t-1 Formula 2

[0137] Υ t ∈{Similarity text ,Similarity law ,Similarity accountbank ,Similarity phone}

[0138] Among them, Υ t is the text similarity Similarity text , the legal person similarity Similarity law , the bank card similarity Similarity accountbank , the mobile phone similarity Similarity phone is an enumerated value of, and t is the number of iterations for adding the clustering similarity. In an optional implementation manner, for the legal person similarity, the bank card similarity, and the mobile phone similarity, if the merchant information is equal, it is 1, otherwise it is 0. The parameter α is the weight ratio and can be set according to experience, for example, set to 0.4. It should be noted that when t = 1, Similarity1 = Υ1 = Similarity text , and in the subsequent iteration process, it can be calculated according to the above formula 2.

[0139] Specifically, when weighting the text similarity and the information similarity, the iterative method listed above can be used, or weights can be set for the text similarity and the information similarity respectively and directly weighted to obtain. Below, mainly taking the formula listed above as an example, the calculation process in the iterative method is introduced:

[0140] It should be noted that when t = 1, Similarity1 = Similarity text ;

[0141] When t = 2, Υ2 = Similarity law , Similarity2 = αSimilarity law +(1 - α)Similarity1;

[0142] When t = 3, Υ3 = Similarity accountbank , Similarity3 = (1 - α)Similarity2 + αSimilarity accountbank ;

[0143] When t = 4, Υ4 = Similarity phone , Similarity4 = αSimilarity phone +(1 - α)Similarity3.

[0144] That is, the similarity between merchant A and merchant B finally obtained is Similarity4.

[0145] In the embodiments of the present application, through the above design calculation of similarity, the similarity between any two merchants can be obtained, such as text similarity, as shown in Table 3:

[0146] Table 3

[0147] Text similarity Luo Moumou Department Store in County A He Moumou Department Store in County A Nai Moumou Department Store in County A Luo Moumou Department Store in County A - 0.667 0.667 He Moumou Department Store in County A 0.667 - 0.667 Nai Moumou Department Store in County A 0.667 0.667 -

[0148] The registration information is as follows, as shown in Table 4:

[0149] Table 4

[0150] Merchant name Legal person information Bank card information Mobile phone number Luo Moumou Department Store in County A 7f122e**53528 fa0717**9cdc8 - He Moumou Department Store in County A 7f122e**53528 fa0717**9cdc8 - Nai Moumou Department Store in County A 7f122e**53528 fa0717**9cdc8 -

[0151] Then the final similarity results between merchants are as follows (α is set to 0.4):

[0152] Table 5

[0153] Similarity Luo Moumou Department Store in County A He Moumou Department Store in County A Nai Moumou Department Store in County A Luo Moumou Department Store in County A - 0.88 0.88 He Moumou Department Store in County A 0.88 - 0.88 Nai Moumou Department Store in County A 0.88 0.88 -

[0154] As can be seen from the above table, since the legal person information and bank card information of [Luo Moumou Department Store in County A] and [He Moumou Department Store in County A] are the same, so Similarity law , Similarity accountbank is set to 1, while the mobile phone numbers are all empty and not equal, so Similarity phone is 0. Set the threshold α to 0.4, that is, the weighted similarity is 0.88. Similarly, the similarity calculation between other merchants can be carried out in the same way.

[0155] In an alternative embodiment, after obtaining the naming pattern of the newly registered merchant, when performing step S22, it is specifically divided into the following situations:

[0156] Situation 1: There is at least one set of candidate objects with the same naming pattern as the target object.

[0157] In this case, it is necessary to obtain the similarity between the target object and each candidate object in each set of candidate objects, and then screen out at least one candidate object that reaches the preset similarity threshold based on the similarity between the objects; select one candidate object from the at least one screened candidate object as the candidate object that matches the target object; and use the set of candidate objects to which the matching candidate object belongs as the set of objects to which the target object belongs. Among them, when selecting one candidate object from the at least one screened candidate object, an alternative embodiment is to select the one with the highest similarity to the target object among these candidate objects. Of course, the second highest can also be selected, etc., and no specific limitation is made here.

[0158] Taking the target object as a newly registered merchant as an example, among them, the candidate object set is the above-listed historical (merchant) gangs, and the candidate object is the historical merchant. When the historical merchant has this naming pattern, obtain the historical merchant and group information under this naming pattern, calculate the similarity between the newly registered merchant and the historical merchant, compare according to the similarity threshold set above, obtain the most similar historical merchant, assign the gang ID to which the historical merchant belongs to the newly registered merchant, and update the group information. Among them, the group information includes the number of merchants in the gang, merchant information, etc.

[0159] Suppose there is a newly registered merchant M with a naming pattern of NR#NR#NN. Query the historical sub-group information under this naming pattern as Figure 5 shown as Figure 5 This is a schematic diagram of a historical gang in an embodiment of the present application, which contains a total of 5 historical gangs, namely GID: 1 to GID: 5. After calculating the similarity between the newly registered merchant M and the historical merchants under this naming pattern, the historical merchant A is the most similar to the newly registered merchant M and meets the threshold setting. Then, the newly registered merchant M is classified into the group of GID: 2 to which the historical merchant A belongs, and the group information is updated.

[0160] Case 2: There are multiple target objects, and for any one of the target objects, there is no candidate object whose similarity to the any one of the target objects reaches the similarity threshold.

[0161] In this case, it is necessary to obtain the similarity between any one of the target objects and each of the other target objects; screen out at least one other target object whose similarity to any one of the target objects reaches the similarity threshold; form a new object set with any one of the target objects and the at least one other target object screened out, and use the new object set as the object set to which the target object belongs.

[0162] Taking the target object as a newly registered merchant as an example, when there are multiple newly registered merchants, for any one of the newly registered merchants, first, it is also necessary to query at least one historical gang with the same naming pattern as the newly registered merchant, then calculate the similarity between the merchant and the historical merchants in the historical gang, and then compare it with the similarity threshold corresponding to this naming pattern. If the similarity threshold is not met, calculate the similarity between the merchants that do not meet the conditions, and use the threshold rule to standardize the threshold. If a connected component is added, and if a merchant is independent of other merchants, store the single vertex alone in the graph, that is, form an independent group.

[0163] Suppose there is a newly registered merchant M i , i≥1, with a naming pattern of NR#NR#NN. Query the historical sub-group information under this naming pattern as Figure 5 shown as, and the newly registered merchant Mi After calculating the similarity with historical merchants under this naming pattern and finding that no historical merchant's similarity meets the preset similarity threshold, the newly registered merchant M will be added. i Calculate pairwise similarities, add new connected components, increase the GID in order (starting from 6), update the undirected graph G, and update the clique information.

[0164] Case 3: There are multiple target objects, and for any one of the target objects, there is no set of candidate objects with the same naming pattern as this arbitrary target object.

[0165] In this case, similar to Case 2, it is also necessary to obtain the similarity between any one target object and each of the other target objects; screen out at least one other target object whose similarity with any one target object reaches the preset similarity threshold; form a new object set with any one target object and the at least one other target object screened out, and use the new object set as the object set to which the target object belongs.

[0166] Taking the newly registered merchant as an example of the target object, when there are multiple newly registered merchants, for any one of the newly registered merchants, it is necessary to use the naming pattern of this newly registered merchant as the new naming pattern, calculate the pairwise similarities of the merchants under this new naming pattern, set the edges using the threshold rule, and perform connected component calculation. Obtain the sub-clique information corresponding to the newly registered merchant and calculate the clique information.

[0167] Suppose the newly registered merchant is M i , i≥1, the naming pattern is VV#VV#VV, and no historical sub-clique information under this naming pattern has been queried yet. Then add this naming pattern, calculate the pairwise similarities of the merchants under this naming pattern, perform connected component calculation according to the threshold setting, assign each merchant its respective clique ID, and update the clique information.

[0168] The following is an example to illustrate the algorithm for incremental connected component calculation:

[0169]

[0170]

[0171] In the above implementation, through the design of graph incremental calculation, it is possible to quickly perform online sub-clustering, meet timeliness requirements, and combine with the risk control platform to conduct online real-time policy strikes.

[0172] It should be noted that in the offline training scheme and online incremental calculation scheme of this application, including but not limited to using connected components for gang identification, unsupervised models such as clustering or deep learning models can also be used for gang identification according to specific application scenarios.

[0173] In the embodiment of the present application, relying on the joint development of the strategy risk control platform by the Anti-Money Laundering and Risk Control Department Analysis Center and the Technology Center, real-time online merchant transaction strikes are carried out every hour through the combination of gang mining and strategies for newly registered merchants. The specific process is as follows Figure 6 shown, which is a schematic diagram of a specific process for carrying out real-time online merchant transaction strikes on newly registered merchants listed in the embodiment of the present application. Figure 6 The solid line in it represents the business flow, corresponding to the application process of the model, and the dotted line represents the data flow, corresponding to the training process of the model.

[0174] First step, whenever a merchant registers on the payment platform, the data is transmitted to the background server through risk control. The status recognition model deployed on the background server will perform incremental calculations every hour, update the merchant set to which the newly registered merchant belongs, and write the latest data into the online database. Second step, when a merchant conducts a transaction, risk control will obtain the variables calculated by the status recognition model for strategy strikes and intercept suspicious merchants.

[0175] Among them, in the first step, when determining the merchant set to which the newly registered merchant belongs based on the status recognition model, it is necessary to perform word segmentation and part-of-speech tagging on the merchant name of the newly registered merchant, match the previously divided historical gangs based on the naming pattern, and then perform similarity calculations, increase the connected components, and update the gangs. The specific process can refer to the above-listed Case 1 or Case 2 or Case 3.

[0176] Among them, CKV (massive distributed storage system) and TSSD are the storage systems listed in the embodiment of the present application, which are used to store two tables, namely merchant sub-group information and group information of each gang.

[0177] It should be noted that the status recognition method in the embodiment of the present application can also be implemented based on machine learning. Specifically, by inputting the object name of the target object into the first status recognition model, performing word segmentation processing on the object name based on the first status recognition model, and performing part-of-speech tagging on each word obtained through word segmentation processing; determining the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word; determining the object set to which the target object belongs according to the naming pattern of the target object; obtaining the target resource status of the target object determined according to the resource status information of the object set output by the first status recognition model; among them, the first status recognition model is trained based on a training sample data set, and the training sample data set includes at least multiple training samples, and each training sample includes the object name of the candidate object.

[0178] Machine learning can be divided into unsupervised learning and supervised learning. In the industrial field, supervised learning is the commonly used method. However, supervised learning requires a relatively high workload of manual label annotation, and the acquisition of anomalies in unknown risks is often less flexible than unsupervised learning. The embodiments of this application adopt the method of unsupervised learning. As Figure 7 shown, it is a schematic flow diagram of offline model training and online model incremental calculation listed in the embodiments of this application. In the application scenario, there are two processes. One is offline model training (the offline part), which includes data collection and cleaning, feature extraction, model training, and model storage. The other process is online model incremental calculation (the online part). The embodiments of this application propose an incremental calculation of the connected components of the graph in the model incremental calculation link to achieve real-time online policy strikes.

[0179] It should be noted that the embodiments of this application use an unsupervised model for real-time online policy strikes. Identify malicious merchant groups, improve the targeting of strikes, and facilitate auditors to locate business risks in a timely manner.

[0180] Refer to Figure 7 shown, in the online model incremental calculation part, it is necessary to obtain merchant registration data in real time, perform word segmentation and part-of-speech tagging on the merchant name, determine the corresponding naming pattern X, further query the merchant and group information with the same naming pattern X, calculate the incremental weighted similarity based on the methods listed above, perform incremental graph connected components, and update the registered merchant sub-groups and group information. For the specific implementation process, refer to the above embodiments, and the repeated parts will not be elaborated here.

[0181] During the offline training process, collect the historical merchant registration data to construct a training sample data set, which is prepared for subsequent data processing and model calculation. The following mainly introduces the offline model training part in detail:

[0182] First, perform word segmentation on the merchant name and then perform part-of-speech tagging. The results are shown in Table 6:

[0183] Table 6

[0184] Merchant name Word segmentation Part-of-speech tagging Luo Moumou Department Store in County A County A#Luo Moumou#Department Store NR#NR#NN He Moumou Department Store in County A County A#He Moumou#Department Store NR#NR#NN Nai Moumou Department Store in County A County A#Nai Moumou#Department Store NR#NR#NN

[0185] Among them, when performing part-of-speech tagging, it can be implemented based on Part-of-speech tagging (Chinese part-of-speech tagging). Part-of-speech tagging refers to a program that tags each word in the word segmentation result with a correct part of speech. Of course, other tagging methods can also be used, and no specific limitation is made here.

[0186] The part-of-speech pattern formed after part-of-speech tagging is the naming pattern mentioned in the embodiments of the present application. Merchants with the same naming pattern are grouped. Specifically, when grouping, the method of calculating connected components is mainly used as an example for illustration. Specifically, merchants with the same naming pattern are used as vertices, the similarity between merchants is used as edges, and a threshold is set. When the similarity is greater than or equal to the threshold, the similarity is regarded as an edge for graph construction, and the connected components are calculated, thereby constructing merchant groups. Then, based on the historical resource processing methods of the merchants in each merchant group, the resource status information corresponding to each candidate object set is determined. Among them, the historical resource processing method of a merchant includes transaction details information during the merchant's transaction process. The resource status information of the candidate object set is analyzed based on the historical resource processing methods of each merchant, and then the resource status information of each candidate object set is output via the status recognition model, such as: suspicious status, normal status, etc. Then, based on the unsupervised learning method, the parameters of the status recognition model are adjusted.

[0187] In the embodiments of the present application, by mining and analyzing the naming characteristics of merchant names, designing naming patterns, and mining groups with high similarity of merchant names, historical registered merchants are clustered into merchant groups. Regarding the reasons for group aggregation, the effect is obvious, the interpretability is high, and it is convenient for business personnel to analyze and locate risks. By grouping merchants and storing group information, it can be used as input for subsequent online real-time incremental grouping of merchant registration.

[0188] Refer to Figure 8 As shown, it is a schematic flowchart of the training process of the status recognition model in an offline training manner in the embodiments of the present application, specifically including the following steps:

[0189] S81: Select at least two training samples from the training sample dataset;

[0190] S82: Input the object names of each candidate object in each training sample into the second status recognition model respectively, and perform word segmentation processing on the object names in each training sample based on the second status recognition model;

[0191] S83: Perform part-of-speech tagging on each word segment obtained after word segmentation processing, and determine the naming pattern corresponding to each candidate object according to the part-of-speech tagging results of each word segment;

[0192] S84: Divide the candidate objects with the same naming pattern according to the similarity between each candidate object, and obtain at least one candidate object set;

[0193] S85: Respectively determine the resource status information corresponding to each candidate object set based on the historical resource processing methods of the candidate objects in each candidate object set;

[0194] S86: Output the resource status information of each candidate object set via the second status recognition model;

[0195] S87: Adjust the parameters of the second status recognition model according to the resource status information output by the second status recognition model;

[0196] S88: Determine whether the second status recognition model after parameter adjustment converges. If so, execute step S89; otherwise, return to step S81;

[0197] S89: Use the second status recognition model after parameter adjustment as the first status recognition model.

[0198] When adopting the online training method, considering that the online strategy strike has high requirements for timeliness and also has high requirements for the real-time performance of the model results, to meet this requirement, an almost real-time deployment method is proposed in the embodiments of the present application. In an alternative embodiment, in order to meet the requirements of real-time online strategy strikes, an incremental calculation scheme is designed for the unsupervised algorithm, so as to meet the requirements of model effect and timeliness online.

[0199] Refer to Figure 9 As shown, it is the implementation flowchart of an almost real-time deployment scheme given in the embodiments of the present application.

[0200] When a merchant transaction triggers a strategy, obtain the merchant dimension information and gang dimension information in CKV and TSSD, and perform online strategy calculation. Among them, the merchant dimension information mainly refers to Sp_key (merchant ID, also known as merchant number), GroupID (ID of the gang to which the merchant belongs), GroupID-ststus(T + 1) (resource status information of the gang to which the merchant belongs). Among them, the gang dimension information adopts the status of T + 1, which means that the gang status obtained by the merchant in the current transaction is the gang status value updated by the merchants in the gang after the previous transaction.

[0201] In an alternative embodiment, for a newly added merchant (target merchant), the resource status information of the merchant gang to which it belongs is determined according to the historical resource processing methods associated with other merchants in the merchant gang except the target merchant. For example Figure 9 The GroupID-ststus(T + 1) listed, that is, the gang status obtained by the merchant in the current transaction is the gang status value updated by the merchants in the gang after the previous transaction; after determining the target resource status of the target merchant according to the resource status information of the merchant gang, when performing resource processing based on the target merchant (that is, when the target merchant conducts a transaction), the resource processing method associated with the target merchant can also be obtained, specifically referring to Figure 9The transaction information shown can be obtained based on the distributed logging system KAFKA. Additionally, it is further necessary to include Sp_key and GroupID as identifiers. Furthermore, according to the resource processing method of the target merchant, update the resource status information of the merchant group to which the target merchant belongs, obtain the updated GroupID - ststus, store it in the group cache, and update the GroupID - ststus corresponding to the merchant cache in CKV. 3. In this way, for each merchant transaction, quasi - real - time variable calculations will be performed to obtain the calculations of TSSD group variables and statuses, and update the status.

[0202] In the above - mentioned embodiment, for the rapid group identification of newly registered merchants, group aggregation is carried out from the dimension of merchant name similarity, and merchant transaction strategies are attacked, providing a dimension for strategy attack for the merchant online transactions of the risk control platform, and effectively identifying merchant name similarity groups in a timely manner. Moreover, the online strategy can obtain merchant sub - group information and suspicious variable information in real time, and conduct strategy attacks on each transaction. By using a solution that combines NLP technology and graph technology, and designing an incremental model calculation, the model can perform strategy attacks in real time, effectively discover black - production groups, and quickly conduct group attacks, reducing the losses of the platform and users in terms of reputation and money.

[0203] Based on the same inventive concept, the embodiment of the present application also provides an object state recognition device. As Figure 10 shown, it is a schematic structural diagram of an object state recognition device 1000 in the embodiment of the present application, and may include:

[0204] A name processing unit 1001, configured to determine the naming pattern corresponding to the target object according to the part - of - speech tagging results of each word segment included in the object name of the target object;

[0205] A set determination unit 1002, configured to determine the object set to which the target object belongs according to the naming pattern, where the object set is obtained by dividing according to the similarity between each object;

[0206] A state recognition unit 1003, configured to determine the target resource state of the target object according to the resource state information of the object set, where the resource state information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

[0207] Optionally, the set determination unit 1002 is specifically configured to:

[0208] If there is at least one candidate object set with the same naming pattern as the target object, obtain the similarity between the target object and each candidate object in the at least one candidate object set;

[0209] Filter out at least one candidate object whose similarity to the target object reaches a preset similarity threshold;

[0210] Select one candidate object from the at least one filtered candidate object as the candidate object matching the target object;

[0211] Use the candidate object set to which the matching candidate object belongs as the object set to which the target object belongs.

[0212] Optionally, if there are multiple target objects, and for any one of the target objects, if there is no candidate object whose similarity to any one of the target objects reaches the similarity threshold, the set determination unit 1002 is specifically configured to:

[0213] Obtain the similarity between any one target object and each other target object;

[0214] Filter out at least one other target object whose similarity to any one target object reaches the similarity threshold;

[0215] Form a new object set by combining any one target object and the at least one other target object filtered out, and use the new object set as the object set to which the target object belongs.

[0216] Optionally, if there are multiple target objects, for any one of the target objects, the set determination unit 1002 is specifically configured to:

[0217] If there is no candidate object set with the same naming pattern as any one target object, obtain the similarity between any one target object and each other target object;

[0218] Filter out at least one other target object whose similarity to any one target object reaches the preset similarity threshold;

[0219] Form a new object set by combining any one target object and the at least one other target object filtered out, and use the new object set as the object set to which the target object belongs.

[0220] Optionally, the naming patterns of the candidate objects in the same candidate object set are the same, and the similarity between any two candidate objects reaches the similarity threshold.

[0221] Optionally, the apparatus further includes:

[0222] A set partitioning unit 1004, configured to obtain the similarity between any two objects before the set determination unit 1002 determines the object set to which the target object belongs according to the naming pattern; for each object with the same naming pattern, query each connected component in the connected graph constructed based on the similarity between each object; respectively form an object set by the objects corresponding to each vertex in each connected component, where the connected graph is constructed by taking each object as a vertex and connecting the vertices corresponding to any two objects whose similarity reaches the similarity threshold; or,

[0223] Obtain the similarity between any two objects; for each object with the same naming pattern, perform clustering on each object based on the similarity between each object to obtain at least one object set.

[0224] Optionally, the set partitioning unit 1004 is specifically configured to:

[0225] Obtain the text similarity between the object names of any two objects; and

[0226] Respectively obtain the information similarity between the registration information corresponding to any two objects;

[0227] Perform weighting on the text similarity and the information similarity corresponding to the registration information to obtain the similarity between any two objects.

[0228] Optionally, the resource status information of the object set is determined according to the historical resource processing methods associated with other objects except the target object in the object set; the apparatus further includes:

[0229] A status update unit 1005, configured to obtain the resource processing method associated with the target object when performing resource processing based on the target object after the status recognition unit determines the target resource status of the target object according to the resource status information of the object set;

[0230] Update the resource status information of the object set to which the target object belongs according to the resource processing method of the target object.

[0231] Optionally, the name processing unit 1001 specifically includes:

[0232] Input the object name of the target object into the first status recognition model, perform word segmentation processing on the object name based on the first status recognition model, and perform part-of-speech tagging on each word obtained through the word segmentation processing;

[0233] Determine the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word;

[0234] The status recognition unit 1003 is specifically configured to:

[0235] Obtain the target resource status of the target object determined according to the resource status information of the object set output by the first status recognition model;

[0236] Among them, the first status recognition model is trained based on a training sample data set, and the training sample data set includes at least multiple training samples, and each training sample includes the object name of the candidate object.

[0237] Optionally, the apparatus further includes:

[0238] The model training unit 1006 is used to train the first status recognition model in the following manner:

[0239] Select at least two training samples from the training sample data set;

[0240] Input the object names of the candidate objects in each training sample into the second status recognition model respectively, and perform word segmentation processing on the object names in each training sample based on the second status recognition model;

[0241] Perform part-of-speech tagging on each word segment obtained through word segmentation processing, and determine the naming pattern corresponding to each candidate object according to the part-of-speech tagging results of each word segment;

[0242] According to the similarity between the candidate objects, divide the candidate objects under the same naming pattern to obtain at least one candidate object set;

[0243] Based on the historical resource processing methods of the candidate objects in each candidate object set respectively, determine the resource status information corresponding to each candidate object set;

[0244] Output the resource status information of each candidate object set via the second status recognition model;

[0245] According to the resource status information output by the second status recognition model, perform at least one parameter adjustment on the second status recognition model to obtain the first status recognition model.

[0246] For the convenience of description, the above parts are divided into each module (or unit) according to functions and described separately. Of course, when implementing the present application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0247] After introducing the object status recognition method and apparatus of the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0248] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0249] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in an embodiment of the present application. The electronic device can be used to identify the state of an object. In one embodiment, the electronic device can be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device can be as Figure 11 shown, including a memory 1101, a communication module 1103, and one or more processors 1102.

[0250] The memory 1101 is used to store the computer program executed by the processor 1102. The memory 1101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0251] The memory 1101 can be a volatile memory (voldtile memory), such as a random access memory (random-access memory, RAM); the memory 1101 can also be a non-volatile memory (non-voldtilememory), such as a read-only memory, a flash memory (flash memory), a hard disk (hard disk drive, HDD), or a solid-state drive (solid-state drive, SSD); or the memory 1101 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 1101 can be a combination of the above memories.

[0252] The processor 1102 can include one or more central processing units (central processing unit, CPU) or be a digital processing unit, etc. The processor 1102 is used to implement the above object state recognition method when calling the computer program stored in the memory 1101.

[0253] The communication module 1103 is used to communicate with the terminal device and other servers.

[0254] In the embodiments of the present application, the specific connection media between the above-mentioned memory 1101, communication module 1103 and processor 1102 are not limited. In the embodiments of the present disclosure, Figure 11 it is shown that the memory 1101 and the processor 1102 are connected through a bus 1104, and the bus 1104 is represented by a thick line in Figure 11 The connection manners between other components are only for illustrative purposes and are not restrictive. The bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only a thick line is used to represent it in

[0255] The memory 1101 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the object state recognition method of the embodiments of the present application. The processor 1102 is used to execute the above-mentioned object state recognition method, as Figure 2 shown.

[0256] In another embodiment, the electronic device can also be other electronic devices, such as Figure 1 shown in the terminal device 110. In this embodiment, the structure of the electronic device can be as Figure 12 shown, including components such as a communication component 1210, a memory 1220, a display unit 1230, a camera 1240, a sensor 1250, an audio circuit 1260, a Bluetooth module 1270, and a processor 1280.

[0257] The communication component 1210 is used to communicate with the server. In some embodiments, it may include a WiFi (Wireless Fidelity) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.

[0258] The memory 1220 can be used to store software programs and data. The processor 1280 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1220. The memory 1220 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1220 stores an operating system that enables the terminal device 110 to operate. In the present application, the memory 1220 can store the operating system and various application programs, and can also store the code for executing the object state recognition method of the embodiments of the present application.

[0259] The display unit 1230 can also be used to display the information input by the user or the information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1230 may include a display screen 1232 disposed on the front of the terminal device 110. Among them, the display screen 1232 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1230 can be used to display the application operation interface of the shopping software in the embodiments of the present application.

[0260] The display unit 1230 can also be used to receive input digital or character information and generate signal inputs related to the user settings and function controls of the terminal device 110. Specifically, the display unit 1230 may include a touch screen 1231 disposed on the front of the terminal device 110, which can collect the touch operations of the user thereon or nearby, such as clicking buttons, dragging scroll boxes, etc.

[0261] Among them, the touch screen 1231 can cover the display screen 1232, or the touch screen 1231 and the display screen 1232 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply referred to as a touch display screen. In the present application, the display unit 1230 can display application programs and corresponding operation steps.

[0262] The camera 1240 can be used to capture static images. The user can send the images captured by the camera 1240 to the chat customer service through the shopping software, or upload them as shopping evaluations, etc. The camera 1240 can be one or more. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1280 to convert it into a digital image signal.

[0263] The terminal device may further include at least one sensor 1250, such as an acceleration sensor 1251, a distance sensor 1252, a fingerprint sensor 1253, a temperature sensor 1254. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0264] The audio circuit 1260, speaker 1261, and microphone 1262 can provide an audio interface between the user and the terminal device 110. The audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, and the speaker 1261 converts it into a sound signal for output. The terminal device 110 can also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1260 and then converted into audio data. The audio data is then output to the communication component 1210 for transmission to, for example, another terminal device 110, or the audio data is output to the memory 1220 for further processing.

[0265] The Bluetooth module 1270 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1270 to perform data interaction.

[0266] The processor 1280 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 1220 and calling data stored in the memory 1220, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1280 may include one or more processing units; the processor 1280 can also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1280. In this application, the processor 1280 can run the operating system, application programs, user interface display, and touch response, as well as the object state recognition method of the embodiments of this application. In addition, the processor 1280 is coupled to the display unit 1230.

[0267] In some possible implementation manners, various aspects of the object state recognition method provided in this application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the object state recognition method according to various exemplary embodiments of this application described above in this specification. For example, the computer device can execute the steps as Figure 2 shown.

[0268] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0269] The program product of the embodiments of the present application may employ a portable compact disk read-only memory (CD-ROM) and include program code and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0270] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0271] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0272] Those of ordinary skill in the art will understand that all or part of the steps for implementing the foregoing method embodiments may be completed by hardware associated with program instructions. The foregoing program may be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the foregoing method embodiments; and the foregoing storage medium includes: various media that can store program code, such as a removable storage device, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disk.

[0273] Alternatively, if the above integrated units in the embodiments of the present application are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.

[0274] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these changes and modifications of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for identifying the state of an object, characterized in that, The method includes: Determining a naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object; Obtaining the similarity between any two objects; for each object with the same naming pattern, dividing the objects into at least one object set based on the similarity between the objects; Determining the object set to which the target object belongs according to the naming pattern; Determining the target resource state of the target object according to the resource state information of the object set, where the resource state information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

2. The method according to claim 1, wherein The determining the object set to which the target object belongs according to the naming pattern of the target object specifically includes: If there is at least one candidate object set with the same naming pattern as the target object, obtaining the similarity between the target object and each candidate object in the at least one candidate object set; Filtering out at least one candidate object whose similarity with the target object reaches a preset similarity threshold; Selecting one candidate object from the filtered at least one candidate object as the candidate object matching the target object; Taking the candidate object set to which the matching candidate object belongs as the object set to which the target object belongs.

3. The method according to claim 2, wherein If there are multiple target objects, and for any one of the target objects, if there is no candidate object whose similarity with the any one of the target objects reaches the similarity threshold, then the determining the object set to which the target object belongs according to the naming pattern of the target object includes: Obtaining the similarity between the any one of the target objects and each other target object; Filtering out at least one other target object whose similarity with the any one of the target objects reaches the similarity threshold; Forming a new object set by the any one of the target objects and the filtered at least one other target object, and taking the new object set as the object set to which the target object belongs.

4. The method according to claim 1, characterized in that, If there are multiple target objects, when determining the object set to which the target object belongs according to the naming pattern of the target object, for any one of the target objects, it specifically includes: If there is no candidate object set with the same naming pattern as the any one of the target objects, obtaining the similarity between the any one of the target objects and each other target object; Filtering out at least one other target object whose similarity with the any one of the target objects reaches a preset similarity threshold; Forming a new object set by the any one of the target objects and the filtered at least one other target object, and taking the new object set as the object set to which the target object belongs.

5. The method according to any one of claims 2 to 4, characterized in that, The naming patterns of the candidate objects in the same candidate object set are the same, and the similarity between any two candidate objects reaches the similarity threshold.

6. The method according to any one of claims 1 to 4, characterized in that, Obtaining the similarity between any two objects; for each object with the same naming pattern, dividing the objects into at least one object set based on the similarity between the objects, including: Obtaining the similarity between any two objects; for each object with the same naming pattern, querying each connected component in the connected graph constructed based on the similarity between the objects; respectively forming an object set by the objects corresponding to the vertices in each connected component, where the connected graph is constructed by taking each object as a vertex and connecting the vertices corresponding to any two objects whose similarity reaches the similarity threshold; or, Obtaining the similarity between any two objects; for each object with the same naming pattern, clustering the objects based on the similarity between the objects to obtain at least one object set.

7. The method according to claim 6, characterized in that The obtaining the similarity between any two objects specifically includes: Obtaining the text similarity between the object names of the any two objects; and Respectively obtaining the information similarity between the registration information corresponding to the any two objects; Weighting the text similarity and the information similarity corresponding to the registration information to obtain the similarity between the any two objects.

8. The method according to claim 1, characterized in that, The resource status information of the object set is determined according to the historical resource processing methods associated with other objects in the object set except the target object; After determining the target resource status of the target object according to the resource status information of the object set, the method further includes: When performing resource processing based on the target object, obtaining the resource processing method associated with the target object; Updating the resource status information of the object set to which the target object belongs according to the resource processing method of the target object.

9. The method according to any one of claims 1 to 4, 7 to 8, characterized in that Determining the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment included in the object name of the target object; determining the object set to which the target object belongs according to the naming pattern of the target object; Determining the target resource status of the target object according to the resource status information of the object set specifically includes: Inputting the object name of the target object into a first status recognition model, performing word segment processing on the object name based on the first status recognition model, and performing part-of-speech tagging on each word segment obtained by the word segment processing; Determining the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word segment; Determining the object set to which the target object belongs according to the naming pattern of the target object; Obtaining the target resource status of the target object determined according to the resource status information of the object set output by the first status recognition model; Wherein, the first status recognition model is trained based on a training sample data set, and the training sample data set at least includes a plurality of training samples, and each training sample includes the object name of a candidate object.

10. The method according to claim 9, wherein The first status recognition model is trained in the following manner: Selecting at least two training samples from the training sample data set; Input the object names of each candidate object in each training sample into the second state recognition model respectively, and perform word segmentation processing on the object names in each training sample based on the second state recognition model; Perform part-of-speech tagging on each word obtained through word segmentation processing, and determine the naming pattern corresponding to each candidate object according to the part-of-speech tagging results of each word; According to the similarity between each candidate object, divide the candidate objects under the same naming pattern to obtain at least one candidate object set; Based on the historical resource processing methods of the candidate objects in each candidate object set respectively, determine the resource status information corresponding to each candidate object set; Output the resource status information of each candidate object set through the second state recognition model; According to the resource status information output by the second state recognition model, perform at least one parameter adjustment on the second state recognition model to obtain the first state recognition model.

11. An object state recognition device, characterized in that, Including: A name processing unit for determining the naming pattern corresponding to the target object according to the part-of-speech tagging results of each word included in the object name of the target object; A set division unit for obtaining the similarity between any two objects; for each object with the same naming pattern, based on the similarity between each object, divide each object into at least one object set; A set determination unit for determining the object set to which the target object belongs according to the naming pattern; A state recognition unit for determining the target resource state of the target object according to the resource status information of the object set, where the resource status information of the object set is determined based on the resource processing methods associated with some or all of the objects in the object set.

12. The device according to claim 11, characterized in that, The set determination unit is specifically used for: If there is at least one candidate object set with the same naming pattern as the target object, obtain the similarity between the target object and each candidate object in the at least one candidate object set; Screen out at least one candidate object whose similarity with the target object reaches a preset similarity threshold; Select one candidate object from the at least one screened candidate object as the candidate object matching the target object; Use the candidate object set to which the matching candidate object belongs as the object set to which the target object belongs.

13. The device according to claim 12, characterized in that, If there are multiple target objects, and for any one of the target objects, if there is no candidate object whose similarity with the any one of the target objects reaches the similarity threshold, then the set determination unit is specifically used for: Obtain the similarity between the any one of the target objects and each other target object; Screen out at least one other target object whose similarity with the any one of the target objects reaches the similarity threshold; Form a new object set with the any one of the target objects and the at least one other target object screened out, and use the new object set as the object set to which the target object belongs.

14. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 10.

15. A computer-readable storage medium, characterized in that, It includes program code, and when the program code runs on an electronic device, the program code is used to cause the electronic device to execute the steps of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Target word determination method, device and storage medium

    CN109271624A

  • Merchant information determination method and device, electronic equipment and nonvolatile storage medium

    CN111833118A

  • Industry recognition model determination method and apparatus

    WO2020143377A1