User identification method and device, equipment and medium
By generating natural language description text and using a large language model for clustering, the problem of the inability to effectively utilize unstructured data in existing technologies is solved, accurate identification of abnormal user groups is achieved, and recognition efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510755165.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, it is impossible to effectively utilize unstructured user behavior data when identifying abnormal user groups, resulting in the inability to accurately identify multiple users with similar abnormal user behaviors.
By obtaining multiple types of user behavior data in the target application, natural language description text is generated, input into the large language model to obtain the description text vector, clustering is performed based on the vector distance, abnormal target clusters are identified, and abnormal user groups are determined.
It achieves effective identification of abnormal user groups, improves the efficiency of utilizing unstructured data, and enhances the accuracy and sensitivity of abnormal user identification.
Smart Images

Figure CN120670873A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of network security technology, and in particular to a user identification method, apparatus, device, and medium. Background Art
[0002] In the field of network security, identifying abnormal user groups is a critical step. Abnormal user groups can be understood as multiple users with similar abnormal user behaviors in an application.
[0003] In related technologies, the transaction behavior of each user is identified, and a node graph is constructed based on the users and transaction behaviors, wherein each node in the node graph is a user or a transaction behavior, and the edges between the nodes are constructed based on the associations between the nodes (such as common devices or IP addresses, etc.). Thus, the node graph is input into a convolutional neural network, and the convolutional neural network identifies multiple users with similar abnormal user behaviors by analyzing the node graphs of multiple users.
[0004] However, the above-mentioned method for identifying multiple users with similar abnormal user behaviors relies on the construction of a node graph, which only includes some structured data (nodes and edges in the graph). Some unstructured data cannot be included in the analysis (for example, user behavior data of users, etc.), resulting in the inability to effectively identify multiple users with similar abnormal user behaviors. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a user identification method, apparatus, device and medium.
[0006] An embodiment of the present disclosure provides a user identification method, the method comprising: obtaining a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple categories of user behavior data of the user in the target application, wherein the multiple categories of user behavior data include at least one category of login behavior data and / or at least one category of application usage behavior data; generating a natural language description text for describing the user behavior data set; inputting the natural language description text into a large language model to obtain a description text vector output by the large language model; clustering all the user behavior data sets within the preset time period according to the vector distance between all the description text vectors within the preset time period to obtain multiple target clustering clusters, wherein each target clustering cluster includes multiple user behavior data sets; identifying whether each category of the user behavior data in the user behavior data set in each target clustering cluster falls within a corresponding preset abnormal behavior data range; determining an abnormal target clustering cluster among the multiple target clustering clusters according to the identification result, wherein all users corresponding to the abnormal target clustering cluster are determined to be an abnormal user group.
[0007] An embodiment of the present disclosure also provides a user identification device, which includes: a first acquisition module, used to obtain a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple categories of user behavior data of the user in the target application, wherein the multiple categories of user behavior data include at least one category of login behavior data, and / or at least one application usage behavior data; a generation module, used to generate a natural language description text for describing the user behavior data set; a second acquisition module, used to input the natural language description text into a large language model to obtain a description text vector output by the large language model; an identification module, used to identify whether each category of the user behavior data in the user behavior data set in each of the target clusters belongs to a corresponding preset abnormal behavior data range; a determination module, used to determine an abnormal target cluster cluster among the multiple target cluster clusters based on the identification result, wherein all users corresponding to the abnormal target cluster cluster are determined to be an abnormal user group.
[0008] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the user identification method provided by the embodiment of the present disclosure.
[0009] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the user identification method provided by the embodiment of the present disclosure.
[0010] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art:
[0011] The user identification solution provided by the embodiment of the present disclosure obtains a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple types of user behavior data of the user in the target application, wherein the multiple types of user behavior data include at least one type of login behavior data and / or at least one type of application usage behavior data, generates a natural language description text for describing the user behavior data set, inputs the natural language description text into a large language model to obtain a description text vector output by the large language model, clusters all user behavior data sets within the preset time period based on the vector distance between all description text vectors within the preset time period to obtain multiple target clusters, wherein each target cluster includes multiple user behavior data sets, identifies whether each type of user behavior data in the user behavior data set in each target cluster falls within a corresponding preset abnormal behavior data range, and then determines an abnormal target cluster cluster from the multiple target cluster clusters based on the identification results, wherein all users corresponding to the abnormal target cluster cluster are determined to be an abnormal user group. In this technical solution, effective identification of abnormal user groups is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0013] Figure 1 A flowchart of a user identification method provided in an embodiment of the present disclosure;
[0014] Figure 2 A schematic diagram of the structure of a user identification device provided in an embodiment of the present disclosure;
[0015] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0017] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0018] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] In order to solve the above problems, the embodiments of the present disclosure provide a user identification method, which is introduced below in conjunction with specific embodiments.
[0023] Figure 1 This is a flow chart of a user identification method provided by an embodiment of the present disclosure. The method can be executed by a user identification device, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. Figure 1 As shown, the method includes:
[0024] Step 101: Obtain a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple categories of user behavior data of the user in the target application, wherein the multiple categories of user behavior data include at least one category of login behavior data and / or at least one category of application usage behavior data.
[0025] The duration of the preset period can be calibrated according to the application scenario. The preset period can be a period of time that is a preset length from the current period, so that user identification is performed once every preset period. The users to be identified include multiple users who log in to the target application within the preset period. The users to be identified can include every user who logs in to the target application within the preset period, or can include some users who log in to the target application within the preset period. The screening method for some users can be set according to the needs of the scenario. For example, some users whose login addresses are within a preset address range can be screened as users to be identified.
[0026] In one embodiment of the present disclosure, a user behavior data set of a user to be identified within a preset time period is obtained, wherein the user behavior data set corresponds one-to-one to the user to be identified, and each data set of each user to be identified includes multiple categories of user behavior data, and the multiple categories of user behavior data include at least one category of login behavior data, and / or at least one application usage behavior data, wherein the login behavior data includes login frequency, device type (virtual login device or real login device, etc.), IP address, etc., and the application usage behavior data includes transaction records, address location, etc.
[0027] Step 102: Generate a natural language description text for describing the user behavior data set.
[0028] In an embodiment of the present disclosure, the user behavior data set is represented by a natural language description text, that is, the user behavior data set is converted into high-dimensional semantics to ensure the accuracy of subsequent identification of abnormal logins.
[0029] In some possible embodiments, a user behavior data set may be input into a large language model, and the large language model may output a natural language description text based on the input user behavior data set.
[0030] In some possible embodiments, a description text segment corresponding to each type of user behavior data may be determined, and multiple description text segments corresponding to multiple types of user behavior data may be combined to generate a natural language description text of the user behavior data set.
[0031] In this embodiment, a corresponding description text segment is generated for each type of user behavior data, further improving the richness of semantic information.
[0032] Among them, in some possible implementations, each type of user behavior data can be matched with multiple preset standard behavior data ranges corresponding to each type of user behavior data, that is, multiple preset standard behavior data ranges corresponding to each type of user behavior data are pre-set, and the preset standard behavior data range for each type of user behavior data that is successfully matched is determined, and the preset description text segment corresponding to the preset standard behavior data range that is successfully matched is determined to be the description text segment. That is, in an embodiment of the present disclosure, a corresponding preset description text segment is set in advance for each first preset behavior data rule. For example, when the user behavior data is login frequency, multiple preset standard behavior data ranges may include: [w1, w2), [w2, w3), [w3, ∞), wherein the preset description text segment corresponding to [w1, w2) is that this is a low-frequency login device, the preset description text segment corresponding to [w2, w3) is that this is a normal login frequency device, and the preset description text segment corresponding to [w3, ∞) is that this is a high-frequency login device, etc. Furthermore, multiple description text segments corresponding to multiple types of user behavior data are combined to generate a natural language description text of the user behavior data set. For example, the obtained natural language description text may be "This is a high-frequency login device, using a virtual device, and accessed through a second-dial IP. The mobile phone is located at ****".
[0033] In step 103, the natural language description text is input into the large language model to obtain a description text vector output by the large language model. In the embodiment of the present disclosure, all user behavior data sets are clustered based on the natural language description text to obtain multiple target clusters. This fully utilizes the unstructured user behavior data for clustering, ensuring the effectiveness of subsequent identification of abnormal user groups.
[0034] In some possible embodiments, the natural language description text may be vectorized to obtain a description text vector. For example, the natural language description text may be input into a large language model (LLM) to obtain the description text vector output by the large language model. The description text vector is used to convert the natural language description text into a set of numbers (i.e., vectors). For example, the natural language description text A may be converted into a set of TF-IDF vectors "0.8, 0.5, 0.7, 0.6, 0.9, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0".
[0035] Step 104 : clustering all user behavior data sets within the preset time period according to the vector distances between all description text vectors within the preset time period to obtain a plurality of target clusters, wherein each target cluster includes a plurality of user behavior data sets.
[0036] The vector distance between the description text vectors corresponding to any two user behavior data sets in each target cluster is less than a first preset threshold, and the vector distance between the description text vectors corresponding to any two user behavior data sets in different target clusters is greater than a second preset threshold, where the second preset threshold is greater than or equal to the first preset threshold. That is, multiple target clusters are generated based on the vector distances between the description text vectors.
[0037] In some possible embodiments, the data difference between user behavior data of different users belonging to the same category can also be calculated separately. When the data difference is less than the corresponding difference threshold, it is determined that the corresponding user behavior data belongs to similar user behavior data, and different users whose number of similar user behavior data is greater than a preset number are clustered into a target cluster, etc.
[0038] Step 105 : determining an abnormal target cluster from the multiple target clusters according to the recognition result, wherein all users corresponding to the abnormal target cluster are determined to be an abnormal user group.
[0039] In one embodiment of the present disclosure, all user behavior data sets within a preset time period are clustered to obtain multiple target clusters, that is, all user behavior data sets are preliminarily clustered, each target cluster includes multiple user behavior data sets, and the user behavior data between each user behavior data set has similarity.
[0040] In one embodiment of the present disclosure, considering that the present disclosure is intended to identify multiple users with similar abnormal user behaviors, when clustering all user behavior data sets within a preset time period, the number of user behavior data sets contained in each candidate clustering cluster can be identified, and the candidate clustering clusters with a number of sets less than a preset number threshold can be filtered out, and only the candidate clustering clusters with a number of sets greater than or equal to the preset number threshold are retained as target clustering clusters.
[0041] In some possible embodiments, a corresponding preset abnormal behavior data range is set in advance for each type of user behavior data, wherein when each type of user behavior data belongs to the corresponding preset abnormal behavior data range, it is considered that the user behavior data may be abnormal user behavior data, wherein the preset abnormal behavior data range can be set by relevant risk control personnel, etc.
[0042] For example, when the user behavior data is login frequency, the preset abnormal behavior data range is [a, +∞). Therefore, when the login frequency is greater than a, it can be considered that the login frequency is high, and the corresponding user behavior data can be considered abnormal user behavior data.
[0043] In this embodiment, the reference number of user behavior data of each type in each target cluster that falls within the preset abnormal behavior data range corresponding to each type of user behavior data is identified, that is, the reference number of abnormal user behavior data of each type in each target cluster is identified. For example, if a target cluster contains three types of user behavior data, the reference number of abnormal user behavior data of each type of user behavior data is identified.
[0044] Furthermore, according to a plurality of reference numbers corresponding to the plurality of types of user behavior data, abnormal target cluster users are determined in the plurality of target clusters.
[0045] As a possible implementation method, the step of determining abnormal target cluster users in multiple target clusters based on multiple reference numbers corresponding to multiple types of user behavior data may include:
[0046] Calculate the ratio of the reference number and the total number of each type of user behavior data in each target cluster. For example, a target cluster contains three types of user behavior data s1, s2, and s3, where the total number corresponding to s1, s2, and s3 is 100, the reference number corresponding to s1 is 20, the reference number corresponding to s2 is 50, and the reference number corresponding to s3 is 30. Then the ratio of the number corresponding to s1 is 20%, the ratio of the number corresponding to s2 is 50%, and the ratio of the number corresponding to s3 is 30%.
[0047] Furthermore, a determination is made as to whether the ratio of the number of user behavior data items in each category is greater than a corresponding ratio threshold, where the ratio threshold can be set based on the scenario. When the number of categories of user behavior data items in the target cluster that are greater than the corresponding ratio threshold exceeds a preset number threshold, the user corresponding to the corresponding target cluster is determined to be an abnormal user. The preset number threshold corresponding to the number of categories can be set based on the communication scenario. For example, continuing with the above example, when the preset number threshold is 2, if the ratio of the number of user behavior data items corresponding to greater than or equal to two categories of user behavior data items in the three categories is greater than the preset ratio threshold, the corresponding target cluster is determined to be an abnormal target cluster.
[0048] In some possible examples, a standard value of normal user behavior data for each type of user behavior data can be pre-set, and the priority of each type of user behavior data can be pre-calibrated. The difference between each type of user behavior data of each user in each target cluster and the corresponding standard value of normal user behavior data is calculated, and the product value of the difference and the corresponding priority is calculated. The sum of all product values corresponding to all types of user behavior data corresponding to each user behavior data is calculated, and it is determined whether the sum value is greater than or equal to a preset sum value threshold. Users whose sum value is greater than or equal to the preset sum value threshold are determined to be candidate abnormal users. The ratio of the number of candidate abnormal users to all users in each cluster is calculated. When the user number ratio is greater than the user number ratio threshold, the target cluster is considered to be an abnormal target cluster.
[0049] Therefore, the identification method of the embodiment of the present disclosure determines the user group with abnormal logins through user behavior data, improves the sensitivity to user behavior when identifying abnormal users, and ensures the effectiveness of abnormal user identification.
[0050] In summary, the user identification method of the embodiment of the present disclosure obtains a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple types of user behavior data of the user in the target application, wherein the multiple types of user behavior data include at least one type of login behavior data and / or at least one type of application usage behavior data, generates a natural language description text for describing the user behavior data set, inputs the natural language description text into a large language model to obtain a description text vector output by the large language model, clusters all user behavior data sets within the preset time period according to the vector distance between all description text vectors within the preset time period to obtain multiple target clusters, wherein each target cluster includes multiple user behavior data sets, and then identifies whether each type of user behavior data in the user behavior data set in each target cluster falls within the corresponding preset abnormal behavior data range, and determines an abnormal target cluster cluster from the multiple target cluster clusters based on the identification results, wherein all users corresponding to the abnormal target cluster cluster are determined to be an abnormal user group. In this technical solution, effective identification of abnormal user groups with similar abnormal login behaviors is achieved.
[0051] In order to implement the above embodiments, the present disclosure also proposes a user identification device.
[0052] Figure 2 This is a schematic diagram of the structure of a user identification device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can generally be integrated into an electronic device. Figure 2 As shown, the device includes: a first acquisition module 210, a generation module 220, a second acquisition module 230, an identification module 240, and a determination module 250, wherein:
[0053] A first acquisition module 210 is configured to acquire a user behavior data set of a target application by a user to be identified within a preset time period, wherein the user behavior data set includes multiple types of user behavior data of the user in the target application, wherein the multiple types of user behavior data include at least one type of login behavior data and / or at least one type of application usage behavior data;
[0054] A generating module 220, configured to generate a natural language description text for describing a user behavior data set;
[0055] A second acquisition module 230 is configured to input the natural language description text into the large language model to obtain a description text vector output by the large language model;
[0056] An identification module 240 is configured to identify whether each type of user behavior data in the user behavior data set in each target cluster falls within a corresponding preset abnormal behavior data range;
[0057] The determination module 250 is configured to determine an abnormal target cluster among the multiple target clusters according to the recognition result, wherein all users corresponding to the abnormal target cluster are determined to be an abnormal user group.
[0058] In one embodiment of the present disclosure, the generating module 220′ is configured to:
[0059] Determining a descriptive text segment corresponding to each category of the user behavior data;
[0060] The multiple description text segments corresponding to the multiple types of user behavior data are combined to generate the natural language description text of the user behavior data set.
[0061] In one embodiment of the present disclosure, the generation module 220 is configured to:
[0062] Matching each category of user behavior data with a plurality of preset standard behavior data ranges corresponding to each category of user behavior data;
[0063] Determine the preset standard behavior data range for each type of user behavior data to successfully match;
[0064] The preset description text segment corresponding to the successfully matched preset standard behavior data range is determined as the description text segment.
[0065] In one embodiment of the present disclosure, the vector distance between the descriptive text vectors corresponding to any two of the user behavior data sets in each of the target clusters is less than a first preset threshold, and the vector distance between the descriptive text vectors corresponding to any two of the user behavior data sets in different target clusters is greater than a second preset threshold, wherein the second preset threshold is greater than or equal to the first preset threshold.
[0066] In one embodiment of the present disclosure, the identification module 240 is configured to identify whether each type of user behavior data corresponding to each target cluster belongs to a preset abnormal behavior data range corresponding to each type of user behavior data.
[0067] In one embodiment of the present disclosure, the identification module 240 is configured to: determine a reference number of user behavior data belonging to a preset abnormal behavior data range corresponding to each category of user behavior data according to each category of user behavior data;
[0068] Abnormal target cluster users are determined in the multiple target clusters according to the multiple reference numbers corresponding to the multiple types of user behavior data.
[0069] In one embodiment of the present disclosure, the identification module 240 is configured to: calculate the ratio of the reference number of user behavior data of each type in each target cluster to the total number of user behavior data;
[0070] Determine whether the ratio of the number of each type of user behavior data is greater than a corresponding ratio threshold;
[0071] When the number of the user behavior data classes in the target cluster that are greater than the corresponding ratio threshold is greater than a preset number threshold, the corresponding target cluster is determined to be the abnormal target cluster.
[0072] The user identification device provided in the embodiments of the present disclosure can execute the user identification method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0073] In order to implement the above embodiments, the present disclosure further proposes a computer program product, including a computer program / instruction, which implements the login identification method in the above embodiments when executed by a processor.
[0074] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.
[0075] The following specific reference Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. The electronic device 300 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0076] like Figure 3 As shown, the electronic device 300 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a memory 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processor 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0077] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0078] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the memory 308, or installed from the ROM 302. When the computer program is executed by the processor 301, the above-mentioned functions defined in the login identification method of the embodiment of the present disclosure are performed.
[0079] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0080] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0081] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0082] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the user identification method.
[0083] The electronic device may write computer program code for performing the operations of the present disclosure in one or more programming languages, or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0085] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0086] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0087] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0088] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0089] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0090] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A user identification method, characterized in that: include: Obtaining a user behavior data set of a to-be-identified user in a target application within a preset time period, wherein the user behavior data set includes multiple types of user behavior data of the user in the target application, wherein the multiple types of user behavior data include at least one type of login behavior data and / or at least one type of application usage behavior data; Generating a natural language description text for describing the user behavior data set; Inputting the natural language description text into a large language model to obtain a description text vector output by the large language model; Clustering all the user behavior data sets within the preset time period according to the vector distances between all the description text vectors within the preset time period to obtain a plurality of target clusters, wherein each target cluster includes a plurality of the user behavior data sets; Identify whether each type of the user behavior data in the user behavior data set in each target cluster belongs to a corresponding preset abnormal behavior data range; An abnormal target cluster is determined from the multiple target clusters according to the recognition result, wherein all users corresponding to the abnormal target cluster are determined to be an abnormal user group.
2. The method according to claim 1, wherein The generating of a natural language description text for describing the user behavior data set includes: Determining a descriptive text segment corresponding to each category of the user behavior data; The multiple description text segments corresponding to the multiple types of user behavior data are combined to generate the natural language description text of the user behavior data set.
3. The method according to claim 2, wherein Determining the description text segment corresponding to each type of user behavior data includes: Matching each category of user behavior data with a plurality of preset standard behavior data ranges corresponding to each category of user behavior data; Determine the preset standard behavior data range for each type of user behavior data to successfully match; The preset description text segment corresponding to the successfully matched preset standard behavior data range is determined as the description text segment.
4. The method according to claim 1, wherein The vector distance between the description text vectors corresponding to any two user behavior data sets in each target cluster is less than a first preset threshold, and the vector distance between the description text vectors corresponding to any two user behavior data sets in different target clusters is greater than a second preset threshold, wherein the second preset threshold is greater than or equal to the first preset threshold.
5. The method according to claim 1, wherein The step of identifying whether each type of user behavior data in the user behavior data set in each target cluster belongs to a user within a corresponding preset abnormal behavior data range includes: Identify whether each type of user behavior data corresponding to each target cluster belongs to a preset abnormal behavior data range corresponding to each type of user behavior data.
6. The method according to claim 5, wherein Determining an abnormal target cluster among the multiple target clusters according to the recognition result includes: A reference number of user behavior data belonging to a preset abnormal behavior data range corresponding to each category of user behavior data; Abnormal target cluster users are determined in the multiple target clusters according to the multiple reference numbers corresponding to the multiple types of user behavior data.
7. The method according to claim 6, wherein The determining of an abnormal target cluster from the plurality of target clusters according to the plurality of reference numbers corresponding to the plurality of types of user behavior data includes: Calculating the ratio of the reference number and the total number of user behavior data corresponding to each category in each target cluster; Determine whether the ratio of the number of each type of user behavior data is greater than a corresponding ratio threshold; When the number of the user behavior data classes in the target cluster that are greater than the corresponding ratio threshold is greater than a preset number threshold, the corresponding target cluster is determined to be the abnormal target cluster.
8. A user identification device, characterized in that: include: A first acquisition module is configured to acquire a user behavior data set of a user to be identified in a target application within a preset time period, wherein the user behavior data set includes multiple types of user behavior data of the user in the target application, wherein the multiple types of user behavior data include at least one type of login behavior data and / or at least one type of application usage behavior data; A generating module, configured to generate a natural language description text for describing the user behavior data set; A second acquisition module is configured to input the natural language description text into a large language model to obtain a description text vector output by the large language model; an identification module, configured to identify whether each type of user behavior data in the user behavior data set in each target cluster falls within a corresponding preset abnormal behavior data range; A determination module is configured to determine an abnormal target cluster among the multiple target clusters according to the recognition result, wherein all users corresponding to the abnormal target cluster are determined to be an abnormal user group.
9. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the user identification method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to execute the user identification method according to any one of claims 1 to 7.