User identification method, apparatus, device, medium, and product

CN118802295BActive Publication Date: 2026-09-15CHINA MOBILE COMM GRP CHONGQING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410454074.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2026-09-15
Estimated Expiration
2044-04-15

AI Technical Summary

Technical Problem

然而,这种方法存在一定的局限性,因为人工验证的准确性取决于分析人员的经验和能力,往往无法达到理想的识别准确度,由此导致异常用户识别的准确性较低

Benefits of technology

[0017] In a user identification method, apparatus, device, medium, and product provided in this application embodiment, firstly, by analyzing user business data, users who may exhibit abnormal online behavior are identified, forming a candidate abnormal user set. Next, a semantic recognition model is used to comprehensively analyze the user's online multimedia data, such as text, voice, video, and images, to obtain the first degree of anomalousness of the candidate abnormal user's online behavior. Finally, by evaluating and filtering the first degree of anomalousness of the candidate abnormal users, further refined screening is performed to ultimately determine the target abnormal user. The entire process, through the semantic recognition model and secondary screening, achieves accurate identification of abnormal users. The semantic recognition model, through in-depth analysis of the user's online multimedia data, provides a basis for identifying the target abnormal user. The semantic recognition model can more comprehensively and accurately capture the characteristics of abnormal behavior, improving the accuracy of abnormal user identification. The secondary screening, through refined evaluation and screening of the candidate abnormal users based on the first degree of anomalousness, further filters out the true target abnormal user from the candidate abnormal users, reducing the possibility of misidentification and further improving the accuracy of abnormal user identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118802295B_ABST
    Figure CN118802295B_ABST
Patent Text Reader

Abstract

The application provides a user identification method, device, equipment, medium and product. According to service information corresponding to each of a plurality of users, the plurality of users are identified to obtain candidate abnormal users, the candidate abnormal users being users with abnormal online behaviors in the plurality of users. Service data streams corresponding to the candidate abnormal users are input into a semantic identification model to obtain a first abnormal degree of online behaviors of the candidate abnormal users, the service data streams including online multimedia data of the users. According to the first abnormal degree, it is determined whether the candidate abnormal users are target abnormal users. The embodiment of the application can improve the accuracy of abnormal user identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a user identification method, apparatus, device, medium and product. Background Technology

[0002] The internet has become an integral part of people's lives and work, serving as a crucial channel for information exchange and data resource generation. However, to ensure network security, we must rigorously monitor and combat behaviors that may cause disruption. Therefore, accurate identification of suspicious users has become paramount.

[0003] However, current methods for identifying anomalous users primarily rely on manual verification, which involves manually analyzing user behavior and classifying it. This method has limitations, as the accuracy of manual verification depends heavily on the experience and skill of the analysts, often falling short of ideal accuracy, resulting in relatively low accuracy in identifying anomalous users. Summary of the Invention

[0004] This application provides a user identification method, apparatus, device, medium, and product that can improve the accuracy of abnormal user identification.

[0005] In a first aspect, embodiments of this application provide a user identification method, the method comprising:

[0006] Based on the business information corresponding to each user, user identification is performed on multiple users to obtain candidate abnormal users. Candidate abnormal users are users among multiple users who have abnormal Internet access behavior.

[0007] The business data stream corresponding to the candidate abnormal user is input into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's Internet access behavior. The business data stream includes the user's Internet multimedia data.

[0008] Based on the first degree of anomaly, determine whether the candidate anomaly user is the target anomaly user.

[0009] Secondly, this application provides a user identification device, the device comprising:

[0010] The first determination module is used to identify multiple users based on their respective business information to obtain candidate abnormal users. Candidate abnormal users are users among the multiple users who have abnormal internet access behavior.

[0011] The second determining module is used to input the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's Internet access behavior. The business data stream includes the user's Internet multimedia data.

[0012] The third determination module is used to determine whether a candidate abnormal user is the target abnormal user based on the first degree of abnormality.

[0013] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions;

[0014] When the processor executes computer program instructions, it implements the user identification method as described in any of the embodiments of the first aspect.

[0015] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the user identification method as described in any of the embodiments of the first aspect.

[0016] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform a user identification method as described in any of the embodiments of the first aspect above.

[0017] In a user identification method, apparatus, device, medium, and product provided in this application embodiment, firstly, by analyzing user business data, users who may exhibit abnormal online behavior are identified, forming a candidate abnormal user set. Next, a semantic recognition model is used to comprehensively analyze the user's online multimedia data, such as text, voice, video, and images, to obtain the first degree of anomalousness of the candidate abnormal user's online behavior. Finally, by evaluating and filtering the first degree of anomalousness of the candidate abnormal users, further refined screening is performed to ultimately determine the target abnormal user. The entire process, through the semantic recognition model and secondary screening, achieves accurate identification of abnormal users. The semantic recognition model, through in-depth analysis of the user's online multimedia data, provides a basis for identifying the target abnormal user. The semantic recognition model can more comprehensively and accurately capture the characteristics of abnormal behavior, improving the accuracy of abnormal user identification. The secondary screening, through refined evaluation and screening of the candidate abnormal users based on the first degree of anomalousness, further filters out the true target abnormal user from the candidate abnormal users, reducing the possibility of misidentification and further improving the accuracy of abnormal user identification. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1This is a flowchart illustrating a user identification method provided in one embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the structure of a user identification system provided in one embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of a user identification device provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0024] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0025] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0026] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0027] Existing methods for identifying anomalous users fall into two categories: one utilizes carrier data mining, requiring the establishment of an analysis platform within the intranet to acquire data from various data sources for modeling, and then constructing user profiles using user characteristics and business usage features. While this method can mine data from the entire user base, data acquisition is difficult, and the determination of anomalous users relies heavily on model fitting, lacking substantial evidence and thus having limited accuracy. The other method utilizes OTT (Over-the-Top) application data mining, but this is limited to some regular applications and cannot acquire data from irregular applications, resulting in an inability to achieve comprehensive mining. Furthermore, this method is computationally intensive, making it difficult to promote and apply across the entire network, thus leading to poor effectiveness in preventing and controlling cybersecurity risks.

[0028] To address the problems existing in the prior art, embodiments of this application provide a user identification method, apparatus, device, medium, and product.

[0029] This application provides a user identification method, apparatus, device, medium, and product. The user identification method provided in this application will be described first. For example... Figure 1 As shown, the method specifically includes the following steps:

[0030] S100 identifies multiple users based on their respective business information to obtain candidate abnormal users. Candidate abnormal users are those users who exhibit abnormal internet access behavior among the multiple users.

[0031] Optionally, in this embodiment of the application, it is first necessary to collect business information of different users, which can be collected from three aspects.

[0032] ① Data collection from the Authentication and Management Function (AMF)

[0033] The specific characteristics of data collected from the AMF include, but are not limited to, timestamps, occupied cells, access type (3rd Generation Partnership Project (3GPP) or non-3GPP), registration status, connection status, connectable status, user access method preference, and user location preference. The timestamp records the time of data collection to determine the time of the event. Occupied cells refer to the wireless communication cell the user is currently in, i.e., the base station the user is connected to. Access type refers to the way the user accesses the network, which can be a 3GPP network or a non-3GPP network, used to determine the user's network environment. Registration status refers to the user's registration status in the network, i.e., whether the registration process has been completed. Connection status refers to the user's current connection status, i.e., whether they are in a connected state. Connectable status refers to whether the user can establish a connection, i.e., whether the conditions for establishing a connection are met. User access method preference infers the user's possible future access method based on the user's past access records and behavioral characteristics. User location preference infers the user's possible future location based on the user's past location information and behavioral characteristics. This information can be used to analyze the user's activity patterns, behavioral habits, and current network conditions.

[0034] ② Collect data from the Session Management Function (SMF)

[0035] The specific characteristic data involved in the data collected from SMF includes, but is not limited to, timestamps, Protocol Data Unit (PDU) session type, Data Network Name (DNN), Public Land Mobile Network (PLMN) number, downlink status, Quality of Service Flow Identifier (QFI) value, PDU session status, session preference, and communication mode preference.

[0036] ③ Collect data from the Operations, Administration, and Maintenance (OAM) system.

[0037] OAM is the operation and maintenance system for the operator's communication network. This application collects relevant data by calling OAM's data service. The dataset involved is mainly user service data (and service information).

[0038] Specifically, this includes, but is not limited to, the following information to characterize a user: user ID (number anonymized), user's city of account opening, user's network tenure, user's account opening location, user's basic plan type (4G / 5G), user's plan name and code, basic plan data allowance, whether a regular data package is activated, data allowance included in the regular data package, off-peak data allowance, total data allowance, billing fee (RMB), out-of-plan data allowance (RMB), mobile TV data allowance, application data allowance, cell with the highest data usage in the current month, cell with the highest number of call minutes in the current month, cell with the most call volume in the current month, data usage generated under 4G network, call minutes generated under 4G network, data usage generated under 5G network, and call volume generated under 5G network. Minutes, participation in contract plans, contract expiration date, whether a user is a group user, whether a user on a shared billing plan, whether a member of a converged service, whether there are family number / cross-promotion / shared service volume relationships with other numbers, number of users contacted during calls, user level, total call duration (minutes), total number of calls in the current month, number of call forwardings, number of calls to other networks, number of days of data usage per month, data usage allowed within the user's main plan, whether the user is a registered user, the maximum data usage across all cell networks in the current month, the maximum number of call minutes across all cell networks in the current month, the maximum number of call minutes across all cell networks in the current month, and whether a universal terminal identification module is used. The information includes the Subscriber Identity Module (USIM) card, gender, whether 4G internet access is activated, whether 5G internet access is activated, user star rating, number of times data traffic is used on the 5G network, number of times voice calls are made on the 5G network, number of times data traffic is used on the 4G network, number of times voice calls are made on the 4G network, whether the 4G data package is activated, total outgoing call duration / minute, revenue per minute for voice calls exceeding the package limit / yuan, number of voice minutes included in the user's main package, free data included in the user's package, number of times data packages are activated in the current month, total voice minutes used in the current month, number of SMS messages sent in the current month, number of SMS messages received in the current month, number of MMS messages sent in the current month, and number of MMS messages received in the current month. It should be noted that the use of 4G and 5G in this application is merely illustrative and is not limited to these technologies; other communication technologies may also be used.

[0039] Optionally, in this embodiment, user ID (number anonymization) refers to the anonymization of a user's ID or identifier. In data processing, a unique user ID or identifier is used to identify and distinguish different users. However, to protect user privacy, these IDs can be anonymized to reduce the risk of sensitive information leakage. This anonymization can be achieved by not displaying the complete user ID, partially masking it, or replacing it.

[0040] Optionally, in one possible implementation of this application, business information of multiple users can be collected from different data sources. This information may include users' internet browsing history, call logs, SMS records, application usage, etc.

[0041] Subsequently, using the collected business information, a user profile is created for each user. User profiles can include information such as user behavior habits, frequently used applications, online time, and call duration.

[0042] By analyzing users' business information, users with abnormal internet behavior can be identified. These abnormal behaviors can be activities that do not conform to normal behavior patterns, such as a large number of unusual accesses, frequent connection attempts, or unusual application usage patterns.

[0043] Users exhibiting unusual internet behavior are identified and designated as potential anomaly users. These potential users may be involved in potential security risks or abnormal situations, requiring further analysis and confirmation. Through the above steps, users can be identified based on their respective business information, and potential anomaly users can be obtained for further analysis and processing.

[0044] S200, input the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's Internet access behavior. The business data stream includes the user's Internet multimedia data.

[0045] Optionally, in one possible implementation of this application, the online multimedia data of the candidate abnormal users is first obtained. This data may include the user's web browsing history, application usage history, downloaded files, etc. Subsequently, the extracted business data stream is preprocessed, including data cleaning, format conversion, noise removal, and other operations to ensure data quality and consistency.

[0046] Subsequently, a suitable semantic recognition model is selected, such as a deep learning-based natural language processing model or an image processing model, and trained. After training, the preprocessed business data stream is input into the constructed semantic recognition model for processing. For text data, it can be directly input into the natural language processing model; for image data, feature extraction and analysis can be performed through the image processing model.

[0047] Subsequently, based on the model's output, the initial degree of anomalousness of the candidate abnormal user's online behavior can be obtained. This initial degree of anomalousness can be represented by the model's output probability, score, or other indicators; a higher value indicates a higher degree of anomalousness. Through these steps, the business data stream corresponding to the candidate abnormal user can be input into the semantic recognition model to obtain the initial degree of anomalousness of the candidate abnormal user's online behavior, providing a reference for subsequent anomaly detection and processing.

[0048] S300, determine whether the candidate abnormal user is the target abnormal user based on the first degree of abnormality.

[0049] Optionally, in one feasible implementation of this application, a threshold or standard can first be set according to business needs and actual conditions. This threshold can be determined based on historical data and experience, or it can be based on the quantiles or statistics of the model output results.

[0050] Subsequently, the initial degree of anomalousness of the candidate anomalous user is compared with a set threshold. If the degree of anomalousness of the candidate anomalous user exceeds or reaches the set threshold, it is identified as a target anomalous user; otherwise, it is determined to be a non-target anomalous user.

[0051] Optionally, in this embodiment of the application, cases identified as target anomalous users may be further analyzed and processed. This may include further abnormal behavior detection, risk assessment, alarm triggering, etc.

[0052] By following the steps above, we can determine whether a candidate abnormal user is the target abnormal user based on the degree of anomaly. This judgment process can be adjusted and optimized according to actual circumstances to improve accuracy and efficiency.

[0053] In a user identification method, apparatus, device, medium, and product provided in this application embodiment, firstly, by analyzing user business data, users who may exhibit abnormal online behavior are identified, forming a candidate abnormal user set. Next, a semantic recognition model is used to comprehensively analyze the user's online multimedia data, such as text, voice, video, and images, to obtain the first degree of anomalousness of the candidate abnormal user's online behavior. Finally, by evaluating and filtering the first degree of anomalousness of the candidate abnormal users, further refined screening is performed to ultimately determine the target abnormal user. The entire process, through the semantic recognition model and secondary screening, achieves accurate identification of abnormal users. The semantic recognition model, through in-depth analysis of the user's online multimedia data, provides a basis for identifying the target abnormal user. The semantic recognition model can more comprehensively and accurately capture the characteristics of abnormal behavior, improving the accuracy of abnormal user identification. The secondary screening, through refined evaluation and screening of the candidate abnormal users based on the first degree of anomalousness, further filters out the true target abnormal user from the candidate abnormal users, reducing the possibility of misidentification and further improving the accuracy of abnormal user identification.

[0054] In one embodiment, step 100 above may specifically be performed as follows:

[0055] S110, based on the business information corresponding to each of the multiple users and the first identification model, identify at least one first user among the multiple users who has abnormal Internet access behavior, the first identification model is trained based on the business information of different users;

[0056] S120, based on the business information and the second identification model corresponding to each of the multiple users, at least one second user is identified from the multiple users. The second identification model is used to identify whether the user is discrete in the cluster.

[0057] S130, at least one first user and at least one second user are identified as candidate abnormal users.

[0058] Optionally, in one specific implementation of this application, a first user set A can be obtained by constructing a regression model (and a first identification model) to profile the acquired user data (and the business data of each user); at the same time, based on the isolated forest, a second user set B that is detached from the whole in the high-dimensional feature space is mined from the acquired user data, and A∪B is taken to obtain candidate abnormal users.

[0059] In these alternative embodiments, by using a first identification model to identify users based on their business information, users who may have abnormal internet behavior can be effectively identified, thereby quickly locating potential abnormal users and improving the efficiency and accuracy of abnormal user identification.

[0060] Using a second identification model to identify whether users are discrete within a cluster helps to further confirm abnormal users, avoid missed detections and false judgments, and improve the comprehensiveness and reliability of anomaly detection.

[0061] Finally, the first and second users were identified as candidate anomalous users. By comprehensively considering the identification results of the two models, potential anomalous users were effectively screened out, providing an important basis for further analysis and processing, thereby improving the accuracy and practicality of anomalous user detection.

[0062] In one embodiment, step 110 above may specifically be performed as follows:

[0063] S111, input the business information corresponding to the first sub-user into the first identification model, determine at least one abnormal probability of the first sub-user, one abnormal probability corresponds to one abnormal type of abnormal internet behavior, and the first sub-user is any one of multiple users;

[0064] S112, if the sum of all abnormal probabilities in at least one abnormal probability is greater than a first threshold, or if the first probability is greater than a second threshold, the first sub-user is determined as the first user. The first probability is any one of the at least one abnormal probability, and the second threshold is determined according to the abnormal type of the abnormal internet access behavior corresponding to the first probability.

[0065] Optionally, in one specific implementation of this application, the abnormal user samples are first formed by associating the existing collected abnormal users as labels with the business data exemplified in S100 above, with the sample size set to U. Then, 5*U of the abnormal users are extracted from other users as normal user samples to form training data.

[0066] The training data is then input into the regression model for model training and fitting. After the regression model is trained, the business data of each user is input into the trained model for profile prediction to determine the anomaly probability. Let the anomaly probabilities corresponding to the predicted anomaly types A, B, C, D, and F of all users be β1, β2, β3, β4, and β5. If β1 + β2 + β3 + β4 + β5 > h (and the first threshold), or any anomaly probability is greater than k (the value of k is associated with the corresponding anomaly type), then the user is considered a candidate anomaly user.

[0067] In these alternative embodiments, the business information of each user is analyzed using a first identification model to determine its probability of anomaly, thereby identifying users who may exhibit abnormal internet behavior. By setting a threshold, users with a high probability of anomaly can be filtered out and identified as primary users for further abnormal user identification and processing, thus improving the accuracy and efficiency of abnormal user detection.

[0068] In one embodiment, step 120 above may specifically be performed as follows:

[0069] S121, input the business information corresponding to each of the multiple users into the second identification model to determine the degree of outlier of each user relative to the set of multiple users;

[0070] S122, users whose outlier level is greater than the third threshold are identified as the second user.

[0071] Optionally, in one specific implementation of this application, candidate anomalous users can also be distinguished from ordinary users by their characteristics, exhibiting outlier behavior in high-dimensional features. Therefore, high-dimensional outlier sample analysis methods can be used to mine candidate anomalous users. For example, this application uses the Isolation Forest algorithm to mine anomalous users from each user's business data.

[0072] Isolation forests are similar to random forests in that they train each tree using randomly sampled data to ensure that the variance of the constructed forest is large enough, meaning the less similar the trees are to each other, the better. When constructing an isolation forest, two parameters need to be set: the number of trees T and the maximum sample size W for each tree.

[0073] Specifically, T = B + k * log 10N, where B is the number of types of business data, N is the sample size (and number of users), and k is the first coefficient set from 0 to 10, which can be set according to expert experience. W = N / (v*T), where v is the sampling coefficient with a value of 10 to 100, which can be set according to expert experience.

[0074] Specifically, the training steps for a single model (single tree) are as follows:

[0075] Input: Sample X, the highest height L, L = ceiling(log2p), p is the sample size.

[0076] Output: A single tree model

[0077] Step 1: Randomly select W points from the training data as subsamples and put them into the root node of an isolated tree;

[0078] Step 2: Randomly specify a dimension, and randomly generate a cut point p within the current node's data range. The cut point is generated between the maximum and minimum values ​​of the specified dimension in the current node's data.

[0079] Step 3: The selection of this cutting point generates a hyperplane that divides the current node's data space into two subspaces. Specifically, points less than p in the currently selected dimension are placed in the left branch of the current node, and points greater than or equal to p are placed in the right branch of the current node.

[0080] Step 4: Recursively repeat steps 2 and 3 on the left and right branches of the node, continuously constructing new leaf nodes until the leaf node has only one data (cannot be cut further) or the tree has grown to the set height L.

[0081] The training steps for an isolated forest model (multiple trees) are as follows:

[0082] Input: Sample X, T trees, number of samples p

[0083] Output: Isolated Forest

[0084] Step 1: Initialize forest

[0085] Step 2: Set the maximum level of the tree L = ceiling(log2p)

[0086] Step 3: Loop from 1 to T, randomly sample Y = sample(X, p) from X.

[0087] Let forest = forest∪tree(Y,L)

[0088] Then, outliers for each user were calculated:

[0089]

[0090] Where s(x,n) is the outlier of user x in user set n, c(n) is a dataset containing n samples, and E(h(x)) is the expected value of the path length of x in multiple trees.

[0091] In these alternative embodiments, a second identification model is used to evaluate the outlier degree of each user relative to the overall user set, thereby identifying users who may be out of the cluster. By setting a third threshold, users with a high degree of outlier can be filtered out and identified as second users, which helps to discover anomalous users whose behavior may differ significantly from normal users, improving the comprehensiveness and accuracy of candidate anomalous user detection.

[0092] In one embodiment, prior to step 200 above, the method may further perform the following steps:

[0093] S201, based on the identity identifier corresponding to the candidate abnormal user, obtain multiple first initial data packets corresponding to the business data streams of different content types of the candidate abnormal user, and one first initial data packet belongs to one content type;

[0094] S202, according to the communication link of the candidate abnormal user, the multiple first initial data packets are parsed to obtain the data packet information corresponding to each first initial data packet. The data packet information includes the data packet transmission path, data packet transmission protocol and data packet sequence number.

[0095] S203, determine the content type of the business data stream corresponding to each first initial data packet based on the data packet transmission path corresponding to each first initial data packet;

[0096] S204, For multiple second initial data packets belonging to the same content type, the multiple second initial data packets are reorganized and sorted according to the data packet sequence number corresponding to each second initial data packet to obtain a reorganized data set, wherein the multiple first initial data packets include multiple second initial data packets;

[0097] S205, according to the data packet transmission protocol corresponding to each second initial data packet in the reassembled data set, perform content parsing on each second initial data packet in the reassembled data set to obtain the business data stream corresponding to the reassembled data set.

[0098] Optionally, in one specific implementation of this application, in order to improve analysis efficiency and accuracy while ensuring the business security of normal users, this application first identifies candidate abnormal users and then further analyzes the business data flow of these candidate abnormal users.

[0099] Specifically, firstly, a splitter is used to collect optical data from the User Plane Function (UPF), and mirror clones are used for voice, video service data streams and data internet access services of the IP Multimedia Subsystem (IMS). Then, an aggregation and splitting device is used to aggregate the optical mirror data from multiple UPFs and perform preliminary data collection and processing.

[0100] Subsequently, based on the protocol specification, reverse parsing of the protocol is first performed from the Physical Layer (L1 layer) to the Protocol Data Unit Layer (PDU layer) to obtain user PDU layer information and application layer data; further application layer analysis, stream reconstruction and decryption operations are performed to restore the original data.

[0101] The specific data stream restoration process is as follows:

[0102] ① Protocol Reverse Analysis

[0103] The data packets are parsed step-by-step from the L1, Data Link Layer (L2), Network Layer (Internet Protocol, IP), User Datagram Protocol (UDP), GPRS Tunneling Protocol-User Plane (GTP-U layer), and PDU layer to obtain PDU layer user information and application layer data. The L1 layer is the physical layer and generally uses fiber optic transmission; it does not include protocol processing.

[0104] The Layer 1 (L1) is the physical layer, responsible for transmitting the raw bit stream, primarily focusing on the physical characteristics of the data on the transmission medium, such as voltage and frequency. Layer 2 provides direct communication between nodes, locating devices via physical addresses, transmitting and receiving data frames, and handling error detection and correction. IP is responsible for addressing and routing within the network, identifying different hosts and networks through IP addresses to enable data packet transmission. UDP:2152 provides a simple, connectionless, transaction-oriented transport service, primarily used for efficient data transmission, but does not guarantee data reliability or order. The GTP-U layer is responsible for encapsulating, forwarding, and decapsulating mobile data packets on the user plane. The PDU layer is responsible for encapsulating and decapsulating upper-layer protocol data units, enabling data exchange and communication between different layers.

[0105] a. Parse and extract the MAC address and multi-layer VLAN information of the data link layer (L2 layer).

[0106] b. Parse and extract the Internet Protocol Version 4 (IPv4) or Internet Protocol Version 6 (IPv6) information from the New Radio Base Station (gNB) N3 IP to the UPF N3 IP network layer. Here, gNB N3 IP refers to the IP network interface used to connect the gNB and AMF in the communication network. UPF N3 IP refers to the IP network interface used to connect the UPF and SMF in the communication network.

[0107] c. Parse and extract the transport layer information from gNB UDP:2152 to UPF UDP:2152.

[0108] d. Parse and extract the Tunnel Endpoint Identifier (TEID) from the gNB-side GTP-U Access Network (AN) to the UPF-side GTP-U CN TEID information. The TEID is an identifier used in the General Packet Radio Service (GPRS) tunneling protocol to uniquely identify the tunnel endpoint. In mobile communication networks, the TEID is used to identify the tunnel for data packets between different nodes.

[0109] e. Parse and extract PDU layer information, including but not limited to: user number, PDU session ID, session type, roaming status information, user equipment (UE) IP information, policy control function (PCF) information, quality of service (QoS) information, tunnel information, destination address, SMF identifier, slice information, AMF information, user location information, session management information, and other related information.

[0110] f. Remove the data headers from layers 1 to 5 to obtain the application layer data (i.e., data packet information).

[0111] ② Application layer analysis, stream reconstruction and decryption

[0112] a. Application Layer Analysis: First, identify the protocol by examining the protocol header information of the data packet, such as the port number and protocol identifier; then extract the application layer data. Different application layer protocols have different data formats and structures, so data extraction and parsing need to be performed according to the protocol specifications.

[0113] b. Stream reassembly: By using information such as the source IP address, destination IP address, source port, and destination port of the data packets, the data packets are grouped to determine that the data packets belong to the same data stream; the acquired data packets are sorted according to the sequence number, timestamp, or other unique identifier in the data packets, and the data packets are reassembled in the correct order; the sequence number and acknowledgment number of the data packets are checked to ensure that the data packets are reassembled in the correct order and that there are no lost or duplicate data packets.

[0114] c. Decryption: Decrypt the encrypted data packets by using a suitable key or algorithm to reverse the encrypted data and restore it to the original plaintext data.

[0115] In these optional embodiments, the user plane protocol functions are acquired through optical splitting, enabling the acquisition of multiple service data streams, including IMS voice, video, and data internet access service data streams, thus covering various user behavior types and improving the comprehensiveness of the analysis. Secondly, protocol reverse parsing is performed based on the protocol specifications, parsing data packets step-by-step from the physical layer to the PDU layer, obtaining rich user information and application layer data, including user numbers, session information, and policy control function information, which is beneficial for in-depth analysis of user behavior. Finally, application layer analysis, stream reconstruction, and decryption operations effectively restore the original data, ensuring data integrity and accuracy, and providing a reliable data foundation for subsequent identification of abnormal users. The beneficial effects of these steps include improved analysis efficiency and accuracy, while ensuring the service security of normal users.

[0116] In one embodiment, the semantic recognition model includes a text recognition model and an image recognition model; step 200 above can specifically be performed as follows:

[0117] S210, when the data type of the online multimedia data is an image, input the online multimedia data into the image recognition model to obtain the first degree of anomaly corresponding to the online multimedia data;

[0118] S220: When the data type of the online multimedia data is voice or text, the online multimedia data is input into the text recognition model to obtain the first anomaly level corresponding to the online multimedia data.

[0119] Optionally, in one specific implementation of this application, semantic recognition for data types such as text, images, videos, and speech requires pre-training two large models: a text semantic recognition model and an image semantic recognition model. Speech semantic recognition first converts speech to text using a speech-to-text model, and then uses the text semantic recognition model for recognition. Semantic recognition of video data types is achieved through an image semantic recognition model (extracting video images) and a text semantic recognition model (first converting speech to text).

[0120] Text semantic recognition models can be pre-trained on large-scale unlabeled text corpora to obtain text encoders with strong expressive power and generalization capabilities. These encoders can be used for natural language understanding tasks (such as text classification, reading comprehension, and information extraction) as well as text generation tasks (machine translation, text summarization, and dialogue generation).

[0121] The specific training process of the text semantic recognition model is as follows:

[0122] ① Unsupervised pre-training

[0123] First, unsupervised pre-training is performed using a large-scale unlabeled text corpus. Specifically, for the unsupervised text sequence U = {u1,...,un}, with a context window size of k, a language model with multiple TransformerDecoder blocks is used to maximize the probability of the sequence:

[0124]

[0125] The conditional probability P is modeled using a neural network with parameters θ, which is learned through stochastic gradient descent. The language model is detailed below:

[0126] h0=UW e +W p

[0127]

[0128]

[0129] Among them, W e W is a word nesting matrix of a text sequence. p The mean value is the nested position value, n is the number of TransformerDecoder blocks used, transformer_block refers to the Transformer Decoder block, and softmax is the normalized exponential function.

[0130] ②Supervised learning adjustment

[0131] The second-stage model is a supervised model designed for a specific task of semantic recognition and classification of various types of data, including different anomaly types and data from normal users. First, the input data (and various types of online multimedia data) is transformed into an ordered sequence that the pre-trained model can process. Then, a linear network layer is added after the pre-trained model for supervised training. Assuming all datasets are C, and the input sequence is x = (x1, ..., xm), with label y, the model's task is to predict the label and maximize the following probabilities:

[0132]

[0133] in W is the output obtained after the first stage of pre-training the model. y This is the weight matrix of the linear layer. Maximize the following objective function:

[0134]

[0135] The two-stage models are integrated and trained to obtain a text semantic recognition model:

[0136] L3(C) = L2(C) + γL1(C)

[0137] Here, γ is a regularization parameter.

[0138] Optionally, in another embodiment of this application, the image semantic recognition model is implemented based on a Contrastive Language-Image Pretraining (CLIP) model. CLIP is designed to enable the model to understand the relationship between text and images, achieving cross-modal semantic understanding by aligning data from different modalities. CLIP uses contrastive learning to train the model. Contrastive learning learns feature representations by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. In CLIP, positive sample pairs are text-image pairings from the same modality, while negative sample pairs are text-image combinations from different modalities.

[0139] The specific training process of the image semantic recognition model is as follows:

[0140] Contrastive Learning Pre-training: CLIP forms training samples by constructing text-image pairs. For each text description, it forms a positive sample pair with positive sample images of the same category, and a negative sample pair with images of other categories. Through this construction, the model is required to bring text-image pairs belonging to the same category closer together in the embedding space, while pushing sample pairs from different categories further apart.

[0141] 1. Data Collection and Labeling: Collect datasets related to different anomaly types and normal users, and label the data. The datasets should include images and corresponding labels or category information.

[0142] 2. Data Processing and Encoding: The image data is preprocessed, including scaling, cropping, and normalization. Then, the processed data is encoded into an embedding vector.

[0143] 3. Construct a task-specific contrastive learning objective function: During the pre-training phase, construct a contrastive learning objective function based on different anomaly types and the classification task for normal users. The objective function should be designed to maximize the similarity of sample pairs within the same class while minimizing the similarity of sample pairs between different classes.

[0144] 4. Task-Specific Pre-training: CLIP is pre-trained using a dataset specific to the recognition task and a constructed contrastive learning objective function. During pre-training, the model is optimized on a large number of text-image pairs, learning semantic representations for the task, ultimately resulting in an image semantic recognition model.

[0145] In these alternative embodiments, semantic understanding of different data types (text, images, video, and speech) is achieved by pre-training large-scale text semantic recognition models and image semantic recognition models, and by converting speech data to text before performing text semantic recognition. Through pre-training and contrastive learning on large-scale unlabeled text corpora and text-image pairs, the model can understand and extract semantic information between data, thereby improving the accuracy and generalization ability of semantic recognition. This approach also ensures the business security of legitimate users while effectively identifying abnormal users, thus improving analysis efficiency and accuracy.

[0146] In one embodiment, abnormal internet access behavior includes multiple abnormal types, and the service data stream includes multiple internet multimedia data; the above step 200 can specifically perform the following steps:

[0147] S230, input multiple Internet multimedia data into the semantic recognition model to obtain multiple first anomaly sub-probabilities corresponding to each of the multiple Internet multimedia data, and each first anomaly sub-probability belongs to an anomaly type and an Internet multimedia data.

[0148] S240, obtain at least one second anomaly subprobability belonging to the first type, where the first type is any one of multiple anomaly types;

[0149] S250, determine the maximum value among at least one second anomaly subprobability as the first anomaly degree of the candidate anomaly user under the first type.

[0150] Optionally, in one specific implementation of this application, the obtained information data stream is processed to retain only four types of data (text, voice, video, and images, as well as online multimedia data), and the sender's mobile phone number is labeled for each data. Then, the text semantic recognition model and image semantic recognition model obtained from the previous training are used to identify the content of online multimedia data of different abnormal types. For a certain type of abnormal problem, the result with the highest tendency ratio among all identified content is taken to obtain the tendency ratio (and first degree of abnormality) of the abnormal online behavior of the user for each abnormal type, which are γ1, γ2, γ3, γ4, and γ5, respectively. Evidence is preserved for online multimedia data identified as the above-mentioned abnormal online information.

[0151] In these optional embodiments, by identifying multiple abnormal types of online behavior and processing the business data stream, only four types of data—text, voice, video, and images—are retained, and the sender's mobile phone number is labeled, thereby achieving more accurate identification and location of abnormal behavior. By using text semantic recognition models and image semantic recognition models to identify the content of different types of online multimedia data, the characteristics of abnormal behavior can be captured more accurately, further improving the detection efficiency and accuracy of abnormal behavior. Simultaneously, preserving evidence of data identified as abnormal online information helps with subsequent review and processing, thereby better protecting network security and user rights.

[0152] In one embodiment, abnormal internet access behavior includes multiple abnormal types; step 200 above can specifically perform the following steps:

[0153] S260, input the Internet multimedia data corresponding to the candidate abnormal user into the semantic recognition model, determine multiple first abnormality levels corresponding to the candidate abnormal user, and one first abnormality level corresponds to one abnormality type;

[0154] In one embodiment, step 300 above may specifically be performed as follows:

[0155] S310, input the business information corresponding to the candidate abnormal user into the first identification model, determine multiple second abnormality levels corresponding to the candidate abnormal user, and one second abnormality level corresponds to one abnormality type;

[0156] S320, obtain the first degree sum between the first degree of the second type corresponding to the first degree of abnormality and the second degree of abnormality corresponding to the second type, where the second type is any one of multiple abnormal types;

[0157] S330, if the first degree is greater than the fourth threshold, the candidate abnormal user is identified as the target abnormal user, and the fourth threshold is determined according to the second type.

[0158] Optionally, in one specific implementation of this application, the anomaly probabilities (and second anomaly degrees) β1, β2, β3, β4, β5 corresponding to different anomaly types of candidate abnormal users obtained according to the first recognition model are summed with the probabilities γ1, γ2, γ3, γ4, γ5 corresponding to different anomaly types obtained through the text semantic recognition model and image semantic recognition model in steps S230-S250, to obtain γ1+β1, γ2+β2, γ3+β3, γ4+β4, γ5+β5. The overall summation value θ is then obtained.

[0159]

[0160] Subsequently, this application can set different judgment thresholds according to the different requirements corresponding to different anomaly types. When the threshold is exceeded, the candidate anomaly user is considered to be the target anomaly user of the corresponding anomaly type, and an overall evaluation threshold is set. If the overall summation value is greater than the overall evaluation threshold, the candidate anomaly user can be considered to be the target anomaly user.

[0161] In these alternative embodiments, by combining the output of the semantic recognition model with the analysis of business information, a multi-dimensional assessment of the anomaly severity of candidate abnormal users is performed, effectively improving the accuracy and precision of anomaly detection. By setting different judgment thresholds according to different anomaly types, different types of abnormal behavior can be identified and processed in a targeted manner, thereby better ensuring network security. At the same time, setting an overall evaluation threshold can comprehensively consider the anomaly severity of multiple anomaly types, ensuring timely detection and processing of overall abnormal behavior, further improving the overall efficiency and response capability of the system.

[0162] In one embodiment, step 300 above may specifically be performed as follows:

[0163] S340, the first degree is determined as the target anomaly subprobability of the candidate anomalous user under the second type;

[0164] S350, obtain the sum of the target anomaly subprobabilities corresponding to each anomaly type for candidate anomaly users;

[0165] S360: If the sum of the target anomaly subprobabilities corresponding to each anomaly type of the candidate anomaly user is greater than the fifth threshold, the candidate anomaly user is determined as the target anomaly user.

[0166] In these optional embodiments, by comprehensively considering the sum of the target anomaly subprobabilities of candidate anomalous users under each anomaly type, anomalous behavior can be effectively comprehensively evaluated and judged. By setting a fifth threshold, the target anomalous user can be quickly and accurately identified based on the comprehensive evaluation results, thereby enabling timely implementation of corresponding security measures to ensure the normal operation of the network and the information security of users. This helps improve the accuracy and real-time performance of anomaly detection, enhances the system's ability to respond to anomalous behavior, and further improves the level of network security.

[0167] Optionally, in this application embodiment, this application also provides a user identification system. This user identification system is based on the core network native network data analytics function (NWDAF) data analysis network element. This application constructs a user profile model based on NWDAF to initially identify candidate abnormal users, and uses parsing methods for text, images, audio, video, and calls of communication network user plane service data to achieve full parsing and restoration of application data. Subsequently, the service data flow is detected based on a pre-trained large model (and semantic recognition model), and finally, the efficiency and accuracy of abnormal user detection are improved through a joint detection method.

[0168] like Figure 2 As shown, this application identifies anomalous users based on a combination of user profiling and content recognition. NWDAF is responsible for user profiling, while a pre-trained large-scale model (and semantic recognition model) is responsible for user content recognition. NWDAF obtains native network data (and service data) from network elements such as AMF, SMF, and OAM to create user profiles. Based on the pre-trained large-scale model, it performs content recognition on the user's voice and video service data in the IMS domain, as well as data internet access services (and internet multimedia data), thereby achieving joint detection. The system architecture of this application includes wireless side, IMS domain, and 5G core network related elements, specifically including wireless base stations, OAM, AMF, SMF, UPF, DN, SBC, a pre-trained large-scale model platform, an anomalous user mining function based on SDN deployment, and a user profiling NWDAF module.

[0169] Specifically, the system architecture of this application includes: an abnormal user detection function, comprising multiple modules, including a data collection, parsing, and restoration module, which, in compliance with relevant regulations, uses techniques such as decryption, stream reconstruction, application layer analysis, and protocol reverse engineering to restore the original data (and business data stream); a large model invocation module, responsible for invoking pre-trained large model semantic recognition models of different types such as text, images, videos, and audio to perform semantic recognition on the above content; and an abnormal user detection function, which only retains the original data of users identified by the pre-trained large model as having abnormal network behavior, and does not retain the original data of any other users.

[0170] The acquisition and processing module consists of optical splitting equipment and aggregation and distribution equipment. It is responsible for mirroring and cloning IMS voice and video service data streams, as well as mirroring and cloning data internet access services. It is also responsible for aggregating the service data streams of multiple UPFs and performing preliminary traffic acquisition.

[0171] The semantic recognition model is used for training, deployment, maintenance, and optimization of large-scale semantic recognition models for text, images, videos, and speech, and provides a model calling interface. This large-scale model platform performs no analysis other than identifying and analyzing abnormal online behavior in users' online multimedia data, and does not retain any user data.

[0172] The data acquisition module is responsible for extracting native communication network data (and service data) required for user profiling from AMF, SMF, and OAM; the data processing module performs multi-dimensional processing and integration of the native data acquired by each network element according to data processing rules; and the data analysis module uses the processed data to build a profiling model, perform data analysis, and output candidate abnormal users.

[0173] In this embodiment, user profiling is first performed to exclude normal users from the network and to identify potential abnormal users exhibiting unusual internet behavior. This reduces a large amount of unnecessary content identification and significantly lowers the processing load. Furthermore, the business data streams of potential candidate abnormal users are traced, and a pre-trained large model is used for identification. Finally, the result of joint detection of user profiling and content matching is output, determining whether the candidate abnormal user is the target abnormal user.

[0174] In these optional embodiments, the abnormal user identification method of this application is deployed on the native network, realizing efficient data collection, processing, and analysis, while also achieving virtualized design, deployment, and management of network functions. Furthermore, while existing methods primarily use user features from the B domain and business usage features from the O domain for mining, this application utilizes not only business data but also raw data for analysis, resulting in a more direct and accurate approach. This application uses pre-trained large models of text and images for identification, leading to superior algorithm performance and overall improved accuracy and precision in abnormal user identification. The joint detection and identification method can further enhance the efficiency and accuracy of abnormal user detection.

[0175] Figure 3 A schematic diagram of the structure of a user identification device provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0176] Reference Figure 3 User identification may include:

[0177] The first determining module 301 is used to identify multiple users based on their respective business information to obtain candidate abnormal users. The candidate abnormal users are users among the multiple users who have abnormal internet access behavior.

[0178] The second determining module 302 is used to input the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's Internet access behavior. The business data stream includes the user's Internet multimedia data.

[0179] The third determination module 303 is used to determine whether a candidate abnormal user is a target abnormal user based on the first degree of abnormality.

[0180] In one embodiment, the first determining module 301 may include:

[0181] The first identification submodule is used to identify at least one first user among multiple users who has abnormal internet access behavior based on the business information corresponding to each of the multiple users and the first identification model. The first identification model is trained based on the business information of different users.

[0182] The second identification submodule is used to identify at least one second user from multiple users based on the business information and the second identification model corresponding to each user. The second identification model is used to identify whether the user is discrete in the cluster.

[0183] The first determination submodule is used to determine at least one first user and at least one second user as candidate abnormal users.

[0184] In one embodiment, the first identification submodule may include:

[0185] The first determining unit is used to input the business information corresponding to the first sub-user into the first identification model and determine at least one abnormal probability of the first sub-user. An abnormal probability corresponds to an abnormal type of abnormal internet access behavior. The first sub-user is any one of multiple users.

[0186] The second determining unit is used to determine the first sub-user as the first user when the sum of all abnormal probabilities in at least one abnormal probability is greater than a first threshold, or the first probability is greater than a second threshold. The first probability is any one of the at least one abnormal probability, and the second threshold is determined according to the abnormal type of the abnormal internet access behavior corresponding to the first probability.

[0187] In one embodiment, the second identification submodule may include:

[0188] The third determining unit is used to input the business information corresponding to each of the multiple users into the second identification model to determine the degree of outlier of each user relative to the set of multiple users.

[0189] The fourth determining unit is used to identify users whose outlier level is greater than the third threshold as the second user.

[0190] In one embodiment, the user identification device may further include:

[0191] The acquisition module is used to acquire multiple first initial data packets corresponding to different content types of business data streams of candidate abnormal users based on the identity identifiers of the candidate abnormal users. Each first initial data packet belongs to one content type.

[0192] The first parsing module is used to parse multiple first initial data packets according to the communication links of the candidate abnormal users, and obtain the data packet information corresponding to each first initial data packet. The data packet information includes the data packet transmission path, data packet transmission protocol and data packet sequence number.

[0193] The fourth determining module is used to determine the content type of the business data stream corresponding to each first initial data packet based on the data packet transmission path corresponding to each first initial data packet.

[0194] The sorting module is used to reorganize and sort multiple second initial data packets belonging to the same content type according to the data packet sequence number corresponding to each second initial data packet, so as to obtain a reorganized data set, wherein the multiple first initial data packets include multiple second initial data packets;

[0195] The second parsing module is used to parse the content of each second initial data packet in the reassembled data set according to the data packet transmission protocol corresponding to each second initial data packet in the reassembled data set, so as to obtain the business data stream corresponding to the reassembled data set.

[0196] In one embodiment, the semantic recognition model includes a text recognition model and an image recognition model; the second determining module 302 may include:

[0197] The second determination submodule is used to input the online multimedia data into the image recognition model when the data type of the online multimedia data is an image, and obtain the first degree of anomaly corresponding to the online multimedia data.

[0198] The third determination submodule is used to input the online multimedia data into the text recognition model when the data type of the online multimedia data is voice or text, and obtain the first degree of anomaly corresponding to the online multimedia data.

[0199] In one embodiment, abnormal internet access behavior includes multiple abnormal types, and the service data stream includes multiple internet multimedia data; the second determining module 302 may further include:

[0200] The fourth determination submodule is used to input multiple Internet multimedia data into the semantic recognition model to obtain multiple first anomaly subprobabilities corresponding to each of the multiple Internet multimedia data. Each first anomaly subprobability belongs to an anomaly type and an Internet multimedia data.

[0201] The first acquisition submodule is used to acquire at least one second anomaly subprobability belonging to the first type, where the first type is any one of multiple anomaly types.

[0202] The fifth determination submodule is used to determine the maximum value among at least one second anomaly subprobability as the first anomaly degree of the candidate anomaly user under the first type.

[0203] In one embodiment, abnormal internet access behavior includes multiple abnormal types; the second determining module 302 may further include:

[0204] The sixth determination submodule is used to input the online multimedia data corresponding to the candidate abnormal user into the semantic recognition model to determine multiple first abnormality levels corresponding to the candidate abnormal user, and each first abnormality level corresponds to an abnormality type.

[0205] In one embodiment, the third determining module 303 may include:

[0206] The seventh determination submodule is used to input the business information corresponding to the candidate abnormal user into the first identification model to determine multiple second abnormality levels corresponding to the candidate abnormal user, and one second abnormality level corresponds to one abnormality type.

[0207] The second acquisition submodule is used to acquire the first degree sum between the first degree of the second type and the second degree of the second type, wherein the second type is any one of multiple abnormal types.

[0208] The eighth determination submodule is used to determine the candidate abnormal user as the target abnormal user when the first degree is greater than the fourth threshold, and the fourth threshold is determined according to the second type.

[0209] In one embodiment, the third determining module 303 may further include:

[0210] The ninth determination submodule is used to determine the first degree and the target anomaly subprobability of the candidate anomaly user under the second type;

[0211] The third acquisition submodule is used to obtain the sum of the target anomaly subprobabilities corresponding to each anomaly type for candidate anomaly users;

[0212] The tenth determination submodule is used to determine the candidate abnormal user as the target abnormal user if the sum of the target abnormal subprobabilities corresponding to each abnormality type of the candidate abnormal user is greater than the fifth threshold.

[0213] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. They are devices corresponding to the above-mentioned battery thermal runaway early warning method. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of this device. For details on its specific functions and the technical effects it brings, please refer to the method embodiment section. It will not be repeated here.

[0214] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0215] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0216] The device may include a processor 401 and a memory 402 storing program instructions.

[0217] When processor 401 executes the program, it implements the steps in any of the above method embodiments.

[0218] For example, the program can be divided into one or more modules / units, one or more of which are stored in memory 402 and executed by processor 401 to complete this application. The one or more modules / units can be a series of program instruction segments capable of performing a specific function, which describe the execution process of the program in the device.

[0219] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0220] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.

[0221] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0222] The processor 401 implements any of the methods described in the above embodiments by reading and executing program instructions stored in the memory 402.

[0223] In one example, the electronic device may also include a communication interface 403 and a bus 410. The processor 401, memory 402, and communication interface 403 are connected via the bus 410 and communicate with each other.

[0224] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0225] Bus 410 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0226] Furthermore, in conjunction with the methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores program instructions; when these program instructions are executed by a processor, they implement any of the methods in the above embodiments.

[0227] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0228] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0229] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0230] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0231] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on machine-readable media or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.

[0232] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0233] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0234] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A user identification method, characterized in that, The method includes: Based on the business information corresponding to each of the multiple users and a first identification model, at least one first user with abnormal internet access behavior is identified among the multiple users. The first identification model is trained based on the business information of different users. Based on the business information corresponding to each of the multiple users and a second identification model, at least one second user is identified from the multiple users. The second identification model is used to identify whether the user is discrete in the cluster. The at least one first user and the at least one second user are determined as candidate abnormal users. The candidate abnormal users are users with abnormal internet access behavior among the multiple users. The business information is collected based on authentication key management function, session management function, and operation management and maintenance system. The first identification model is a regression model, and the second identification model is a model built based on the isolated forest algorithm. Based on the identity identifiers corresponding to the candidate abnormal users, a splitter is used to perform optical acquisition on the User Plane Protocol Function (UPF). From the data obtained by optical acquisition, optical mirror data of multiple UPFs is obtained by mirroring and cloning the data streams of the IP Multimedia Subsystem (IMS) voice, video services, and data access services. The optical mirror data of multiple UPFs is aggregated and preliminarily acquired using a convergence and splitting device to obtain processed optical mirror data. Based on the protocol specifications, the processed optical mirror data is reverse-parsed from the physical layer to the protocol data unit to obtain multiple first initial data packets corresponding to the service data streams of different content types of the candidate abnormal users. Each first initial data packet belongs to one content type. According to the communication link of the candidate abnormal user, the plurality of first initial data packets are parsed to obtain data packet information corresponding to each first initial data packet. The data packet information includes data packet transmission path, data packet transmission protocol and data packet sequence number. Based on the data packet transmission path corresponding to each first initial data packet, determine the content type of the business data stream corresponding to each first initial data packet; For multiple second initial data packets belonging to the same content type, the multiple second initial data packets are reorganized and sorted according to the data packet sequence number corresponding to each second initial data packet to obtain a reorganized data set, wherein the multiple first initial data packets include the multiple second initial data packets; Based on the data packet transmission protocol corresponding to each second initial data packet in the reconstructed data set, the content of each second initial data packet in the reconstructed data set is parsed to obtain the business data stream corresponding to the reconstructed data set; The business data stream corresponding to the candidate abnormal user is input into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's online behavior. The business data stream includes the user's online multimedia data. Based on the first degree of abnormality, determine whether the candidate abnormal user is the target abnormal user.

2. The method according to claim 1, characterized in that, The step of identifying at least one first user among the multiple users who exhibits abnormal internet access behavior based on the business information corresponding to each of the multiple users and the first identification model includes: Input the business information corresponding to the first sub-user into the first identification model to determine at least one abnormal probability of the first sub-user. An abnormal probability corresponds to an abnormal type of abnormal online behavior. The first sub-user is any one of the multiple users. If the sum of all the abnormal probabilities in the at least one abnormal probability is greater than a first threshold, or if the first probability is greater than a second threshold, the first sub-user is identified as the first user. The first probability is any one of the at least one abnormal probabilities, and the second threshold is determined based on the abnormal type of the abnormal internet behavior corresponding to the first probability.

3. The method according to claim 1, characterized in that, The step of identifying at least one second user from the plurality of users based on the business information and the second identification model corresponding to each of the plurality of users includes: The business information corresponding to each of the multiple users is input into the second identification model to determine the degree of outlier of each user relative to the set of the multiple users; Users whose outlier level is greater than the third threshold are identified as the second user.

4. The method according to claim 1, characterized in that, The semantic recognition model includes a text recognition model and an image recognition model; The step of inputting the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's online behavior includes: When the data type of the online multimedia data is an image, the online multimedia data is input into the image recognition model to obtain the first degree of anomaly corresponding to the online multimedia data; When the data type of the online multimedia data is voice or text, the online multimedia data is input into the text recognition model to obtain the first anomaly level corresponding to the online multimedia data.

5. The method according to claim 1, characterized in that, The abnormal internet access behavior includes multiple abnormal types, and the service data stream includes multiple internet multimedia data. The step of inputting the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's online behavior further includes: The multiple online multimedia data are input into the semantic recognition model to obtain multiple first anomaly sub-probabilities corresponding to each of the multiple online multimedia data. Each first anomaly sub-probability belongs to an anomaly type and an online multimedia data. Obtain at least one second anomaly subprobability belonging to the first type, wherein the first type is any one of the plurality of anomaly types; The maximum value among the at least one second anomaly sub-probability is determined as the first anomaly degree of the candidate anomaly user under the first type.

6. The method according to claim 1, characterized in that, The abnormal internet access behavior includes several abnormal types; The step of inputting the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's online behavior further includes: The online multimedia data corresponding to the candidate abnormal user is input into the semantic recognition model to determine multiple first abnormality levels corresponding to the candidate abnormal user, and each first abnormality level corresponds to an abnormality type. The step of determining whether the candidate abnormal user is the target abnormal user based on the first degree of abnormality includes: The business information corresponding to the candidate abnormal user is input into the first identification model to determine multiple second abnormality levels corresponding to the candidate abnormal user, and each second abnormality level corresponds to an abnormality type. Obtain the first degree sum between the first degree of abnormality corresponding to the second type and the second degree of abnormality corresponding to the second type, wherein the second type is any one of the plurality of abnormality types; If the first degree is greater than the fourth threshold, the candidate abnormal user is identified as the target abnormal user, and the fourth threshold is determined according to the second type.

7. The method according to claim 6, characterized in that, The step of determining whether the candidate abnormal user is the target abnormal user based on the first degree of abnormality further includes: The first degree is determined as the target anomaly subprobability of the candidate anomalous user under the second type; Obtain the sum of the target anomaly subprobabilities corresponding to each anomaly type for the candidate anomaly users; If the sum of the target anomaly subprobabilities corresponding to each anomaly type of the candidate anomaly user is greater than the fifth threshold, the candidate anomaly user is determined as the target anomaly user.

8. A user identification device, characterized in that, The device includes: The first determining module is used to identify at least one first user among the multiple users who exhibit abnormal internet access behavior based on the business information corresponding to each of the multiple users and a first identification model, wherein the first identification model is trained based on the business information of different users; and to identify at least one second user among the multiple users based on the business information corresponding to each of the multiple users and a second identification model, wherein the second identification model is used to identify whether the user is discrete in the cluster; and to determine the at least one first user and the at least one second user as candidate abnormal users, wherein the candidate abnormal users are users among the multiple users who exhibit abnormal internet access behavior, wherein the business information is collected based on authentication key management function, session management function, and operation management and maintenance system, the first identification model is a regression model, and the second identification model is a model built based on the isolated forest algorithm; The acquisition module is used to perform optical acquisition of the User Plane Protocol Function (UPF) using a splitter based on the identity identifier corresponding to the candidate abnormal user. From the data obtained by optical acquisition, it obtains optical mirror data of multiple UPFs for the IP Multimedia Subsystem (IMS) voice, video service data streams, and data internet access service mirror clones. It then uses a convergence and splitting device to converge the optical mirror data of multiple UPFs and perform preliminary acquisition processing to obtain processed optical mirror data. Based on the protocol specification, it performs protocol reverse parsing from the physical layer to the protocol data unit to obtain multiple first initial data packets corresponding to the service data streams of different content types of the candidate abnormal user. Each first initial data packet belongs to one content type. The first parsing module is used to parse the plurality of first initial data packets according to the communication links of the candidate abnormal users, and obtain data packet information corresponding to each first initial data packet. The data packet information includes data packet transmission path, data packet transmission protocol and data packet sequence number. The fourth determining module is used to determine the content type of the business data stream corresponding to each first initial data packet based on the data packet transmission path corresponding to each first initial data packet. The sorting module is used to reorganize and sort multiple second initial data packets belonging to the same content type according to the data packet sequence number corresponding to each second initial data packet, so as to obtain a reorganized data set, wherein the multiple first initial data packets include multiple second initial data packets; The second parsing module is used to parse the content of each second initial data packet in the reassembled data set according to the data packet transmission protocol corresponding to each second initial data packet in the reassembled data set, so as to obtain the business data stream corresponding to the reassembled data set. The second determining module is used to input the business data stream corresponding to the candidate abnormal user into the semantic recognition model to obtain the first degree of abnormality of the candidate abnormal user's online behavior, wherein the business data stream includes the user's online multimedia data; The third determining module is used to determine whether the candidate abnormal user is the target abnormal user based on the first degree of abnormality.

9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the user identification method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the user identification method as described in any one of claims 1-7.

11. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device causes the electronic device to perform the user identification method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Abnormal user identification method and device, electronic equipment and storage medium

    CN111641608A

  • Abnormal object identification method and device, equipment and storage medium

    CN117649234A