Data processing method and device, computer equipment and computer readable storage medium
By collecting and generating content information and statistical information from multiple data attribute dimensions and combining it with time behavior series, the problem of low accuracy in abnormal account detection in existing technologies is solved, and the precise screening and identification of abnormal accounts in Internet applications is achieved.
Patent Information
- Application Number
- CN202410280077.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-12
AI Technical Summary
The data processing methods used in existing technologies to detect abnormal accounts in Internet applications have low accuracy and are difficult to identify accounts that exploit system vulnerabilities to conduct malicious activities.
By collecting Internet data from multiple data attribute dimensions, generating content information and statistical information, conducting the first anomaly screening, and conducting the second screening based on time behavior series, combining preset strategy rules and neural network models to identify abnormal accounts.
It improves the accuracy of abnormal account detection, avoids data information omission, and enables detailed screening of abnormal accounts.
Smart Images

Figure CN120639320A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technology, and in particular to a data processing method, apparatus, computer equipment, and computer-readable storage medium. Background Art
[0002] Currently, in many internet applications, some accounts exploit vulnerabilities in application systems or business rules to engage in harmful business and dissemination activities, such as disseminating inappropriate content, engaging in abnormal resource exchange, and obtaining benefits at extremely low costs. Therefore, in order to protect the security of the network environment and the interests of other accounts, it is extremely important to identify abnormal accounts in applications through data processing methods.
[0003] However, the data processing methods used in related technologies to detect abnormal accounts in Internet applications have low accuracy. Summary of the Invention
[0004] The embodiments of the present disclosure provide a data processing method, apparatus, computer device, and computer-readable storage medium. The data processing method can improve the accuracy of abnormal account detection in applications.
[0005] A first aspect of the present disclosure provides a data processing method, the method comprising:
[0006] Acquire internet data to be detected associated with a target scenario in a target internet application, wherein the internet data to be detected includes internet data collected from multiple data attribute dimensions;
[0007] Determining a target data attribute dimension from the multiple data attribute dimensions based on the target scenario, and generating content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension;
[0008] Performing anomaly screening on each to-be-detected Internet data according to the content information and the statistical information, and determining a plurality of first abnormal accounts in the target Internet application according to the anomaly screening results;
[0009] Obtaining a time behavior sequence corresponding to each of the first abnormal accounts;
[0010] A behavior sequence comparison is performed on the multiple first abnormal accounts based on the time behavior sequence, and multiple second abnormal accounts are determined from the multiple first abnormal accounts according to the comparison result.
[0011] A second aspect of the present disclosure provides a data processing device, the device comprising:
[0012] A first acquisition unit is configured to acquire to-be-detected Internet data associated with a target scenario in a target Internet application, wherein the to-be-detected Internet data includes Internet data collected from multiple data attribute dimensions;
[0013] a generating unit, configured to determine a target data attribute dimension from the plurality of data attribute dimensions based on the target scenario, and generate content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension;
[0014] a screening unit, configured to perform anomaly screening on each to-be-detected Internet data based on the content information and the statistical information, and determine a plurality of first abnormal accounts in the target Internet application based on the anomaly screening results;
[0015] a second acquiring unit, configured to acquire a time behavior sequence corresponding to each of the first abnormal accounts;
[0016] The comparison unit is configured to perform a behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence, and determine a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result.
[0017] Optionally, in some embodiments, the screening unit is specifically configured to:
[0018] Performing anomaly matching on the statistical information according to preset strategy rules to obtain a first screening result;
[0019] Performing abnormality classification on the content information based on a preset neural network model to obtain a second screening result;
[0020] determining a plurality of abnormal Internet data from the plurality of Internet data to be detected according to the first screening result and the second screening result;
[0021] A plurality of first abnormal accounts in the target Internet application is determined according to the plurality of abnormal Internet data.
[0022] Optionally, in some embodiments, the screening unit is specifically configured to:
[0023] Obtaining a first weight coefficient corresponding to the first screening result, and obtaining a second weight coefficient corresponding to the second screening result;
[0024] Performing weighted calculation on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result;
[0025] A plurality of abnormal Internet data are determined from the plurality of Internet data to be detected according to the target screening result.
[0026] Optionally, in some embodiments, the screening unit is specifically configured to:
[0027] determining a first score corresponding to the first screening result and a second score corresponding to the second screening result;
[0028] Performing weighted calculation on the first score and the second score based on the first weight coefficient and the second weight coefficient to obtain a target score;
[0029] A target screening result is determined according to the target score.
[0030] Optionally, in some embodiments, the screening unit is specifically configured to:
[0031] Obtaining account information corresponding to each abnormal Internet data and the target Internet application;
[0032] A plurality of first abnormal accounts in the target Internet application is determined according to the account information.
[0033] Optionally, in some embodiments, the comparison unit is specifically configured to:
[0034] Get the reference time behavior sequence;
[0035] Calculating the similarity between the time behavior sequence and the reference time behavior sequence to obtain a comparison result;
[0036] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison result.
[0037] Optionally, in some embodiments, the comparison unit is specifically configured to:
[0038] Extracting features from the time behavior sequence based on a preset sequence feature generation model to obtain a first sequence feature;
[0039] Extracting features from the reference time behavior sequence based on the preset sequence feature generation model to obtain a second sequence feature;
[0040] Calculate the cosine similarity between the first sequence feature and the second sequence feature, and determine the alignment result according to the cosine similarity.
[0041] Optionally, in some embodiments, the data processing device provided by the present application further includes:
[0042] a first constructing unit, configured to construct a spectrum behavior sequence corresponding to each of the first abnormal accounts according to the time behavior sequence;
[0043] The comparison unit is specifically used for:
[0044] A behavior sequence comparison is performed on the multiple first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence, and multiple second abnormal accounts are determined from the multiple first abnormal accounts according to the comparison result.
[0045] Optionally, in some embodiments, the comparison unit is specifically configured to:
[0046] Obtaining a reference time behavior sequence and a reference spectrum behavior sequence;
[0047] Determining a first sub-comparison result based on the similarity between the time behavior sequence and the reference time behavior sequence, and determining a second sub-comparison result based on the similarity between the spectrum behavior sequence and the reference spectrum behavior sequence;
[0048] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the first sub-comparison result and the second sub-comparison result.
[0049] Optionally, in some embodiments, the data processing device provided by the present application further includes:
[0050] a determining unit, configured to determine a plurality of abnormal data attribute nodes corresponding to the plurality of second abnormal accounts, and associated data attribute nodes corresponding to the plurality of abnormal data attribute nodes;
[0051] A second construction unit is configured to construct a graph network using the plurality of second abnormal account numbers, the plurality of abnormal data attribute nodes, and the associated data attribute nodes;
[0052] A partitioning unit is used to partition the graph network into multiple subgraphs based on the closeness of the node relationships in the graph network, and to determine multiple abnormal account sets based on the multiple subgraphs.
[0053] Optionally, in some embodiments, the data processing device provided by the present application further includes:
[0054] A third acquisition unit is used to obtain abnormal feedback information of Internet data in multiple data attribute dimensions of the target Internet application;
[0055] A screening unit is used to screen out the Internet data to be detected from the application data of the target Internet application based on the abnormal feedback information.
[0056] A third aspect of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data processing method as described in the first aspect is implemented.
[0057] A fourth aspect of the present disclosure provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the data processing method as described in the first aspect when executing the computer program.
[0058] A fifth aspect of the present disclosure provides a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the data processing method as described in the first aspect.
[0059] The data processing method provided by the embodiments of the present disclosure obtains internet data to be detected that is associated with a target scenario in a target internet application, where the internet data to be detected includes internet data collected from multiple data attribute dimensions; determines a target data attribute dimension from the multiple data attribute dimensions based on the target scenario, and generates content information and statistical information for each internet data to be detected based on the target data attribute dimension; performs anomaly screening on each internet data to be detected based on the content information and the statistical information, and determines multiple first abnormal accounts in the target internet application based on the anomaly screening results; obtains a time behavior sequence corresponding to each first abnormal account; performs behavior sequence comparison on the multiple first abnormal accounts based on the time behavior sequence, and determines multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results.
[0060] In this way, account anomaly identification is not only achieved by evaluating the content and behavior of the account itself, but also by obtaining the Internet data to be detected from multiple data attribute dimensions, and using the target data attribute dimensions that need to be paid attention to in the target scenario to generate the content information and statistical information of the Internet data to be detected for the first anomaly screening. Therefore, the first anomaly screening can be performed based on the attribute dimensions that the target scenario focuses on, and the first abnormal account obtained by the screening can be obtained based on more comprehensive multiple data attribute dimensions. After the first anomaly screening, a second anomaly screening is performed based on the time behavior sequence of the first abnormal account to achieve a more detailed screening of the first abnormal account from the perspective of the relationship between time and account behavior. In summary, the data processing method provided by the embodiment of the present disclosure obtains Internet data based on more comprehensive data attribute dimensions, avoids omissions of data information, and performs abnormal account detection in a targeted manner in the target data dimensions that the target scenario focuses on, and then achieves more detailed abnormal account screening based on the relationship between time and account behavior. In this way, the accuracy of abnormal account detection in the application is improved.
[0061] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present disclosure. The purposes and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.
[0063] Figure 1 A system architecture diagram for the data processing method according to an embodiment of the present disclosure;
[0064] Figure 2A-2B This is a schematic diagram of a scenario in which the data processing method according to an embodiment of the present disclosure is applied to a scenario in which a group of abnormal accounts is detected;
[0065] Figure 3 A flowchart of the data processing method provided by the present disclosure;
[0066] Figure 4 This is a schematic diagram of generating content information and statistical information of Internet data to be detected based on dimensional relationships according to an embodiment of the present disclosure;
[0067] Figure 5 This is a schematic diagram of obtaining a first abnormal account after performing abnormal screening on statistical information and content information respectively according to an embodiment of the present disclosure;
[0068] Figure 6 is a schematic diagram of a temporal behavior sequence of an account according to one embodiment of the present disclosure;
[0069] Figure 7 is a diagrammatic representation of a temporal behavior sequence according to an embodiment of the present disclosure;
[0070] Figure 8 is a schematic diagram of a spectrum behavior sequence according to an embodiment of the present disclosure;
[0071] Figure 9 This is a schematic diagram of obtaining multiple second abnormal accounts after performing sequence comparison on a time behavior sequence and a spectrum behavior sequence according to an embodiment of the present disclosure;
[0072] Figure 10 is a schematic diagram of a graph network according to an embodiment of the present disclosure;
[0073] Figure 11 is a schematic diagram of dividing a graph network into multiple subgraphs according to an embodiment of the present disclosure;
[0074] Figure 12 This is a schematic diagram of obtaining a target physical address and a target mobile phone number based on sorting and summarizing each node in a subgraph according to an embodiment of the present disclosure;
[0075] Figure 13 This is a schematic diagram of mining organizational information behind abnormal behavior groups through associations between nodes in a subgraph according to an embodiment of the present disclosure;
[0076] Figure 14 Another flowchart of the data processing method provided by the present disclosure;
[0077] Figure 15 is an implementation detail diagram of a data processing method provided according to an embodiment of the present disclosure;
[0078] Figure 16 is a schematic diagram of an overall implementation of a data processing method provided according to an embodiment of the present disclosure;
[0079] Figure 17 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure;
[0080] Figure 18 is a diagram of a terminal structure for implementing various methods according to an embodiment of the present disclosure;
[0081] Figure 19 It is a server structure diagram for implementing various methods according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.
[0083] Before further explaining the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:
[0084] Internet Protocol Address (IP address): It assigns a logical address to each network and each host on the Internet to mask the differences in physical addresses.
[0085] Cosine similarity: A method for measuring the similarity between two vectors, which is calculated based on the cosine value between the vectors. The value range of cosine similarity is [-1,1]. The closer the value is to 1, the more similar the two vectors are, the closer the value is to -1, the less similar the two vectors are, and the value is 0, which means the two vectors are unrelated. The geometric meaning of cosine similarity is: in n-dimensional space, the cosine similarity of two vectors is equal to the cosine value of the angle between them. When the directions of two vectors are closer, their angle is smaller, the cosine value is larger, and the similarity is higher; when the directions of two vectors are opposite, their angle is 180 degrees, the cosine value is -1, and the similarity is the lowest. In fields such as natural language processing, computer vision, and recommendation systems, cosine similarity is often used to measure the similarity between vectors such as text and images.
[0086] In related technologies, data processing methods for detecting abnormal accounts often use a single data dimension. For example, they only test the account dimension, using the account's content and statistical features to determine whether an account is abnormal. However, detecting from a single data dimension may miss some relatively hidden abnormal accounts. For example, to circumvent the system's detection of the account dimension, multiple accounts may be registered on the same device and collaborate to engage in malicious activities. The system's abnormality detection results for each account may all be normal accounts. Therefore, the accuracy of abnormal account detection using a single data dimension is low.
[0087] In order to solve the problem of low detection accuracy, the present disclosure provides a data processing method, so as to improve the accuracy of abnormal account detection in applications.
[0088] System architecture and scenario description of the application of the embodiments of the present disclosure
[0089] Figure 1 This is a system architecture diagram for the data processing method according to an embodiment of the present disclosure, which includes a terminal 140, the Internet 130, a gateway 120, a server 110, and the like.
[0090] Terminal 140 includes various device forms, such as desktop computers, laptops, PDAs (personal digital assistants), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, intelligent voice interaction devices, smart home appliances, or aircraft. Furthermore, it can be a single device or a collection of multiple devices. Terminal 140 can communicate with Internet 130 via wired or wireless means to exchange data.
[0091] Server 110 is a computer system that provides certain services to terminal 140. Compared to ordinary terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a high-performance computer (such as a virtual machine), or a combination of portions of multiple high-performance computers (such as virtual machines).
[0092] Gateway 120, also known as a gateway or protocol converter, implements network interconnection at the transport layer and is a computer system or device that performs a conversion function. It acts as a translator between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 140 to server 110 are sent through gateway 120 to the corresponding server 110. Messages sent from server 110 to terminal 140 are also sent through gateway 120 to the corresponding terminal 140.
[0093] The data processing method provided by the embodiments of the present disclosure may be implemented in the terminal 140 , or in the server 110 , or may be implemented partially in the terminal 140 and partially in the server 110 .
[0094] When the data processing method provided by the embodiment of the present disclosure is implemented in the terminal 140, the terminal 140 obtains the Internet data to be detected that is associated with the target scenario in the target Internet application; determines the target data attribute dimension among multiple data attribute dimensions based on the target scenario, and generates content information and statistical information for each Internet data to be detected based on the target data attribute dimension; further, the terminal 140 performs anomaly screening on each Internet data to be detected based on the content information and statistical information, and determines multiple first abnormal accounts in the target Internet application based on the anomaly screening results; finally, the terminal 140 obtains the time behavior sequence corresponding to each first abnormal account and performs behavior sequence comparison on the multiple first abnormal accounts based on the time behavior sequence, and determines multiple second abnormal accounts among the multiple first abnormal accounts based on the comparison results.
[0095] When the data processing method provided by the embodiment of the present disclosure is implemented in the server 110, the server 110 obtains the Internet data to be detected that is associated with the target scenario in the target Internet application; determines the target data attribute dimension among multiple data attribute dimensions based on the target scenario, and generates content information and statistical information for each Internet data to be detected based on the target data attribute dimension; further, the server 110 performs anomaly screening on each Internet data to be detected based on the content information and statistical information, and determines multiple first abnormal accounts in the target Internet application based on the anomaly screening results; finally, the server 110 obtains the time behavior sequence corresponding to each first abnormal account and performs behavior sequence comparison on the multiple first abnormal accounts based on the time behavior sequence, and determines multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results.
[0096] When the data processing method provided by the embodiment of the present disclosure is partially implemented in the terminal 140 and partially implemented in the server 110, the terminal 140 obtains the Internet data to be detected that is associated with the target scenario in the target Internet application; determines the target data attribute dimension among multiple data attribute dimensions based on the target scenario, and generates content information and statistical information for each Internet data to be detected based on the target data attribute dimension; then, the server 110 performs anomaly screening on each Internet data to be detected based on the content information and statistical information, and determines multiple first abnormal accounts in the target Internet application based on the anomaly screening results; finally, the server 110 obtains the time behavior sequence corresponding to each first abnormal account and performs behavior sequence comparison on the multiple first abnormal accounts based on the time behavior sequence, and determines multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results.
[0097] The data processing method of the embodiment of the present disclosure can be applied in various scenarios, such as Figure 2A-2B In the scenario shown in the figure, a group of abnormal accounts are detected.
[0098] like Figure 2A As shown, app A1 hosted an "XXX event," in which each registered account could only participate once. To profit from the event, accounts A, B, C, D, and E, registered on the host network with IP address XXXX, participated in the event together. This constitutes an act that undermines the fairness of the event.
[0099] In order to maintain the fairness of the activities, Figure 2B As shown, abnormal accounts in the activity scenario are detected in real time. Figure 2BThe data shown in Figure 1 includes checking the internet data of accounts a through e at the account level and XXXX at the IP address level. After screening the internet data for each data attribute dimension for anomalies, the anomaly screening results show that accounts a through e are normal accounts, but IP address XXXX is identified as an anomalous IP address due to operating multiple accounts simultaneously. Based on the IP addresses, accounts a through e are identified as anomalous accounts.
[0100] pass Figure 2A-2B It can be seen that Accounts a to e have jointly engaged in group bad behavior, but only performing anomaly detection on individual accounts cannot lead to the conclusion that the accounts are abnormal. The data processing method of the disclosed embodiment can collect Internet data from multiple data attribute dimensions for anomaly detection, and is not limited to a single account dimension. Therefore, after obtaining Internet data in the IP address dimension, the conclusion that the IP address is abnormal can be drawn based on the content information and statistical information of the IP address, and then based on the IP address, it can be concluded that Accounts A to E are abnormal accounts. It can be seen that the data processing method of the disclosed embodiment can perform abnormal account detection from Internet data in multiple data attribute dimensions, thereby improving the accuracy of abnormal account detection.
[0101] The above examples do not limit the scope of protection of this case.
[0102] General description of the embodiments of the present disclosure
[0103] According to one embodiment of the present disclosure, a data processing method is provided. Figure 3 FIG. 1 is a flow chart of a data processing method provided by the present disclosure. The method can be applied to a data processing device, which can be integrated into a computer device, which can specifically be a terminal or a server. The data processing method may include:
[0104] Step 310: Obtain the Internet data to be detected in the target Internet application and associated with the target scenario.
[0105] Target internet applications can be of various types, such as content push apps, resource exchange apps, and instant messaging apps. Each type of application may contain anomalous accounts that could harm the application environment. For example, accounts in content push apps disseminate harmful content, accounts in resource exchange apps exchange resources at extremely low costs, and accounts in instant messaging apps send harmful messages to other accounts. Therefore, accurately detecting anomalous accounts is crucial for each internet application.
[0106] Internet applications can include transactions across multiple scenarios, with different business processing logic applied to different scenarios. For example, content push applications can include content push scenarios and object interaction scenarios. Resource exchange applications can offer a variety of resource exchange activities, each with different implementation rules, allowing each activity to be divided into multiple scenarios. Therefore, the methods used by anomalous accounts to engage in illicit activities can vary across different scenarios within an internet application. For example, in content push applications, anomalous accounts can upload inappropriate content within the content push scenario. In object interaction scenarios, they can exploit features like "comment" and "like" to increase content exposure and recommendation rankings through informal means to profit. In resource exchange applications, when a scenario limits the number or number of exchanges per account, objects can exchange for more resources by registering accounts in bulk. When a scenario has tasks and corresponding rewards, anomalous accounts can use informal means to complete a large number of tasks in a short period of time to obtain higher rewards.
[0107] Because abnormal account behavior may vary in different scenarios, when detecting abnormal accounts in internet applications, we can conduct targeted detection on the target scenario. We collect a large amount of internet data from the target scenario and detect abnormal accounts based on the business rules of the target scenario to improve the accuracy of abnormal account detection.
[0108] The Internet data to be detected associated with the target scenario includes Internet data collected from multiple data attribute dimensions. The data attribute dimension can be an individual dimension used for account detection, such as the account dimension, IP address dimension, device number dimension, and registered mobile phone number dimension. The Internet data can be the traffic data generated based on account behavior in the target Internet application. The Internet data to be detected associated with the target scenario can be the traffic data generated by the individual behavior corresponding to each data attribute dimension in the target scenario of the target Internet application. For example, when collecting Internet data from the account dimension, the account login data and behavior data can be collected; when collecting Internet data from the IP address dimension, the registration request data and login request data issued under the IP address can be collected; when collecting Internet data from the registered mobile phone number dimension, the account registration request data and login request data corresponding to the registered mobile phone number can be collected.
[0109] The purpose of collecting internet data to be tested is to identify anomalous accounts within the data. For example, internet data based on the account dimension can be used to determine whether the corresponding account is anomalous; internet accounts based on the IP address dimension can be used to determine whether the account registered or logged in under the IP address is anomalous. Obtaining the internet data to be tested associated with the target scenario within the target internet application can be achieved by obtaining traffic data generated within the target scenario, recorded on the target internet application's backend server.
[0110] Because target internet applications generate a large amount of internet data, including accounts, IP addresses, and registered mobile phone numbers, a small percentage of these data come from anomalous accounts or other data attributes related to anomalous accounts. Therefore, to improve the efficiency of identifying anomalous accounts, the internet data in the target internet applications can be screened.
[0111] In one embodiment, before obtaining the to-be-detected Internet data associated with the target scenario in the target Internet application, the data processing method provided by the embodiment of the present disclosure further includes:
[0112] Obtaining abnormal feedback information on Internet data across multiple data attribute dimensions in target Internet applications;
[0113] Based on the abnormal feedback information, the Internet data to be detected is filtered out from the application data of the target Internet application.
[0114] Abnormal feedback information can be used to preliminarily determine whether an individual in a certain data attribute dimension within the target internet application has exhibited abnormal behavior. For example, an account may have been browsing pages 24 hours a day; a device corresponding to a device number logged into multiple accounts within an hour; or a single IP address issued multiple account registration requests within 20 minutes. These behaviors are all considered abnormal behaviors within the target internet application, but this does not necessarily mean that the individuals in the data attribute dimension corresponding to these abnormal behaviors are abnormal individuals. When abnormal feedback information is obtained for internet data, it indicates that the corresponding individual has exhibited abnormal behavior and may be an abnormal individual. The associated account may also be an abnormal account, requiring further testing of the internet data to determine this.
[0115] In one embodiment, obtaining abnormal feedback information on Internet data of multiple data attribute dimensions in the target Internet application can be obtained through an abnormality detection model in the target Internet application.
[0116] Anomaly detection models are used to perform real-time monitoring of data traffic generated by target internet applications, promptly identifying anomalous behavior within the applications. The anomaly detection model uses preset anomaly detection rules to detect data traffic. These rules can be manually set based on experience. For example, the number of registration requests from an IP address exceeding a predetermined value within a predetermined time period, or the number of access data generated by an account exceeding a predetermined value within a predetermined time period, can be detected. When the anomaly detection model detects anomalous behavior in internet data, it can generate anomaly feedback information. This anomaly feedback information generated by the anomaly detection model can also be manually reviewed to ensure its accuracy.
[0117] In another embodiment, in addition to the target internet application internally detecting internet data to obtain abnormality feedback information, abnormality feedback information can also be obtained through external data information of the target internet application. Specifically, the ways to obtain abnormality feedback information through the target internet application's external data information may include: feedback information from an external management organization, object feedback information of the target internet application, and abnormal data information of the external application.
[0118] Feedback from the external management organization originates from the external organization or institution responsible for managing the target internet application. Its purpose is to maintain the security and stability of the entire internet environment. Anomalous behavior data within the target internet application can sometimes be missed by internal system detection. External management organizations can conduct further inspections of internet data generated by the target internet application. When anomalous data is detected, anomaly feedback information is generated and provided to the target internet application for more accurate anomaly detection.
[0119] The target internet application's subject feedback is generated by subjects' feedback on unusual behavior while using the target internet application. If a subject discovers unusual behavior from another subject or account while using the target internet application, they can submit feedback about the unusual behavior through the feedback channels provided by the target internet application. Upon receiving feedback from the subject, anomaly feedback information is generated and anomaly detection is performed on the corresponding internet data.
[0120] External application anomaly data information originates from applications other than the target internet application. Some anomalous IP addresses or phone numbers may be involved in malicious activities across multiple applications. Therefore, different applications can share information. When an application other than the target internet application detects an anomalous IP address, it can send the information corresponding to this IP address to the target internet application. The target internet application can then generate anomaly feedback information specific to this IP address and perform anomaly detection on the corresponding internet data.
[0121] After obtaining the abnormal feedback information, the target internet application's application data can be filtered to identify the internet data to be tested based on the abnormal feedback information. Because the abnormal feedback information may include the specific individual who exhibited the abnormal behavior, such as an account number, IP address, and mobile phone number, the internet data corresponding to the abnormal individual in the abnormal feedback information can be obtained from the internet application's application data as the internet data to be tested. For example, if the abnormal feedback information includes abnormal feedback for IP address A1, the internet data related to IP address A1 can be obtained from the target internet application as the internet data to be tested.
[0122] Using abnormal feedback information to filter the Internet data to be detected can narrow the scope of abnormal account detection, avoid investing a lot of time and computing power in normal accounts, and improve the efficiency of abnormal account detection.
[0123] Step 320: Determine a target data attribute dimension from a plurality of data attribute dimensions based on the target scenario, and generate content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension.
[0124] Because the behavior of abnormal accounts may vary in different scenarios, the target data attribute dimension of the target scenario can be determined when performing targeted anomaly detection on the target scenario. The target data attribute dimension can be the data attribute dimension that requires attention when detecting abnormal accounts within the target scenario. For example, when a scenario in a resource exchange application limits the number of resources that can be exchanged by each account, users often register multiple accounts under the same IP address to obtain more resources. In this scenario, the data attribute dimension that requires the most attention when detecting abnormal accounts may be the IP address dimension. By detecting abnormal behavior in the IP address dimension, it can be determined whether there is a group of abnormal accounts, so the target data attribute dimension is the IP address dimension. When another scenario in a resource exchange application sets tasks that require accounts to complete and corresponding rewards, users who want to obtain more rewards can complete a large number of tasks through informal means. In this scenario, since abnormal behavior is primarily initiated by accounts, the data attribute dimension that requires the most attention when detecting abnormal accounts is the account dimension, so the target data attribute dimension is the account dimension.
[0125] To determine the target data attribute dimension from among multiple data attribute dimensions based on the target scenario, you can identify the data attribute dimension with a higher probability of abnormal behavior within the target scenario and use this data attribute dimension as the target data attribute dimension. For example, in a scenario where most abnormal behavior is initiated by IP address, the target data attribute dimension would be the IP address dimension.
[0126] The internet data to be tested acquired in step 310 includes internet data collected across multiple data attribute dimensions, such as account data, IP address data, and device ID data. Because the target scenario has a target data attribute dimension of interest, it is difficult to determine whether anomalies exist in the corresponding internet data to be tested within the target scenario using data attribute dimensions other than the target data attribute dimension. Therefore, when performing anomaly detection on each piece of internet data to be tested, content information and statistical information for each piece of internet data to be tested can be generated based on the target data attribute dimension.
[0127] The content information of the internet data to be tested can be text information related to the internet data to be tested, such as account registration information, device login information, business activity information, and account relationship information. Statistical information can be numerical information related to the internet data to be tested, such as the number of account logins in a day, the account online time, and the number of accounts logged in from the same IP address.
[0128] In one embodiment, content information and statistical information of each to-be-detected Internet data are generated based on the target data attribute dimension, including:
[0129] If the data attribute dimension corresponding to the Internet data to be detected is the target data attribute dimension, obtaining content information and statistical information from the Internet data to be detected;
[0130] If the data attribute dimension corresponding to the Internet data to be detected is not the target data attribute dimension, obtain the dimensional relationship between the data attribute dimension corresponding to the Internet data to be detected and the target data attribute dimension, and obtain content information and statistical information of the target data attribute dimension in the Internet data based on the dimensional relationship.
[0131] If the data attribute dimension of the internet data to be tested is the target data attribute dimension, anomaly detection can be performed based on the content and statistical information of the internet data to be tested. For example, if the target data attribute dimension is the IP address, and the data attribute dimension of the internet data to be tested is also the IP address, then the content and statistical information corresponding to the IP address can be obtained from the internet data to be tested.
[0132] If the data attribute dimension corresponding to the Internet data to be detected is not the target data attribute dimension, the dimensional relationship between the data attribute dimension and the target data attribute dimension can be obtained. Figure 4An example of a dimensional relationship is shown in , where IP addresses and devices can be associated using time. For example, at time T1, IP address A1 is used to log in to device E1. At time T2, the user logs in to device E2, but the device's IP address has previously used both IP addresses A1 and A2. At time T3, the user logs in to device E1, but the device's IP address has previously used both IP addresses A1 and A2. Devices and accounts can be directly associated. For example, account a, account b, and account c are logged in to device E1; and account a, account b, and account d are logged in to device E2. Similarly, IP addresses and accounts can also be directly associated, which will not be discussed here.
[0133] Therefore, when the data attribute dimension corresponding to the internet data to be tested is not the target data attribute dimension, content information and statistical information of the internet data to be tested can be generated based on the dimensional relationship. For example, if the target data attribute dimension is an account number, but the data attribute dimension corresponding to the internet data to be tested is an IP address, then the content information of the internet data to be tested can include registration information, login information, activity information, and relationship information of accounts registered based on IP addresses, and statistical information can include the online time, number of logins, and total number of logins of accounts registered based on IP addresses.
[0134] Generating content information and statistical information based on the target data attribute dimensions can assist in anomaly detection of the Internet data to be detected, thereby improving the accuracy of anomaly detection in the target scenario.
[0135] Step 330: Perform anomaly screening on each to-be-detected Internet data based on the content information and the statistical information, and determine a plurality of first abnormal accounts in the target Internet application based on the anomaly screening results.
[0136] Anomaly screening of the Internet data to be detected based on content information and statistical information enables a comprehensive judgment of whether the Internet data to be detected is abnormal based on the content and behavior of the Internet data to be detected, which is conducive to improving the accuracy of anomaly detection.
[0137] In one embodiment, performing anomaly screening on each to-be-detected Internet data based on content information and statistical information, and determining a plurality of first abnormal accounts in the target Internet application based on the anomaly screening results, includes:
[0138] Perform anomaly matching on statistical information according to preset strategy rules to obtain a first screening result;
[0139] Classify the content information as abnormal based on a preset neural network model to obtain a second screening result;
[0140] Determining a plurality of abnormal Internet data from the plurality of Internet data to be detected according to the first screening result and the second screening result;
[0141] A plurality of first abnormal accounts in a target Internet application is determined based on the plurality of abnormal Internet data.
[0142] The preset policy rules may be rules pre-set based on the behavior of historical abnormal Internet data, and are used to verify whether the statistical information corresponding to the Internet data to be detected is abnormal. There may be multiple preset policy rules, and the preset policy rules may be different for the Internet data to be detected in different data attribute dimensions. For example, for the Internet data to be detected in the IP address data attribute dimension, the preset policy rules may include: the number of account registrations and logins initiated within one hour is greater than 10; the account has logged in multiple times and the average online time is less than five minutes; the number of accounts online at the same time is greater than 8, etc. When the statistical information of the Internet data to be detected matches one or a combination of multiple rules in the preset policy rules, the first screening result may indicate that there is an abnormality in the statistical information of the Internet data to be detected.
[0143] Although statistical information is numerical information about the internet data to be tested, it also has multiple attributes, specifically statistical information about static attributes, behavioral attributes, and relationship attributes. Statistical information about static attributes can be information represented by numerical values that does not change over time in the internet data to be tested, such as the registration time and device number of an account. Statistical information about behavioral attributes can be information collected based on the behavior of the internet data to be tested, such as the number of logins and online time of an account. Statistical information about relationship attributes can be information used to describe the relationship between the internet data to be tested and other internet data, such as the number of interactions and the number of associated accounts.
[0144] For different types of statistical information, the preset policy rules are also different. Therefore, in one embodiment, the statistical information is matched with anomalies according to the preset policy rules to obtain a first screening result, including:
[0145] Obtaining a first sub-policy rule corresponding to the statistical information of the static attribute, a second sub-policy rule corresponding to the statistical information of the behavioral attribute, and a third sub-policy rule corresponding to the statistical information of the relationship attribute;
[0146] Anomaly matching is performed on the statistical information based on the first sub-strategy rule, the second sub-strategy rule, and the third sub-strategy rule to obtain a first screening result.
[0147] The first sub-policy rule specifies the conditions for abnormal statistical information of static attributes, such as accounts associated with the same IP address having similar registration times, or a device number that has been marked as an abnormal device.
[0148] The second sub-policy rule specifies the conditions for abnormal statistical information of behavioral attributes, such as the number of accounts online simultaneously under the IP address is greater than 10, or the number of times a device logs into an account is greater than 8 times within 10 minutes.
[0149] The third sub-policy rule specifies the conditions for abnormal statistical information of relationship attributes, such as the number of associated accounts being greater than 1,000, or interacting with more than 20 associated accounts at the same time.
[0150] When performing anomaly matching on statistical information based on the first, second, and third sub-policy rules, the first sub-policy rule can be used to perform a first match on the statistical information of static attributes; the second sub-policy rule can be used to perform a second match on the statistical information of behavioral attributes; and the third sub-policy rule can be used to perform a third match on the statistical information of relationship attributes. A first screening result is obtained based on the results of the first, second, and third matches. When two or more of the results of the first, second, and third matches indicate abnormal information, the first screening result may indicate that the statistical information is abnormal. For example, if the statistical information of behavioral attributes and relationship attributes displays abnormality after abnormal matching, then the first screening result indicates that the statistical information is abnormal.
[0151] When performing anomaly matching on statistical information based on the first sub-strategy rule, the second sub-strategy rule and the third sub-strategy rule, the first sub-strategy rule, the second sub-strategy rule and the third sub-strategy rule can also be combined; and anomaly matching on the statistical information is performed based on the combined result to obtain a first screening result.
[0152] By combining the first, second, and third sub-policy rules using various logical relationships, you can create various policy combinations. For example, a policy combination might include: the device number has been marked as an abnormal device, and the device has logged into an account more than 8 times within 10 minutes.
[0153] Strategy combinations can be divided into hit strategy combinations and filtering strategy combinations. When the statistical information meets the rules of the hit strategy combination, the first screening result can indicate that the internet data to be tested is anomalous internet data; when the statistical information meets the filtering strategy combination, the first screening result can indicate that the internet data to be tested is not anomalous internet data. When using the strategy combination to determine the first screening result, an anomaly score can also be output based on the hit status of the hit strategy combination and the filtering strategy combination. The more hit strategy combinations the statistical information meets, the higher the anomaly score can be, indicating a greater likelihood that the statistical information is anomalous.
[0154] Different policy rules are obtained for different types of statistical information, and exception matching is performed based on multiple policy rules, so that more accurate policy rules can be set for different types of statistical information, which is conducive to improving the accuracy of exception matching.
[0155] While using the preset strategy rules to perform anomaly matching on the statistical information, the content information can also be classified as abnormal based on the preset neural network model. The second screening result obtained by the abnormal classification indicates whether the content information of the Internet data to be detected is abnormal information.
[0156] The preset neural network model for performing abnormal classification on content information may be a pre-trained neural network model, such as a Resnet convolutional neural network model.
[0157] In one embodiment, the content information is classified as abnormal based on a preset neural network model to obtain a second screening result, including:
[0158] Vectorize the content information to obtain a content vector;
[0159] The content vector is input into the preset neural network model for anomaly classification to obtain the second screening result.
[0160] During data processing, content information, as textual information, is difficult to directly recognize and process using computer language. Therefore, before performing anomaly screening, the content can be vectorized to a form that can be processed by a neural network model.
[0161] Vectorizing content information can be done by inputting it into a model using embedding technology, such as the BERT model, to generate content vectors. Embedding technology vectorizes text or discrete data into continuous vectors to capture the underlying relationships and structures between them.
[0162] The preset neural network model can perform anomaly classification based on the content vector to obtain a second screening result.
[0163] Content information can be text information in the internet data to be detected. It also has multiple attributes, specifically statistical information of static attributes, behavioral attributes, and relationship attributes. Static attribute content information can be text information related to the individual in the internet data to be detected that does not change over time, such as account names and object-defined account descriptions. Behavioral attributes can be text information related to actions in the internet data to be detected, such as registration instructions and login instructions. Relationship attributes can be text information used to describe the relationship between the internet data to be detected and other internet data, such as account relationships and device relationships.
[0164] Before using the preset neural network model to classify content information anomalies, the content information of different attributes can also be vectorized separately. After vectorization, the content vectors of each attribute are concatenated to form a concatenated vector containing all the features of the content information. The concatenated vector is input into the preset neural network model for anomaly classification to obtain a second screening result. The second screening result can directly indicate whether the content information is anomaly in the form of a label, or it can also use an anomaly score to indicate the likelihood of the content information being anomaly.
[0165] After obtaining the first screening result and the second screening result, a plurality of abnormal Internet data can be determined from the plurality of Internet data to be detected based on the first screening result and the second screening result.
[0166] In one embodiment, the first screening result indicates whether the statistical information is abnormal. When the first screening result is 1, it indicates that the statistical information is abnormal; otherwise, when the first screening result is 0, it indicates that the statistical information is normal. The second screening result indicates whether the content information is abnormal. When the second screening result is 1, it indicates that the content information is abnormal; otherwise, it indicates that the content information is normal.
[0167] When at least one of the first screening result and the second screening result corresponding to the Internet data to be detected indicates that the corresponding information is abnormal, it can be determined that the Internet data to be detected is abnormal Internet data.
[0168] In another embodiment, determining a plurality of abnormal Internet data from a plurality of Internet data to be detected based on the first screening result and the second screening result includes:
[0169] Obtaining a first weight coefficient corresponding to the first screening result, and obtaining a second weight coefficient corresponding to the second screening result;
[0170] Performing weighted calculation on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result;
[0171] According to the target screening results, multiple abnormal Internet data are determined from multiple Internet data to be detected.
[0172] The first weight coefficient and the second weight coefficient can respectively indicate the contribution of the first screening result and the second screening result to the detection of anomalies in internet data. Furthermore, they can respectively indicate the contribution of anomalies in statistical information and anomalies in content information to the detection of anomalies in internet data. The first weight coefficient and the second weight coefficient can be pre-set based on the business requirements of the target scenario.
[0173] The first screening result and the second screening result may be weightedly calculated according to the first weight coefficient and the second weight coefficient.
[0174] In one embodiment, weighted calculation is performed on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result, including:
[0175] determining a first score corresponding to the first screening result and a second score corresponding to the second screening result;
[0176] Performing weighted calculation on the first score and the second score based on the first weight coefficient and the second weight coefficient to obtain a target score;
[0177] The target screening results are determined based on the target scores.
[0178] The first score corresponding to the first screening result can indicate the likelihood that the statistical information is abnormal, and the second score corresponding to the second screening result can indicate the likelihood that the content information is abnormal. The target score obtained by weighted calculation using the first weight coefficient and the second weight coefficient can indicate the likelihood that the detected internet data is abnormal.
[0179] For example, the first score corresponding to the first screening result is 60, the second score corresponding to the second screening result is 70, the first weight coefficient is 0.6, and the second weight coefficient is 0.4. Therefore, the target score is 60*0.6+70*0.4=64.
[0180] After determining the target screening result based on the target score, the target screening result can be further determined based on the score threshold. The target screening result indicates whether the internet data to be tested is anomalous. If the target score is greater than or equal to the score threshold, the target screening result indicates that the internet data to be tested is anomalous; otherwise, the internet data to be tested is not anomalous.
[0181] Determining the target screening results according to the target score enables anomaly detection of the Internet data to be detected in a quantitative manner, thereby improving the accuracy of anomaly detection.
[0182] After obtaining the target screening result, multiple abnormal Internet data can be determined from the multiple Internet data to be detected based on the target screening result. Specifically, the target screening results corresponding to the multiple Internet data to be detected that are abnormal can be determined as abnormal Internet data.
[0183] The target screening result is obtained by weighted calculation of the first screening result and the second screening result, and different weights can be set for abnormal Internet data based on the business needs of the target scenario, so that the target screening result is more in line with the scenario requirements, which is conducive to improving the flexibility and accuracy of Internet data anomaly detection.
[0184] After determining the plurality of abnormal Internet data, a plurality of first abnormal accounts in the target Internet application may be determined based on the plurality of Internet data.
[0185] In one embodiment, determining a plurality of first abnormal accounts in a target Internet application based on a plurality of abnormal Internet data includes:
[0186] Obtain the account information corresponding to each abnormal Internet data and the target Internet application;
[0187] A plurality of first abnormal accounts in the target Internet application is determined based on the account information.
[0188] The account information corresponding to the abnormal Internet data and the target Internet application can be an account related to the abnormal Internet data in the target Internet application. For example, if the abnormal Internet data is an IP address, then the account information can be an account that has logged in or registered the target Internet application at the IP address; if the abnormal Internet data is an account, then the account information can be the account itself.
[0189] After determining the multiple account information, the accounts corresponding to the account information may be determined as the multiple first abnormal accounts in the target Internet application.
[0190] In summary, in this embodiment, the process of performing abnormal screening on the Internet data to be detected based on the content information and statistical information to obtain multiple first abnormal accounts can be as follows: Figure 5 As shown. Statistical information in the internet data to be detected is screened using policy rules, and content information is screened using a preset neural network. For statistical information, static attributes can be screened using preset policy rules R1, R2, and R3; behavioral attributes can be screened using preset policy rules R4, R5, and R6; and relationship attributes can be screened using preset policy rules R7, R8, and R9. Policy rules for different attributes can be combined to obtain policy combinations C1, C2, and C3. Based on the policy combinations, a first screening result corresponding to the statistical information can be obtained.
[0191] For content information, the static, behavioral, and relational attributes of the content can be vectorized separately, and then the multiple vectors can be concatenated into a single vector. The concatenated vectors are input into a preset neural network for anomaly classification, resulting in a second screening result corresponding to the content information. Based on the first and second screening results, it can be determined whether the internet data to be tested is abnormal internet data. If the internet data to be tested is determined to be abnormal internet data, a first abnormal account can be identified based on the abnormal internet data.
[0192] Abnormal screening is performed using the first screening result obtained by anomaly matching the statistical information and the second screening result obtained by anomaly classification of the content to obtain multiple first abnormal accounts in the abnormal Internet data. After anomaly detection is performed on the statistical information and content information respectively, a comprehensive analysis of the detection results of the two is performed to determine whether the Internet data to be detected is abnormal, ensuring that relatively accurate conclusions can be obtained in the anomaly detection of the statistical information and content information, thereby improving the accuracy of anomaly detection on the Internet data to be detected.
[0193] Step 340: Obtain the time behavior sequence corresponding to each first abnormal account.
[0194] The time behavior sequence can be a sequence of behaviors of the first abnormal account at different times arranged in chronological order. Figure 6 As shown, account a performed behavior M1 at time t1, behavior M2 at time t2, and behavior M3 at time t5; account b performed behavior M4 at time t1, behavior M5 at time t2, and behavior M6 at time t3; account c performed behavior M7 at time t2, behavior M8 at time t3, and behavior M9 at time t4. Sorting the behaviors of each account by occurrence time yields the corresponding time behavior sequence.
[0195] Time behavior series can also be represented by graphs, for example Figure 7 The following table shows the time behavior sequence corresponding to the four first abnormal accounts. Account behavior can be mapped to behavior identifiers, and a time behavior sequence diagram can be generated based on time and corresponding behavior identifiers.
[0196] The first abnormal account can be an abnormal account determined based on the statistical information and content information of the internet data to be tested. Because the internet data to be tested is obtained from multiple data attribute dimensions, if the internet data to be tested is abnormal in a certain data attribute dimension, the associated first abnormal account is likely abnormal, but requires further testing. Therefore, to improve the accuracy of abnormal account detection, the first abnormal account can be further determined to be abnormal from the perspective of account behavior based on the time behavior sequence of the first abnormal account.
[0197] The time series of behavior corresponding to the first abnormal account can be obtained from the account behavior log recorded in the backend system of the target internet application. The account behavior log contains the account's behavior within a predetermined time period, for example, logging into the application at 10:21, browsing content S1 at 10:22, and submitting a resource exchange request at 10:23.
[0198] In one embodiment, after obtaining the time behavior sequence corresponding to each first abnormal account, the data processing method provided by the embodiment of the present disclosure further includes: constructing a spectrum behavior sequence corresponding to each first abnormal account based on the time behavior sequence.
[0199] Spectral behavior sequences can be used to describe the periodicity and trend of time behavior sequences in the frequency domain. Periodicity can describe the repetitive behavior within a predetermined period in a time behavior sequence, while trends can be used to analyze behavioral changes in a time behavior sequence.
[0200] To construct a spectral behavior sequence corresponding to the first abnormal account based on the time behavior sequence, spectral analysis methods can be used to convert the time behavior sequence data from the time domain to the frequency domain, generating a spectral behavior sequence containing multiple frequency components. Spectral analysis methods can include Fourier transform, short-time Fourier transform, and wavelet transform.
[0201] The spectrum behavior sequence can also be represented by a graph, for example Figure 8 The spectrum behavior sequence shown in Figure 8 In the figure, the horizontal axis represents the behavior identifier corresponding to the behavior, the vertical axis represents the frequency of the behavior, and each broken line represents the frequency of each first abnormal account in performing different behaviors.
[0202] Constructing a spectrum behavior sequence corresponding to the first abnormal account based on the temporal behavior sequence can be used to analyze the periodicity and trend of the first abnormal account's behavior. Anomaly analysis based on the periodicity and trend of the first abnormal account's behavior helps improve the accuracy of abnormal account detection.
[0203] Step 350: Perform a behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence, and determine a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result.
[0204] The second abnormal account may be an account that is determined to be abnormal. Based on the time behavior sequence, a behavior sequence comparison is performed on multiple first abnormal accounts, and the first abnormal account with an abnormal comparison result may be determined as the second abnormal account.
[0205] In one embodiment, a behavior sequence comparison is performed on a plurality of first abnormal accounts based on a time behavior sequence, and a plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison results, including:
[0206] Get the reference time behavior sequence;
[0207] Calculate the similarity between the time behavior sequence and the reference time behavior sequence to obtain the comparison result;
[0208] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison result.
[0209] The reference time behavior sequence serves as a reference for determining whether the time behavior sequence of the first abnormal account is abnormal. It can be derived from the time behavior sequences of a large number of non-abnormal accounts and can be used to represent the behavioral trends of common non-abnormal accounts. If the time behavior sequence of the first abnormal account differs significantly from the reference time behavior sequence, the behavior of the first abnormal account can be considered abnormal.
[0210] Reference time behavior sequences can be obtained based on the historical time behavior sequences of accounts in the target internet application. These sequences are the time behavior sequences of normal accounts over a period of time before the current time. By analyzing and summarizing these time behavior sequences, a reference time behavior sequence can be obtained that represents the general behavior patterns of normal accounts in the target internet application.
[0211] Reference temporal behavior sequences can also be obtained based on external data from the target internet application. External data can include open source data and third-party data. Open source data can come from publicly available databases, from which the temporal behavior sequences of normal accounts published by multiple applications similar to the target internet application can be obtained. Third-party data can come from other internet applications that are informationally connected to and similar to the target internet application. For example, internet applications A1 and A2 are both resource exchange applications with similar business processes. Therefore, the temporal behavior sequences of accounts in internet applications A1 and A2 are also similar, and these two applications can be informationally connected to expand the reference data. In other words, the temporal behavior sequences of normal accounts in other internet applications can be obtained from third-party data. By analyzing and summarizing the temporal behavior sequences of normal accounts derived from external data, a reference temporal behavior sequence can be obtained that represents the common behavioral patterns of normal accounts in most applications similar to the target internet application.
[0212] After obtaining the reference time behavior sequence, the similarity between each first abnormal account and the reference time behavior sequence may be calculated.
[0213] In one embodiment, calculating the similarity between the time behavior sequence and the reference time behavior sequence to obtain a comparison result includes:
[0214] Extract features from the time behavior sequence based on a preset sequence feature generation model to obtain the first sequence feature;
[0215] Extract features from the reference time behavior sequence based on a preset sequence feature generation model to obtain a second sequence feature;
[0216] Calculate the cosine similarity between the first sequence feature and the second sequence feature, and determine the alignment result based on the cosine similarity.
[0217] A preset sequence feature generation model is used to extract features from sequences. After extracting the first sequence features from the time behavior sequence of the first abnormal account and obtaining the second sequence features from the reference time behavior sequence, the distance between the time behavior sequence of the first abnormal account and the reference time behavior sequence is determined by calculating the cosine similarity between the first and second sequence features. A higher cosine similarity value indicates a closer relationship between the first and second sequence features, and a more similar relationship between the time behavior sequence of the first abnormal account and the reference time behavior sequence.
[0218] When determining the comparison result value based on cosine similarity, the cosine similarity can be compared to a similarity threshold. The similarity threshold is used to assess whether the time behavior sequence of the first abnormal account is abnormal. If the cosine similarity is less than the similarity threshold, the comparison result can be that the time behavior sequence is abnormal compared to the reference time behavior sequence.
[0219] The cosine similarity between sequence features is used to calculate the similarity between the time behavior sequence and the reference time behavior sequence. The calculation efficiency is high and the calculation results are relatively accurate for indefinite length vectors.
[0220] When multiple second abnormal accounts are determined from multiple first abnormal accounts based on the comparison results, the first abnormal account whose time behavior sequence is indicated by the comparison result as abnormal can be determined as the second abnormal account.
[0221] Determining the second abnormal account by using the similarity between the reference time behavior sequence and the time behavior sequence of the first abnormal account provides a reference for anomaly detection of the time behavior sequence of the first abnormal account, which is conducive to improving the accuracy of determining the second abnormal account based on the time behavior sequence.
[0222] In the aforementioned embodiment, after obtaining the time behavior sequence corresponding to each first abnormal account, a spectrum behavior sequence corresponding to each first abnormal account is constructed based on the time behavior sequence. Based on this, in one embodiment, a behavior sequence comparison is performed on multiple first abnormal accounts based on the time behavior sequence, and multiple second abnormal accounts are determined from the multiple first abnormal accounts based on the comparison results, including: performing a behavior sequence comparison on multiple first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence, and determining multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results.
[0223] Sequence comparison based on temporal behavior sequences can determine the similarity between the first abnormal account and normal accounts in terms of their behavior sequences; sequence comparison based on spectral behavior sequences can determine the similarity between the first abnormal account and normal accounts in terms of their behavior cycles and trends. Therefore, comparing the behavior sequences of multiple first abnormal accounts based on temporal behavior sequences and spectral behavior sequences can improve the accuracy of identifying the second abnormal account.
[0224] The process of determining multiple second abnormal accounts from multiple first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence can be as follows: Figure 9 First, for each of the multiple first abnormal accounts, a corresponding time behavior sequence is obtained, and a spectrum behavior sequence is constructed based on the time behavior sequence. Second, a sequence comparison is performed based on the time behavior sequence and a sequence comparison is performed based on the spectrum behavior sequence. Finally, based on the comparison results of the time behavior sequence and the spectrum behavior sequence, multiple second abnormal accounts are screened out from the multiple first abnormal accounts.
[0225] In one embodiment, a behavior sequence comparison is performed on a plurality of first abnormal accounts based on a time behavior sequence and a spectrum behavior sequence, and a plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison results, including:
[0226] Obtaining a reference time behavior sequence and a reference spectrum behavior sequence;
[0227] Determining a first sub-matching result based on the similarity between the time behavior sequence and the reference time behavior sequence, and determining a second sub-matching result based on the similarity between the spectrum behavior sequence and the reference spectrum behavior sequence;
[0228] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the first sub-comparison result and the second sub-comparison result.
[0229] The reference time behavior sequence can be a time behavior sequence that can reflect the regularity of normal account behavior; similarly, the reference spectrum behavior sequence can be a spectrum behavior sequence that can reflect the normal account behavior cycle or behavior trend. Acquiring the reference time behavior sequence has been described in detail in the aforementioned embodiment and will not be repeated here. The reference spectrum behavior sequence can be acquired in the same manner as the reference time behavior sequence, or after acquiring the reference time behavior sequence, the reference spectrum behavior sequence can be constructed using the reference time behavior sequence.
[0230] The method for calculating the similarity between the time behavior sequence and the reference time behavior sequence has been described in detail in the above embodiment. The method for calculating the similarity between the spectrum behavior sequence and the reference spectrum behavior sequence is the same as the method for calculating the similarity between the time behavior sequence and the reference time behavior sequence, and will not be repeated here.
[0231] The first sub-comparison result determined based on the similarity between the time behavior sequence and the reference time behavior sequence can directly indicate whether the time behavior sequence is abnormal. When the first sub-comparison result is 1, it indicates that the time behavior sequence is abnormal, and when the first sub-comparison result is 0, it indicates that the time behavior sequence is normal. The first sub-comparison result can also be the similarity score between the time behavior sequence and the reference time behavior sequence. For example, if the similarity between the time behavior sequence and the reference time behavior sequence is calculated to be 0.6, then the first sub-comparison result is 0.6. The second sub-comparison result is similar to the first sub-comparison result. It can directly indicate whether the spectrum behavior sequence is abnormal, or it can be the similarity score between the spectrum behavior sequence and the reference spectrum behavior sequence. It will not be repeated here. The form of the first sub-comparison result and the second sub-comparison result can be consistent.
[0232] When the first sub-comparison result and the second sub-comparison result directly indicate whether the corresponding behavior sequence is abnormal, multiple second abnormal accounts are determined from the multiple first abnormal accounts based on the first sub-comparison result and the second sub-comparison result, including: for each first abnormal account, when the first sub-comparison result and the second sub-comparison result indicate that at least one of the time behavior sequence and the spectrum behavior sequence is abnormal, the first abnormal account is determined as the second abnormal account.
[0233] That is, only when the first sub-comparison result indicates that the temporal behavior sequence is normal and the second sub-comparison result also indicates that only the spectral behavior sequence is normal, the first abnormal account can be determined to be a normal account.
[0234] When the first sub-comparison result and the second sub-comparison result indicate a similarity score, determining a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the first sub-comparison result and the second sub-comparison result includes:
[0235] Calculating an overall similarity score of the first abnormal account based on the first similarity score of the first sub-comparison result and the second similarity score of the second sub-comparison result;
[0236] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the overall similarity scores.
[0237] The overall similarity score of the first abnormal account can be calculated by adding the first similarity score and the second similarity score. For example, if the first similarity score is 0.6 and the second similarity score is 0.5, the overall similarity score is 0.6+0.5=1.1.
[0238] To calculate the overall similarity score of the first abnormal account, weights may be assigned to the first and second similarity scores, and a weighted sum of the first and second similarity scores may be calculated based on the weights. For example, if the weight of the first similarity score is 0.6 and the weight of the second similarity score is 0.4, the overall similarity score is 0.6*0.6+0.5+0.4=1.26.
[0239] After calculating the overall similarity score, multiple second abnormal accounts are identified from the multiple first abnormal accounts based on the overall similarity score. First abnormal accounts whose corresponding overall similarity scores are less than a predetermined threshold can be identified as second abnormal accounts. For example, if the predetermined threshold is 1, the overall similarity scores of first abnormal account a, first abnormal account b, first abnormal account c, and first abnormal account d are 1.2, 0.9, 1.05, and 0.88, respectively. Therefore, the first abnormal accounts identified as second abnormal accounts include: first abnormal account b and first abnormal account d.
[0240] After performing anomaly detection on the time behavior sequence and the spectrum behavior sequence respectively, a comprehensive assessment is made based on the comparison results of the two to determine whether the first abnormal account has abnormal behavior, which is conducive to improving the accuracy of determining the second abnormal account.
[0241] After determining the second abnormal account, the relevant information of the second abnormal account can also provide a reference basis for subsequent account abnormality detection. Therefore, in one embodiment, after comparing the behavior sequences of multiple first abnormal accounts based on the time behavior sequence, and determining multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results, the data processing method provided by the embodiment of the present disclosure further includes:
[0242] Obtain abnormal data associated with the second abnormal account in the target data attribute dimension;
[0243] The abnormal data corresponding to the multiple second abnormal accounts are summarized and summarized to generate an abnormal information set.
[0244] Since the target data attribute dimension is the data attribute dimension of interest in the target scenario, after obtaining the second abnormal account in the target scenario, in order to analyze the abnormal situation in the target scenario, abnormal data of the target data attribute dimension corresponding to the second abnormal account can be obtained. For example, if the target data attribute dimension is the IP address dimension, then after obtaining the second abnormal account, the IP address corresponding to the second abnormal account can be determined, and the Internet data of the IP address can be obtained as abnormal data.
[0245] Abnormal data may include basic attribute information, abnormal situation information, and abnormal time information. Basic attribute information can be the basic information of the individual corresponding to the target data attribute dimension. For example, if the target data attribute dimension is an IP address, the basic attribute information may include the physical address corresponding to the IP address, the network type, and the backbone network. Abnormal situation information can indicate why the individual corresponding to the target data attribute dimension is abnormal, such as whether it matches a policy rule or the content information is determined to be abnormal. Abnormal time information can indicate the time when the abnormality occurred for the individual corresponding to the target data attribute dimension.
[0246] After obtaining the abnormal data, the abnormal data corresponding to multiple second abnormal accounts can be summarized and summarized. Since the target data attribute dimensions corresponding to multiple second abnormal accounts may be the same, when summarizing and summarizing, the abnormal data with the same target data attribute dimensions can be grouped together, using the target data attribute dimension as a unit, so that abnormal information can be obtained from the perspective of the target data attribute dimension as quickly as possible. The generated abnormal information set can be presented in a table, specifically, as shown in Table 1:
[0247]
[0248] Table 1
[0249] As can be seen from the example in Table 1, the second abnormal accounts corresponding to IP address A1 include account a and account b, etc. Table 1 also presents basic attribute information, abnormal situation information, and abnormal time information of IP address A1.
[0250] By summarizing and analyzing abnormal data, we can intuitively obtain abnormal data generated in the target scenario of the target Internet application. This abnormal data can be used to quickly handle abnormal events in the target scenario and provide data reference for subsequent abnormal account detection.
[0251] Abnormal accounts in the target Internet application sometimes engage in bad activities in the form of a group. Therefore, after determining the second abnormal account, other undetected abnormal accounts can be checked based on the second abnormal account.
[0252] In one embodiment, after comparing the behavior sequences of multiple first abnormal accounts based on the time behavior sequence and determining multiple second abnormal accounts from the multiple first abnormal accounts based on the comparison results, the data processing method provided by the embodiment of the present disclosure further includes:
[0253] Determining a plurality of abnormal data attribute nodes corresponding to the plurality of second abnormal accounts, and associated data attribute nodes corresponding to the plurality of abnormal data attribute nodes;
[0254] Constructing a graph network with the plurality of second abnormal accounts, the plurality of abnormal data attribute nodes, and the associated data attribute nodes;
[0255] The graph network is divided into multiple subgraphs based on the closeness of the node relationships in the graph network, and multiple sets of abnormal accounts are determined based on the multiple subgraphs.
[0256] The abnormal data attribute node corresponding to the second abnormal account may be a data attribute dimension corresponding to the internet data to be detected corresponding to the second abnormal account. For example, if the second abnormal account is detected based on internet data corresponding to an IP address, then the corresponding abnormal data attribute node is the IP address corresponding to the second abnormal account.
[0257] The associated data attribute nodes corresponding to the abnormal data attribute node can be other data attribute nodes associated with the abnormal data attribute node. For example, if the abnormal data attribute node is an IP address, then the associated data attribute nodes can be the device and mobile phone number related to the IP address.
[0258] A graph network is constructed based on the relationship between the plurality of second abnormal accounts, the corresponding plurality of abnormal data attribute nodes, and the associated data attribute nodes. Figure 10 The diagram shown is a diagram of a graph network. Figure 10 It can be seen intuitively that the second abnormal account a is associated with device E1 and IP address A1; IP address A1 is also associated with the second abnormal account b and the second abnormal account c; the second abnormal account b is also associated with mobile phone number P1 and IP address A2; in addition to this, there is other information, which will not be repeated here.
[0259] After obtaining the graph network corresponding to each node, the graph network can be divided into multiple subgraphs according to the closeness of the node relationship. Specifically, nodes with a higher degree of closeness of node relationship can be divided into a subgraph. The closeness of the node relationship can be determined by calculating the association relationship between nodes. Specifically, the Fast Unfolding algorithm can be used. Using the Fast Unfolding algorithm to divide the graph network into multiple subgraphs can specifically include: treating each node in the graph network as a separate subgraph; calculating the subgraph modularity, where the modularity is an indicator to measure the quality of the community structure, and the higher the value, the tighter and more stable the community structure; traversing each node in the graph network, trying to move it from the current subgraph to another subgraph, and determining whether the subgraph structure has changed based on the modularity before and after the move; repeating the above steps until the modularity of each subgraph no longer increases after moving the node or merging the subgraph. Each of the multiple subgraphs divided by the Fast Unfolding algorithm has a high degree of closeness. As Figure 11As shown in the figure, after the graph network is divided into subgraphs, three subgraphs are obtained, and the nodes in the three subgraphs are represented by circles, squares and triangles respectively.
[0260] The nodes in a subgraph can be identified as an abnormal behavior group, and the accounts in each subgraph can be extracted to construct the abnormal account set corresponding to this abnormal behavior group.
[0261] After identifying the abnormal behavior group, we can also lock the physical address and mobile phone number information of the object that engages in harmful behavior in the target Internet application by sorting and summarizing each node in the subgraph, so as to take management measures against the object that engages in harmful behavior and prevent it from continuing to engage in harmful behavior. Figure 12 As shown in the figure, based on a subgraph after partitioning, devices E1 through E4 represent devices used by a group to engage in harmful behavior; mobile phone numbers P1 through P5 represent mobile phone numbers used by a group to engage in harmful behavior; accounts a through d represent registered accounts within the group, and their corresponding IP addresses include A1 through A3; and accounts associated with these IP addresses include f and g. Analyzing this information reveals that the physical addresses of the anomalous group include K1, K2, and K3, and the mobile phone numbers include P6 and P7. Based on these physical addresses and mobile phone numbers, the anomalous group can be linked to curb their harmful behavior.
[0262] Furthermore, through the association between nodes in the subgraph, we can also explore the organizational information behind the abnormal behavior groups. Figure 13 As shown, the accounts in the abnormal behavior group are uniformly managed and allocated by organizational accounts M1, M2, and M3. Organization accounts M1 and M2 are registered and logged in at IP address A1, while organization account M3 is registered and logged in at IP address A2. Organization accounts M4, M5, and M6 can also be mined from IP addresses A1 and A2. These three accounts do not directly engage in harmful behavior in the target internet application, making them difficult to detect. However, by mining the account relationships in the subgraph layer by layer, it can be determined that these three accounts also belong to the abnormal behavior group. Organization account M4 is responsible for supplying and managing abnormal accounts to engage in harmful behavior, organization account M5 is responsible for providing technical support to the abnormal behavior group, and organization account M6 is responsible for providing assistance to the abnormal behavior group.
[0263] pass Figure 12 and Figure 13It can be seen that the data processing method based on the embodiment of the present disclosure can conduct in-depth mining of the information behind the abnormal behavior group after determining the abnormal behavior group, so as to fundamentally control the bad behavior occurring in the target Internet application, which is conducive to improving the stability and security of the network environment.
[0264] Using a graph construction method to find a set of abnormal accounts based on the second abnormal account is conducive to determining the group information of abnormal events, so as to centrally handle the abnormal behavior group, thereby ensuring that no abnormal account is missed, and improving the accuracy and comprehensiveness of abnormal account detection.
[0265] The embodiments of the present disclosure are described in detail in conjunction with specific application scenarios.
[0266] like Figure 14 FIG. 1 is another flow chart of the data processing method provided by the present disclosure. The method specifically includes the following steps:
[0267] Step 1410: Acquire multiple pieces of to-be-detected Internet data of multiple data dimensions of the target Internet application, where the data dimensions are obtained based on the basic attribute types of the Internet data.
[0268] Internet data can be traffic data generated in the target Internet application. The data dimensions can be divided based on the basic attribute types that generate the Internet data, which can specifically include: device, mobile phone number, account number and IP address, etc.
[0269] The internet data to be tested can be internet data generated by various data dimensions, such as device-related internet data, mobile phone number-related internet data, account-related internet data, and IP address-related internet data. Multiple pieces of internet data to be tested across multiple data dimensions can be obtained based on the backend system's recording of internet data within the target internet application.
[0270] Step 1420: Obtain reference Internet data, and screen the Internet data to be detected based on the reference Internet data and preset matching rules to obtain first abnormal Internet data, where the reference Internet data is Internet data with known detection results.
[0271] like Figure 15 As shown, reference internet data can come from both internal and external data. Internal data can be feedback from objects in the target internet application, while external data can be feedback from external management organizations or abnormal data from external applications. These reference internet data are internet data that have been preliminarily determined to be abnormal.
[0272] Preset matching rules can be manually set based on experience and are used to test the internet data to be tested. If the internet data to be tested matches the preset matching rules, it can be preliminarily determined to be an anomaly. For example, the number of registration requests from a single IP address within a predetermined time period exceeds a predetermined value.
[0273] When the Internet data to be detected is screened using reference Internet data and preset matching rules, if the Internet data to be detected matches the reference Internet data, then the Internet data to be detected can be determined as the first abnormal Internet data; if the Internet data to be detected can match the preset matching rules, then the Internet data to be detected can also be determined as the first abnormal Internet data.
[0274] Step 1430: Determine the target data dimension according to the target application scenario, obtain the environmental data of the first abnormal Internet data, and extract the content information and statistical information of each first abnormal Internet data based on the environmental data.
[0275] The target data attribute dimension can be the data attribute dimension that needs to be paid attention to when detecting abnormal accounts in the target application scenario. There may be correlations between Internet data of different data dimensions, such as Figure 15 As shown in the internet data relationships across multiple data attribute dimensions, account N1 is associated with device E1 and IP address A1, account N3 is associated with mobile phone number P1, IP address A1, IP address A2, and so on. Therefore, the context data for the first abnormal internet data can be obtained based on the relationship between the data attribute dimension to which the first abnormal internet data belongs and other data attribute dimensions. For example, if the first abnormal internet data is internet data from device E1, the corresponding context data can be the internet data from mobile phone number P1 and the internet data from account N1.
[0276] When extracting content and statistical information from the first abnormal internet data based on the environmental data, the first step is to obtain internet data related to the target data attribute dimension from the environmental data. For example, if the first abnormal internet data is related to the IP address dimension, the target data dimension is the account dimension. Therefore, information about the account registered with the IP address, the time the IP address was registered, and the device information used to register the IP address can be obtained. Content and statistical information can be obtained based on the internet data related to the target data attribute dimension.
[0277] The content information of the first abnormal internet data may be text information related to the first abnormal internet data, such as account registration information, device login information, business activity information, and account relationship information. The statistical information may be numerical information related to the first abnormal internet data, such as the number of account logins in a day, the account online duration, and the number of accounts logged in from the same IP address. The content information and statistical information of the first abnormal internet data can be extracted based on the information provided in the environmental data.
[0278] Step 1440: Screen the statistical information based on the policy rule set, and determine the first evaluation result according to the screening hit situation; extract and classify the content information based on the preset neural network model to obtain the second evaluation result.
[0279] The policy rule set is a set of preset rules that are strongly related to the target application scenario. The policy rules define the conditions under which behavioral information in the internet data is considered abnormal. When the statistical information matches the policy rules, the first evaluation result can determine that the statistical information in the first abnormal internet data is abnormal.
[0280] The content information can be evaluated for anomalies using a preset neural network. Before entering the preset neural network, the content information can be vectorized using embedding technology to generate a content feature vector. The concatenated content feature vectors are then input into the preset neural network for anomaly evaluation, generating a second evaluation result. The second evaluation result indicates whether the content information of the first abnormal internet data contains anomalies.
[0281] Step 1450: Determine a plurality of second abnormal Internet data in the first abnormal Internet data according to the first evaluation result and the second evaluation result, and determine a plurality of first abnormal accounts based on the plurality of second abnormal Internet data.
[0282] The first evaluation result can directly indicate whether the statistical information contains an anomaly, and the second evaluation result can directly indicate whether the content information contains an anomaly. Therefore, when at least one of the first evaluation result and the second evaluation result corresponding to the first abnormal internet data indicates an anomaly in the corresponding information, the first abnormal internet data can be determined as the second abnormal internet data.
[0283] The first evaluation result may also be a first anomaly score indicating an anomaly in the statistical information, and the second evaluation result may also be a second anomaly score indicating an anomaly in the content information. Therefore, based on the first and second evaluation results, multiple second internet data items are determined within the first anomaly internet data. A target score can be calculated based on the first and second anomaly scores, and multiple second anomaly internet data items are determined based on the target scores of the multiple first anomaly internet data items. Specifically, the multiple second anomaly internet data items can be determined based on a score threshold. When the target score is greater than or equal to the score threshold, the first anomaly internet data item is determined as the second anomaly internet data item.
[0284] After determining the plurality of second abnormal Internet data, account information corresponding to each second abnormal Internet data and the target Internet application may be obtained, and the plurality of first abnormal accounts may be determined based on the account information.
[0285] Step 1460: Obtain the time behavior sequence corresponding to each first abnormal account, and compare the time behavior sequence to obtain a first comparison result; perform time-frequency conversion on the time behavior sequence to obtain a spectrum behavior sequence; and compare the spectrum behavior sequence to obtain a second comparison result.
[0286] The time behavior sequence is a sequence that arranges the behaviors of the first abnormal account at different times in chronological order. Figure 15 As shown in the time behavior sequence in , the time behavior sequence of account a includes behaviors M1, M2, and M3. Specifically, the time behavior sequence can be compared with a reference time behavior sequence from a normal account for similarity, and a first comparison result can be obtained based on the similarity between the time behavior sequence and the reference time behavior sequence.
[0287] The spectrum behavior sequence obtained by performing time-frequency conversion on the time behavior sequence can be used to describe the periodicity and trend of the time behavior sequence in the frequency domain. To construct the spectrum behavior sequence corresponding to the first abnormal account based on the time behavior sequence, spectrum analysis methods can be used to convert the time behavior sequence data from the time domain to the frequency domain, generating a spectrum behavior sequence containing multiple frequency components. The spectrum behavior sequence can be compared by performing a similarity comparison with a reference spectrum behavior sequence from a normal account, and a second comparison result can be obtained based on the similarity comparison result.
[0288] Step 1470: Determine multiple second abnormal accounts from the multiple first abnormal accounts based on the first comparison result and the second comparison result.
[0289] The first comparison result can provide a conclusion on whether the temporal behavior sequence corresponding to the first abnormal account is abnormal. The second comparison result can provide a conclusion on whether the spectral behavior sequence corresponding to the first abnormal account is abnormal. Therefore, when at least one of the first comparison result and the second comparison result concludes that the corresponding behavior sequence is abnormal, the first abnormal account is determined to be the second abnormal account. When both the first comparison result and the second comparison result conclude that the corresponding behavior sequence is not abnormal, the first abnormal account can be determined to be a normal account.
[0290] Step 1480: Obtain multiple dimensions of environmental information associated with multiple second abnormal accounts; construct a first graph network based on the multiple second abnormal accounts and the multiple dimensions of environmental information; and divide the first graph network into subgraphs to obtain multiple abnormal account sets.
[0291] The multiple dimensions of environmental information associated with the second abnormal account may be information of other dimensions related to the second abnormal account, for example, the device used to log in to the second abnormal account, the mobile phone number and IP address associated with the second abnormal account, etc.
[0292] The multiple second abnormal accounts and the environmental information associated with each second abnormal account are used as multiple nodes, and the first graph network is constructed based on the relationship between each node. After obtaining the first graph network, the first graph network can be divided into multiple subgraphs according to the closeness of the node relationship. Specifically, the nodes with a higher degree of closeness of the node relationship can be divided into a subgraph. The closeness of the node relationship can be determined by calculating the association relationship between the nodes. Specifically, the Fast Unfolding algorithm can be used. When the account nodes contained in each subgraph are extracted, the abnormal account set corresponding to the subgraph can be formed. Figure 15 As shown in the abnormal account set in [ ], it includes account a, account b, account c, and account d. The accounts in the abnormal account set belong to a group, and based on the abnormal account set, we can further explore relevant information about the group. The data processing method of the disclosed embodiment ultimately obtains abnormal information, including abnormal account information and abnormal group information. Based on this abnormal information, we can fundamentally constrain abnormal behavior in the target Internet application.
[0293] The overall implementation of the data processing method of the embodiment of the present disclosure is specifically divided into four stages: Figure 16As shown, the following are: data acquisition stage, data processing stage, data analysis stage, and analysis result obtaining stage. In the data acquisition stage, multiple pieces of Internet data to be detected of multiple data dimensions of the target Internet application are obtained; in the data processing stage, reference Internet data is mined, and data screening is performed on the Internet data to be detected according to the reference Internet data and the preset matching rules; in the data analysis stage, a policy rule set and a neural network model are used to preliminarily analyze whether there are anomalies in the Internet data to be detected, and suspected abnormal accounts are obtained from the abnormal Internet data, and then the time behavior sequence and the spectrum behavior sequence are used to determine whether the suspected abnormal accounts are abnormal; in the analysis result obtaining stage, abnormal accounts can be obtained, and abnormal data information can be obtained based on the abnormal accounts. The abnormal data information can provide assistance for subsequent abnormal Internet data detection. Through the above process, the data processing method provided by the embodiment of the present disclosure obtains the Internet data to be detected based on a more comprehensive data attribute dimension, and uses policy model analysis and behavior sequence analysis to perform multiple detections on abnormal accounts, thereby improving the accuracy of abnormal account detection in Internet applications.
[0294] Description of the apparatus and device of the present disclosure
[0295] It is to be understood that, although the steps in the above-mentioned flowcharts are shown in sequence according to the arrow representations, these steps are not necessarily performed in sequence according to the order represented by the arrows. Unless otherwise specified in the present embodiment, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flowcharts may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0296] It should be noted that in each specific embodiment of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object, such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of such data will comply with the relevant laws, regulations and standards of the relevant region. In addition, when the embodiment of the present application needs to obtain the attribute information of the target object, the target object's separate permission or separate consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the target object's separate permission or separate consent, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.
[0297] Figure 17This is a schematic diagram of the structure of a data processing device 1700 provided in an embodiment of the present disclosure. The data processing device 1700 includes:
[0298] A first acquiring unit 1710 is configured to acquire to-be-detected Internet data associated with a target scenario in a target Internet application, where the to-be-detected Internet data includes Internet data collected from multiple data attribute dimensions;
[0299] A generating unit 1720 is configured to determine a target data attribute dimension from a plurality of data attribute dimensions based on a target scenario, and generate content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension;
[0300] A screening unit 1730 is configured to perform anomaly screening on each to-be-detected Internet data based on the content information and the statistical information, and to determine a plurality of first abnormal accounts in the target Internet application based on the anomaly screening results;
[0301] The second acquiring unit 1740 is configured to acquire a time behavior sequence corresponding to each first abnormal account;
[0302] The comparison unit 1750 is configured to compare the behavior sequences of the plurality of first abnormal accounts based on the time behavior sequence, and determine a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result.
[0303] Optionally, in some embodiments, the screening unit 1730 is specifically configured to:
[0304] Perform anomaly matching on statistical information according to preset strategy rules to obtain a first screening result;
[0305] Classify the content information as abnormal based on a preset neural network model to obtain a second screening result;
[0306] Determining a plurality of abnormal Internet data from the plurality of Internet data to be detected according to the first screening result and the second screening result;
[0307] A plurality of first abnormal accounts in a target Internet application is determined based on the plurality of abnormal Internet data.
[0308] Optionally, in some embodiments, the screening unit 1730 is specifically configured to:
[0309] Obtaining a first weight coefficient corresponding to the first screening result, and obtaining a second weight coefficient corresponding to the second screening result;
[0310] Performing weighted calculation on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result;
[0311] According to the target screening results, multiple abnormal Internet data are determined from multiple Internet data to be detected.
[0312] Optionally, in some embodiments, the screening unit 1730 is specifically configured to:
[0313] determining a first score corresponding to the first screening result and a second score corresponding to the second screening result;
[0314] Performing weighted calculation on the first score and the second score based on the first weight coefficient and the second weight coefficient to obtain a target score;
[0315] The target screening results are determined based on the target scores.
[0316] Optionally, in some embodiments, the screening unit 1730 is specifically configured to:
[0317] Obtain the account information corresponding to each abnormal Internet data and the target Internet application;
[0318] A plurality of first abnormal accounts in the target Internet application is determined based on the account information.
[0319] Optionally, in some embodiments, the comparison unit 1750 is specifically configured to:
[0320] Get the reference time behavior sequence;
[0321] Calculate the similarity between the time behavior sequence and the reference time behavior sequence to obtain the comparison result;
[0322] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison result.
[0323] Optionally, in some embodiments, the comparison unit 1750 is specifically configured to:
[0324] Extract features from the time behavior sequence based on a preset sequence feature generation model to obtain the first sequence feature;
[0325] Extract features from the reference time behavior sequence based on a preset sequence feature generation model to obtain a second sequence feature;
[0326] Calculate the cosine similarity between the first sequence feature and the second sequence feature, and determine the alignment result based on the cosine similarity.
[0327] Optionally, in some embodiments, the data processing device 1700 provided by the present application further includes:
[0328] A first constructing unit (not shown) is configured to construct a spectrum behavior sequence corresponding to each first abnormal account according to the time behavior sequence;
[0329] The comparison unit 1750 is specifically used for:
[0330] A behavior sequence comparison is performed on the plurality of first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence, and a plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison result.
[0331] Optionally, in some embodiments, the comparison unit 1750 is specifically configured to:
[0332] Obtaining a reference time behavior sequence and a reference spectrum behavior sequence;
[0333] Determining a first sub-matching result based on the similarity between the time behavior sequence and the reference time behavior sequence, and determining a second sub-matching result based on the similarity between the spectrum behavior sequence and the reference spectrum behavior sequence;
[0334] A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the first sub-comparison result and the second sub-comparison result.
[0335] Optionally, in some embodiments, the data processing device 1700 provided by the present application further includes:
[0336] a determining unit (not shown), configured to determine a plurality of abnormal data attribute nodes corresponding to the plurality of second abnormal accounts, and associated data attribute nodes corresponding to the plurality of abnormal data attribute nodes;
[0337] A second construction unit (not shown) is configured to construct a graph network using a plurality of second abnormal account numbers, a plurality of abnormal data attribute nodes, and associated data attribute nodes;
[0338] A partitioning unit (not shown) is used to partition the graph network into multiple subgraphs based on the closeness of the node relationships in the graph network, and determine multiple abnormal account sets based on the multiple subgraphs.
[0339] Optionally, in some embodiments, the data processing device 1700 provided by the present application further includes:
[0340] A third obtaining unit (not shown), configured to obtain abnormal feedback information of Internet data in multiple data attribute dimensions of the target Internet application;
[0341] The screening unit (not shown) is used to screen out the Internet data to be detected from the application data of the target Internet application based on the abnormal feedback information.
[0342] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0343] Reference Figure 18 , Figure 18 The following is a block diagram of the structure of a terminal 140 for implementing the data processing method according to an embodiment of the present disclosure. The terminal 140 includes: a radio frequency (RF) circuit 1410, a memory 1815, an input unit 1830, a display unit 1840, a sensor 1850, an audio circuit 1860, a wireless fidelity (WiFi) module 1870, a processor 1880, and a power supply 1890. It will be understood by those skilled in the art that Figure 18 The illustrated structure of the terminal 140 does not limit the structure of a mobile phone or a computer, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0344] The RF circuit 1810 may be used for receiving and sending signals during information transmission or calls. In particular, after receiving downlink information from the base station, it is sent to the processor 1880 for processing. In addition, the designed uplink data is sent to the base station.
[0345] The memory 1815 may be used to store software programs and modules. The processor 1880 executes various functional applications and document editing of the terminal by running the software programs and modules stored in the memory 1815 .
[0346] The input unit 1830 may be configured to receive input digital or character information and generate key signal input related to terminal settings and function control. Specifically, the input unit 1830 may include a touch panel 1831 and other input devices 1832 .
[0347] The display unit 1840 may be configured to display input information or provided information and various menus of the terminal. The display unit 1840 may include a display panel 1841 .
[0348] The audio circuit 1860 , the speaker 1861 , and the microphone 1862 may provide an audio interface.
[0349] In this embodiment, the processor 1880 included in the terminal 140 can execute the data processing method of the previous embodiment.
[0350] The terminal 140 in the embodiment of the present disclosure includes but is not limited to a mobile phone, a computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.
[0351] Figure 19 A structural block diagram of a portion of a server 110 for implementing the data processing method of an embodiment of the present disclosure. The server 110 may vary greatly due to different configurations or performance, and may include one or more central processing units (CPUs) 1922 (for example, one or more processors) and a storage device 1932, and one or more storage media 1930 (for example, one or more mass storage devices) for storing application programs 1942 or data 1944. The storage device 1932 and the storage medium 1930 may be temporary storage or permanent storage. The program stored in the storage medium 1930 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server 110. Furthermore, the central processing unit 1922 may be configured to communicate with the storage medium 1930 to execute a series of instruction operations in the storage medium 1930 on the server 110.
[0352] The server 110 may also include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input and output interfaces 1958, and / or one or more operating systems 1941, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0353] The central processing unit 1922 in the server 110 may be configured to execute the data processing method of the embodiment of the present disclosure.
[0354] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the data processing methods of the aforementioned embodiments.
[0355] The present disclosure also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned data processing method.
[0356] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0357] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0358] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.
[0359] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0360] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0361] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0362] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned computer-readable storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program codes.
[0363] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0364] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire internet data to be detected associated with a target scenario in a target internet application, wherein the internet data to be detected includes internet data collected from multiple data attribute dimensions; Determining a target data attribute dimension from the multiple data attribute dimensions based on the target scenario, and generating content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension; Performing anomaly screening on each to-be-detected Internet data according to the content information and the statistical information, and determining a plurality of first abnormal accounts in the target Internet application according to the anomaly screening results; Obtaining a time behavior sequence corresponding to each of the first abnormal accounts; A behavior sequence comparison is performed on the multiple first abnormal accounts based on the time behavior sequence, and multiple second abnormal accounts are determined from the multiple first abnormal accounts according to the comparison result.
2. The method according to claim 1, characterized in that The performing anomaly screening on each to-be-detected Internet data according to the content information and the statistical information, and determining a plurality of first abnormal accounts in the target Internet application according to the anomaly screening results, includes: Performing anomaly matching on the statistical information according to preset strategy rules to obtain a first screening result; Performing abnormality classification on the content information based on a preset neural network model to obtain a second screening result; determining a plurality of abnormal Internet data from the plurality of Internet data to be detected according to the first screening result and the second screening result; A plurality of first abnormal accounts in the target Internet application is determined according to the plurality of abnormal Internet data.
3. The method according to claim 2, characterized in that The determining of a plurality of abnormal Internet data from the plurality of Internet data to be detected according to the first screening result and the second screening result includes: Obtaining a first weight coefficient corresponding to the first screening result, and obtaining a second weight coefficient corresponding to the second screening result; Performing weighted calculation on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result; A plurality of abnormal Internet data are determined from the plurality of Internet data to be detected according to the target screening result.
4. The method according to claim 3, characterized in that The performing weighted calculation on the first screening result and the second screening result based on the first weight coefficient and the second weight coefficient to obtain a target screening result includes: determining a first score corresponding to the first screening result and a second score corresponding to the second screening result; Performing weighted calculation on the first score and the second score based on the first weight coefficient and the second weight coefficient to obtain a target score; A target screening result is determined according to the target score.
5. The method according to claim 2, characterized in that The determining, based on the plurality of abnormal Internet data, a plurality of first abnormal accounts in the target Internet application includes: Obtaining account information corresponding to each abnormal Internet data and the target Internet application; A plurality of first abnormal accounts in the target Internet application is determined according to the account information.
6. The method according to claim 1, wherein The performing behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence, and determining a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result, includes: Get the reference time behavior sequence; Calculating the similarity between the time behavior sequence and the reference time behavior sequence to obtain a comparison result; A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the comparison result.
7. The method according to claim 6, characterized in that The calculating the similarity between the time behavior sequence and the reference time behavior sequence to obtain a comparison result includes: Extracting features from the time behavior sequence based on a preset sequence feature generation model to obtain a first sequence feature; Extracting features from the reference time behavior sequence based on the preset sequence feature generation model to obtain a second sequence feature; Calculate the cosine similarity between the first sequence feature and the second sequence feature, and determine the alignment result according to the cosine similarity.
8. The method according to claim 1, characterized in that After obtaining the time behavior sequence corresponding to each of the first abnormal accounts, the method further includes: Constructing a spectrum behavior sequence corresponding to each of the first abnormal accounts according to the time behavior sequence; The performing behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence, and determining a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result, includes: A behavior sequence comparison is performed on the multiple first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence, and multiple second abnormal accounts are determined from the multiple first abnormal accounts according to the comparison result.
9. The method according to claim 8, characterized in that The performing behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence and the spectrum behavior sequence, and determining a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result, includes: Obtaining a reference time behavior sequence and a reference spectrum behavior sequence; Determining a first sub-comparison result based on the similarity between the time behavior sequence and the reference time behavior sequence, and determining a second sub-comparison result based on the similarity between the spectrum behavior sequence and the reference spectrum behavior sequence; A plurality of second abnormal accounts are determined from the plurality of first abnormal accounts according to the first sub-comparison result and the second sub-comparison result.
10. The method according to claim 1, characterized in that After comparing the behavior sequences of the multiple first abnormal accounts based on the time behavior sequence and determining multiple second abnormal accounts from the multiple first abnormal accounts according to the comparison results, the method further includes: Determining a plurality of abnormal data attribute nodes corresponding to the plurality of second abnormal accounts, and associated data attribute nodes corresponding to the plurality of abnormal data attribute nodes; Constructing a graph network with the plurality of second abnormal accounts, the plurality of abnormal data attribute nodes, and the associated data attribute nodes; The graph network is divided into a plurality of subgraphs based on the closeness of the node relationships in the graph network, and a plurality of abnormal account sets are determined based on the plurality of subgraphs.
11. The method according to claim 1, wherein Before obtaining the to-be-detected Internet data associated with the target scenario in the target Internet application, the method further includes: Obtaining abnormal feedback information on Internet data across multiple data attribute dimensions in target Internet applications; Internet data to be detected is screened out from application data of the target Internet application based on the abnormal feedback information.
12. A data processing device, characterized in that: The device comprises: A first acquisition unit is configured to acquire to-be-detected Internet data associated with a target scenario in a target Internet application, wherein the to-be-detected Internet data includes Internet data collected from multiple data attribute dimensions; a generating unit, configured to determine a target data attribute dimension from the plurality of data attribute dimensions based on the target scenario, and generate content information and statistical information of each to-be-detected Internet data based on the target data attribute dimension; a screening unit, configured to perform anomaly screening on each to-be-detected Internet data based on the content information and the statistical information, and determine a plurality of first abnormal accounts in the target Internet application based on the anomaly screening results; a second acquiring unit, configured to acquire a time behavior sequence corresponding to each of the first abnormal accounts; The comparison unit is configured to perform a behavior sequence comparison on the plurality of first abnormal accounts based on the time behavior sequence, and determine a plurality of second abnormal accounts from the plurality of first abnormal accounts according to the comparison result.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 11 is implemented.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the data processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the data processing method according to any one of claims 1 to 11.