User behavior risk identification method and device, equipment and medium
By performing structured analysis on mouse trajectory data, constructing a multi-dimensional dataset, and integrating page context and user environment data, multi-dimensional behavioral features are extracted, solving the problem of data isolation in user behavior risk identification and achieving higher identification accuracy and stability.
Patent Information
- Application Number
- CN202610524100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-14
AI Technical Summary
Existing technologies suffer from isolated data dimensions and missing correlations in user behavior risk identification, which limits the accuracy and stability of risk identification results.
By performing structured analysis on mouse trajectory data, a multi-dimensional dataset is constructed. The relationships between mouse behavior data, page context data, and user environment data are integrated to extract multi-dimensional behavioral features for risk identification.
It improves the accuracy and reliability of risk identification for user interaction behavior, and can more accurately distinguish between the behavior patterns of humans and machines, solving the problems of unstable identification results and limited accuracy caused by data isolation.
Smart Images

Figure CN122394883A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network information security technology, and in particular to a method, apparatus, device, and medium for identifying user behavior risks. Background Technology
[0002] In internet applications, user behavior risk identification is an important means of ensuring business security and preventing fraud attacks. It is often used to determine whether the operator is a real person by analyzing the continuous behavioral data generated by the user during the interaction process.
[0003] Existing methods typically extract statistical features from single-dimensional behavioral sequences and process these features as independent information. However, behavioral data in actual interactions often originates from multiple sources and is intrinsically linked to the scene state and environmental parameters during the interaction, but current technologies lack a modeling framework for these cross-dimensional relationships. This fragmented processing of behavioral data limits feature extraction to isolated data sources, making it difficult to form a comprehensive representation of the complete context of user behavior.
[0004] In summary, existing technologies for identifying user behavior risks suffer from isolated data dimensions and a lack of correlation, which limits the accuracy and stability of risk identification results. Summary of the Invention
[0005] The purpose of this application is to solve the above-mentioned problems by providing a method, apparatus, device, and medium for identifying user behavior risks.
[0006] According to one aspect of this application, a method for identifying user behavior risks is provided, comprising the following steps: Acquire mouse trajectory data generated by the target user during webpage interaction; The mouse trajectory data is subjected to structured parsing to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the relationship between mouse behavior data, page context data, and user environment data. Multidimensional behavioral features are extracted from the multidimensional dataset, and the multidimensional behavioral features characterize the target user's behavioral characteristics in different dimensions during web page interaction; Based on the multidimensional behavioral characteristics, risk identification is performed on the interaction behavior of the target user to determine the corresponding risk type of the target user.
[0007] According to another aspect of this application, a user behavior risk identification device is provided, comprising: The trajectory acquisition module is used to acquire mouse trajectory data generated by the target user during web page interaction; The parsing and association module is used to perform structured parsing on the mouse trajectory data to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the association between mouse behavior data, page context data, and user environment data. The feature extraction module is used to extract multi-dimensional behavioral features from the multi-dimensional dataset, wherein the multi-dimensional behavioral features characterize the target user's behavioral features in different dimensions during webpage interaction; The risk identification module is used to identify risks in the interaction behavior of the target user based on the multidimensional behavioral characteristics, so as to determine the risk type corresponding to the target user.
[0008] According to another aspect of this application, an electronic device is provided, including a central processing unit and a memory, wherein the central processing unit is configured to invoke and run a computer program stored in the memory to perform the steps of the user behavior risk identification method described in this application.
[0009] According to another aspect of this application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the user behavior risk identification method in the form of computer-readable instructions, wherein the computer program, when invoked by a computer, executes the steps included in the method.
[0010] Compared to existing technologies, firstly, by performing structured analysis on mouse trajectory data, the previously isolated recorded mouse behavior data, page context data, and user environment data are associated and mapped to construct a multi-dimensional data set with a standardized structure. This data set solves the problem of isolated data dimensions in existing technologies, enabling each mouse trajectory coordinate to correspond to a specific page interaction intent and physical environment constraint. This provides a data foundation with a complete semantic background for subsequent feature extraction, ensuring that the extracted features are no longer abstract coordinate sequences, but carry contextual information such as "in what scenario and in what environment" the operation takes place. Secondly, the multidimensional behavioral features extracted from the aforementioned multidimensional dataset, by incorporating page context and user environment parameters, can characterize the degree to which the target user's behavioral patterns conform to physical laws under specific interaction constraints. For example, combining page context can distinguish between natural pauses and aimless mechanical drifts while reading content, and combining user environment data can determine whether mouse movement acceleration conforms to the physical characteristics of a real input device. Compared to the statistical features of existing technologies that rely solely on a single trajectory sequence, the multidimensional behavioral features extracted in this application significantly enhance the distinguishability between humans and machines. Therefore, when risk identification is performed based on such multidimensional behavioral features, it is possible to more accurately determine whether the target user's interactive behavior belongs to an abnormal risk type, ultimately solving the technical problem of unstable identification results and limited accuracy due to missing data associations. Attached Figure Description
[0011] Figure 1 This is an exemplary network architecture suitable for applying the user behavior risk identification method of this application; Figure 2 This is a flowchart illustrating one embodiment of the user behavior risk identification method of this application. Figure 3 This is a schematic block diagram of the user behavior risk identification device of this application; Figure 4 This is a schematic diagram of the structure of an electronic device used in this application. Detailed Implementation
[0012] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0013] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0014] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0015] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDAs (Personal Digital Assistants) that may include radio frequency receivers, pagers, internet / intranet access, web browsers, notepads, calendars, and / or GPS (Global Positioning System) receivers; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0016] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0017] It should be noted that the concept of "server" used in this application can also be extended to apply to server clusters. Based on network deployment principles as understood by those skilled in the art, the servers should be logically divided; physically, these servers can be independent yet accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method described in this application.
[0018] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client for access.
[0019] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0020] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0021] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0022] To facilitate understanding of the various embodiments of this application, exemplary network architectures and application scenarios will be introduced first.
[0023] like Figure 1 As shown, the network architecture of this application aims to build a user behavior risk identification system. The network architecture includes a client 80 and a server 81, wherein the client 80 and the server 81 are connected via the Internet.
[0024] Specifically, client 80 responds to the actions generated by the target user during webpage interaction, collects mouse trajectory data, and sends the mouse trajectory data to server 81. Server 81 receives the mouse trajectory data, performs structured parsing on the mouse trajectory data, and obtains a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the correlation between mouse behavior data, page context data, and user environment data. Server 81 is also used to extract multi-dimensional behavioral features from the multi-dimensional data set. The multi-dimensional behavioral features represent the target user's behavioral characteristics in different dimensions during webpage interaction. Server 81 is also used to identify risks in the target user's interactive behavior based on the multi-dimensional behavioral features to determine the corresponding risk type of the target user.
[0025] Accordingly, the network architecture of this application realizes distributed collection and centralized structured analysis of behavioral data through the collaborative interaction between client 80 and server 81, ensuring the standardization and integrity of data during transmission and processing, thereby improving the accuracy and reliability of risk identification of user interaction behavior.
[0026] In one exemplary application scenario, the user behavior risk identification method of this application can be applied to various Internet scenarios that require user authentication or behavior security monitoring, such as e-commerce platform independent websites and online financial transactions. The following explanation uses an e-commerce platform independent website as an example: In the context of independent e-commerce platforms, users need to perform key operations such as accessing product detail pages, logging into their accounts, adding items to their shopping carts, or making payments through a client (e.g., a web browser). Because independent websites have diverse page structures that may dynamically adjust with marketing campaigns, traditional risk identification methods often struggle to adapt to different page layouts, leading to insufficient accuracy in distinguishing between normal user actions and machine script attacks.
[0027] When applying the solution of this application, when a user interacts on the client, the client is responsible for collecting mouse trajectory data and sending it to the server. After receiving the data, the server does not make judgments based directly on isolated data points, but first performs structured parsing of the data, associating the user's operation behavior with the current page context (such as the position information of function buttons) and the operating environment, constructing a multi-dimensional data set with a standardized structure. Subsequently, based on this association, the server extracts multi-dimensional behavioral features that can represent the user's actual operating habits, identifies whether the interaction behavior is normal or risky, and determines the corresponding risk type.
[0028] This solution enables the server to accurately determine the type of risk based on standardized, correlated data, even when the page structure changes or the environment fluctuates. This not only ensures account security during e-commerce transactions and effectively defends against automated attacks, but also avoids unnecessary interference with legitimate users, thus improving the user experience.
[0029] Through the detailed description of the network architecture and application scenarios described above, a better understanding of the specific application of the user behavior risk management method of this application in the field of network information security technology can be achieved. The following will elaborate on the detailed description of various specific embodiments of this application based on these exemplary contents.
[0030] like Figure 2 As shown, in one embodiment, the user behavior risk identification method of this application includes: Step S5100: Obtain mouse trajectory data generated by the target user during webpage interaction.
[0031] The client responds to the target user's actions on the webpage, collecting event information during mouse movement and encapsulating it into mouse trajectory data. The server receives the mouse trajectory data sent by the client. This mouse trajectory data records the target user's operation path and time sequence during webpage interaction, serving as the foundational data source for subsequent structured analysis and risk identification.
[0032] Mouse trajectory data collected by the client can be temporarily stored in the client's local cache, and when preset trigger conditions are met, such as when the user completes a specific page operation or when the data sending interval is reached, the mouse trajectory data is sent to the server through a network request.
[0033] The server can listen for network requests through a preset data receiving interface, parse the request body, and extract mouse trajectory data.
[0034] Step S5200: Perform structured parsing on the mouse trajectory data to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the relationship between mouse behavior data, page context data, and user environment data.
[0035] The server performs structured parsing on mouse trajectory data to obtain a multi-dimensional dataset with a standardized structure. The structured parsing process reorganizes the raw data into a standardized format based on its inherent organizational form. This multi-dimensional dataset represents the relationships between mouse behavior data, page context data, and user environment data, establishing logical connections between user actions, interface content, and system state. Through structured parsing, discrete data is transformed into a standard data body with semantic relationships.
[0036] When the client encapsulates event information into mouse trajectory data, it can do so according to a standardized structure, using a layered nested structure. For example, the first logical layer is metadata, the second logical layer is mouse behavior data, the third logical layer is page context data, and the fourth logical layer is user environment data.
[0037] Standardized structure refers to data following predefined nested format specifications; multi-dimensional datasets refer to data carriers that integrate multiple types of information; mouse behavior data layer records operation trajectories; page context data layer records interface information; user environment data layer records system state; and relationships refer to the connections established between different data dimensions based on interaction logic, used to represent the context in which behavior occurs. The server can parse the internal structure of mouse trajectory data according to the mouse trajectory data format specification, separate behavior, page, and environment-related information, and establish logical connections between different dimensions of information using the time or space attributes contained in the data. Through integration processing, a multi-dimensional data set containing correlations is generated. The data integration process ensures the integrity of data records within the same session.
[0038] Mouse trajectory data can be formatted in various ways, as long as it has a structured nested feature.
[0039] Step S5300: Extract multidimensional behavioral features from the multidimensional dataset. These multidimensional behavioral features characterize the target user's behavioral characteristics in different dimensions during webpage interaction.
[0040] The server extracts multidimensional behavioral features from a multidimensional dataset. The extraction process is based on mouse behavior data, page context data, and user environment data contained in the multidimensional dataset, as well as the relationships between them. These multidimensional behavioral features characterize the different dimensions of a target user's behavior during webpage interaction. By utilizing the relationships between data points for feature calculation, the raw data is transformed into quantitative indicators that reflect user operating habits and behavioral patterns.
[0041] Extracting multidimensional behavioral features from multidimensional datasets allows us to leverage the relationships between data to capture the interaction characteristics between user actions and the page environment and system state. Compared to features extracted solely from raw trajectory data, features extracted from correlated data possess richer semantic information and greater discriminative power, more accurately reflecting the biomechanical characteristics and operational intentions of human behavior.
[0042] The specific algorithm for multidimensional behavioral feature extraction can be adjusted according to business needs, including but not limited to statistical calculation, physical model calculation, frequency domain analysis, or entropy calculation.
[0043] Step S5400: Based on multidimensional behavioral characteristics, identify risks in the target user's interactive behavior to determine the corresponding risk type for the target user.
[0044] The server identifies risks in target users' interactive behaviors based on multidimensional behavioral characteristics to determine the corresponding risk type. By utilizing multidimensional behavioral characteristics for identification, quantitative indicators are transformed into security assessment conclusions, ensuring that the risk determination results are supported by data.
[0045] Risk identification refers to the process of assessing the security of interactive behaviors. Risk type refers to the classification of the security level of a behavior, used to distinguish between normal operation and potential threats. Multidimensional behavioral features contain various aspects of user operations, reflecting the operating habits and intentions behind the behavior. Identification based on multidimensional behavioral features can comprehensively utilize information from motion, space, and time dimensions, improving the accuracy of the judgment.
[0046] The server can input multidimensional behavioral features into a pre-trained risk identification model, output a risk score, and determine the risk type based on the risk score. For example, different score ranges can be set to correspond to different risk levels; when the score exceeds a preset threshold, the target user is identified as a high-risk user. Alternatively, a rule engine can match multidimensional behavioral features with known risk patterns to directly output the risk type classification result.
[0047] As can be seen from the above embodiments, this application achieves the following outstanding effects: First, by performing structured analysis on mouse trajectory data, a multi-dimensional dataset with a standardized structure is constructed, establishing the correlation between mouse behavior data, page context data, and user environment data. This transforms discrete operation records into a correlated data body with spatiotemporal consistency, solving the problems of isolated data dimensions and lack of semantic context in existing data, and providing a complete data foundation for risk identification.
[0048] Secondly, by extracting multidimensional behavioral features based on the correlation in multidimensional datasets, it can integrate the interaction information between operational behavior, interface content, and system status, capture the biomechanical characteristics and operational intentions of human behavior, solve the problem that existing feature extraction relies only on a single trajectory dimension and has insufficient discriminative power, and significantly improve the feature's resistance to simulation by automated scripts.
[0049] Furthermore, risk identification and risk type determination based on multi-dimensional behavioral characteristics enable the risk judgment logic to be based on comprehensive and in-depth behavioral information, solving the problem of insufficient accuracy in existing risk identification. Even with dynamic changes in page structure or fluctuations in the network environment, the stability and reliability of risk judgment can still be maintained.
[0050] Based on any embodiment of the method in this application, the mouse trajectory data is subjected to structured parsing to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the relationship between mouse behavior data, page context data, and user environment data, including: Step S5210: Identify the standardized structural framework of mouse trajectory data and determine the nesting level of mouse trajectory data.
[0051] The server identifies a standardized structural framework for mouse trajectory data and determines the nesting level of the data. The standardized structural framework is a predefined format specification that the data follows, and the nesting level is the hierarchical distribution relationship within the data. By identifying the structural framework and nesting level, a structural basis is provided for subsequent data extraction.
[0052] The identification process aims to understand the data organization in order to accurately separate information from different dimensions.
[0053] The server can read the header information or metadata fields of mouse trajectory data, identify the data format type, match the key-value pair distribution in the data according to a predefined structure template, analyze the containment relationship of data objects, and determine the hierarchical distribution of root nodes, child nodes, and leaf nodes. For example, by parsing the JavaScript object representation format, it can identify top-level fields and nested objects, verify whether the data conforms to the expected structural specifications, and confirm the validity of the nesting level.
[0054] The data format is not limited to JavaScript object representation; other structured formats can also be used.
[0055] The hierarchy can be determined through recursive traversal or schema verification.
[0056] Step S5220: Extract mouse behavior data, page context data, and user environment data from the mouse trajectory data according to the nesting level.
[0057] The server extracts mouse behavior data, page context data, and user environment data from mouse trajectory data based on the nesting hierarchy. The extraction operation maps data fields at different levels to corresponding data categories based on path identifiers in the data structure. The nesting hierarchy corresponds to different business information dimensions: mouse behavior data records operation trajectories, page context data records interface information, and user environment data records system status.
[0058] The server can traverse the object tree of mouse trajectory data, identify specific key names or path markers, and classify the matched data blocks into mouse behavior data, page context data, and user environment data. For mouse behavior data, the server can extract trajectory point sequences and fine-grained trajectories of functional areas. The trajectory point sequence includes the x-coordinate, y-coordinate, event occurrence time, and event type identifier of the trajectory points. Event types include movement, click, and hover. The fine-grained trajectory of functional areas records the user's operation sequence near specific functional elements. For page context data, the server can extract data including the height and width of the visible area, the top-left corner of the target button, and the... The system includes bottom-right corner coordinates, button geometric center coordinates, button semantic identifiers, and scroll bar position change records. The visible area size is used to define the coordinate space, button coordinates are used for spatial mapping, semantic identifiers are used to distinguish multi-button scenarios, and scroll records are used to analyze the temporal relationship between page scrolling and mouse movement. For user environment data, the server can extract browser identifiers, total page height and width, the interval between page interactivity and the first mouse movement, and page loading performance metrics. Browser identifiers are used to identify client type, total page size is used to calculate relative position, the interval reflects user reading or thinking time, and performance metrics are used to identify abnormal environments.
[0059] The server can also classify the matched data blocks into metadata. For metadata, the server can extract data format version number, data type identifier, anti-tamper signature, and base timestamp of data packet generation. The version number is used to identify the data structure version, the data type identifier is used to distinguish business scenarios, the anti-tamper signature is used for integrity verification, and the base timestamp of data packet generation serves as the benchmark for time alignment.
[0060] Step S5230: Associate and integrate mouse behavior data, page context data, and user environment data to obtain a multi-dimensional data set.
[0061] The server associates, binds, and integrates mouse behavior data, page context data, and user environment data to obtain a multi-dimensional data set. Logical connections are established between data layers through association and binding, and the integration process merges the scattered data layers into a unified data body with spatiotemporal relationships.
[0062] Association binding refers to the process of establishing corresponding relationships between different data dimensions. Integration refers to storing the associated data into a unified structure. Mouse behavior data, page context data, and user environment data are initially independent after extraction. Association binding solves the problem of data isolation and gives trajectory data a spatiotemporal context.
[0063] The association and binding method can be based on timestamp and coordinate mapping, or on sequence number matching or relative position identifier, as long as the logical correspondence between data can be established.
[0064] Integration can be achieved through database association, memory object reference, or file merging, as long as a unified data body can be formed.
[0065] As can be seen from the above embodiments, this application achieves the effect of standardizing data organization and eliminating data format chaos by identifying the standardized structural framework of mouse trajectory data and determining the nesting level; by extracting mouse behavior data, page context data, and user environment data according to the nesting level, it achieves the effect of separating coupled data into business data blocks with clear semantics; by associating and binding multiple types of data and integrating them to obtain a multi-dimensional data set, it achieves the effect of establishing spatiotemporal correlations between behavior, page, and environment information and solving the problem of data isolation; on this basis, subsequent feature extraction can be based on a complete data foundation with semantic context, rather than relying on isolated data points, thereby significantly improving the accuracy and reliability of user interaction behavior risk identification.
[0066] Based on any embodiment of the method in this application, mouse behavior data, page context data, and user environment data are associated, bound, and integrated to obtain a multi-dimensional data set, including: Step S5310: Time-align mouse behavior data and user environment data based on a preset reference timestamp to establish a temporal correlation between mouse behavior data and user environment data.
[0067] Mouse behavior data and user environment data are time-aligned based on a preset baseline timestamp to establish a temporal correlation between them. The time alignment operation unifies time information from different data sources to the same time baseline. The temporal correlation represents the logical correspondence between mouse behavior events and user environment states in the time dimension.
[0068] Time alignment refers to the process of eliminating time reference differences between different data modules. Mouse behavior data includes the event occurrence time, and user environment data includes environmental state time parameters. Because data acquisition times may differ or relative time counting may be used, direct comparison can lead to logical errors. A reference timestamp serves as a unified benchmark, used to calibrate the time information of each data layer. After establishing time correlations, the server can determine mouse operations occurring under specific environmental conditions, providing a spatiotemporally consistent data foundation for subsequent feature calculations.
[0069] The server can extract the baseline timestamp generated from the data packet from the metadata, read the event occurrence time of each trajectory point in the mouse behavior data, calculate the time difference between the event occurrence time and the baseline timestamp, or calibrate the event occurrence time to the absolute time series of the baseline timestamp. It can also read time parameters from the user environment data, such as the interval between page interactivity and the first mouse movement or page load performance metrics. Based on the baseline timestamp and time parameters, the server calculates the absolute time point when the user environment state takes effect and matches the calibrated mouse behavior time with the user environment state time. If the mouse behavior event occurs within the time range when the user environment state takes effect, a time correlation is established. The server can record these time correlations in a multi-dimensional dataset.
[0070] Time difference calculation can be performed with millisecond-level precision, and time matching can be performed by comparing time windows or judging by the equality of timestamps.
[0071] The baseline timestamp is not limited to metadata extraction; it can also be the session start timestamp or the server reception timestamp, as long as it can serve as a unified baseline.
[0072] Step S5320: Extract the visible area size and interface element coordinates from the page context data, and map the mouse behavior data to the coordinate space defined by the visible area size to establish a spatial relationship between the mouse behavior data and the interface element coordinates.
[0073] The server extracts the visible area size and UI element coordinates from the page context data, and maps the mouse behavior data to the coordinate space defined by the visible area size to establish a spatial relationship between the mouse behavior data and the UI element coordinates. This spatial mapping operation incorporates the mouse behavior data into a unified spatial reference system, and the spatial relationship represents the geometric correspondence between the mouse operation position and the position of the UI functional element.
[0074] The visible area size defines the effective interactive range within the browser window and serves as the boundary of the coordinate space. Interface element coordinates identify the specific location of functional components on the page and act as the target reference point for spatial association. Mouse behavior data contains spatial location information of the operation trajectory. Due to differences in screen resolution and page layout across different client devices, the raw coordinate data lacks a unified spatial semantics. By mapping to the coordinate space defined by the visible area size, spatial deviations caused by device differences are eliminated. After establishing spatial associations, the server can determine whether the mouse trajectory is targeting a specific interface element.
[0075] The server can read the visible area height and visible area width fields from the page context data, read the coordinate fields of interface elements, such as the coordinates of the top left and bottom right corners of a button, or the geometric center coordinates of a button, and read the horizontal and vertical coordinates of the trajectory points from the mouse behavior data.
[0076] The server can determine whether the x and y coordinates of the trajectory point are within the range defined by the height and width of the visible area. It calculates the Euclidean distance or containment relationship between the trajectory point and the coordinates of the interface elements. If the trajectory point is within the rectangular area defined by the coordinates of the interface elements, or if the distance between the trajectory point and the geometric center coordinates of the button is less than a preset pixel threshold, a spatial relationship can be determined, and this spatial relationship can be marked in a multi-dimensional data set, such as by adding a spatial relationship identifier field. The origin of the coordinate space can be set to the top left corner of the page or the top left corner of the visible area.
[0077] Step S5330: Based on time and spatial relationships, integrate mouse behavior data, page context data, and user environment data to obtain a multi-dimensional data set.
[0078] The server integrates mouse behavior data, page context data, and user environment data based on temporal and spatial relationships to obtain a multi-dimensional data set. This integration process merges spatiotemporally related data layers into a unified data structure. The multi-dimensional data set includes mouse behavior data, page context data, and user environment data, along with their associated identifiers. Through integration, a data object with complete semantic context is formed.
[0079] After alignment and mapping, mouse behavior data, page context data, and user environment data still need to be integrated into structured objects. Multi-dimensional datasets are the fundamental carriers for feature extraction, and the integrity of their internal structure directly affects the accuracy of feature calculation. Integration based on relationships ensures that the data maintains spatiotemporal consistency in subsequent processing, avoiding feature distortion caused by data fragmentation.
[0080] The server can create data structure objects that aggregate multi-dimensional data. Time-aligned mouse behavior data and user environment data can be written into these data structure objects, as can spatially mapped page context data. Based on data identification information in metadata, index links between data layers can be established. For example, a correlation field can be added to the data structure object to record the correspondence between mouse behavior data and page context data.
[0081] The server can verify the integrity of the integrated multi-dimensional data set based on the tamper-proof signature in the metadata. If the verification passes, the multi-dimensional data set can be stored in memory or a database; if the verification fails, the multi-dimensional data set can be marked as abnormal.
[0082] The structured objects of multi-dimensional data collections can be in JavaScript object representation or binary format, the index links can be in pointer references or foreign key associations, and the integrity verification can be in hash value comparison or signature verification algorithm.
[0083] As can be seen from the above embodiments, this application establishes a temporal correlation by aligning mouse behavior data and user environment data based on a reference timestamp, thereby eliminating differences in data time bases and ensuring consistency between behavioral events and environmental states in the time dimension. It also establishes a spatial correlation by mapping mouse behavior data to coordinate space based on the visible area size and interface element coordinates, thus eliminating the impact of device resolution differences and assigning specific page semantics to trajectory points. Furthermore, by integrating multiple types of data based on temporal and spatial correlations, a multi-dimensional data set is obtained, solving the problem of data isolation and constructing a data foundation with complete spatiotemporal context. On this basis, subsequent feature extraction can be performed based on the correlated multi-dimensional data, rather than relying on isolated data points, thereby significantly improving the accuracy and reliability of user interaction behavior risk identification.
[0084] Based on any embodiment of the method in this application, multidimensional behavioral features are extracted from a multidimensional dataset. These multidimensional behavioral features characterize the target user's behavioral features in different dimensions during webpage interaction, including: Step S5410: Obtain the trajectory point sequence from the mouse behavior data of the multi-dimensional dataset. The trajectory point sequence includes the trajectory point coordinates and trajectory point event time.
[0085] The server obtains a sequence of trajectory points from mouse behavior data in a multi-dimensional dataset. The sequence of trajectory points includes the coordinates of the trajectory points and the event times of the trajectory points. The acquisition operation includes accessing the mouse behavior data, reading the stored path data, and parsing out the coordinates of the trajectory points and the event times of the trajectory points.
[0086] The server can read trajectory point sequences or fine-grained trajectory fields for functional areas. The field data format is an array or list, where each array element represents a trajectory point. The server can iterate through each element in the array, parsing the x-coordinate, y-coordinate, and timestamp values contained in each element. The parsed coordinate and timestamp values are stored in a trajectory point sequence object, and the sequence is sorted according to the timestamp values to ensure chronological order. If the data contains event type identifiers, the server can extract them and associate them with the corresponding trajectory points.
[0087] Step S5420: Based on the coordinates of the trajectory points and the event time of the trajectory points, obtain the kinematic features, statistical features, time series features and geometric features respectively.
[0088] Based on the trajectory point coordinates and trajectory point event times, the server acquires kinematic features, statistical features, time series features, and geometric features, respectively. The server calculates motion state indicators using trajectory point coordinates and trajectory point event times, calculates statistical distribution indicators using trajectory point distribution, calculates frequency domain indicators using time series changes, and calculates geometric morphology indicators using the spatial relationship between trajectory point coordinates and interface element coordinates.
[0089] Kinematic features characterize the physical motion of mouse movement. Statistical features characterize the spatial or temporal distribution of trajectory points. Time series features characterize the dynamic pattern of behavior changing over time. Geometric features characterize the spatial geometric relationship between the trajectory shape and the target element.
[0090] Step S5430: Obtain composite features based on the interaction relationships between kinematic features, statistical features, time series features, and geometric features.
[0091] The server extracts composite features based on the interactions between kinematic, statistical, time-series, and geometric features. These interactions refer to the correlations, ratios, or combinational logic between feature values of different dimensions. The extraction process involves performing mathematical operations or vector concatenation on the four basic feature classes to generate new feature vectors representing multi-dimensional coupling patterns.
[0092] Composite features represent higher-order behavioral patterns that cannot be captured by single-dimensional features. Single features may be affected by noise or lack discriminative power. By exploring the interaction relationships between features, we can capture the deep logic behind behavior. For example, the coupling between speed and trajectory curvature may reflect operational proficiency.
[0093] In one embodiment, the server can read instantaneous velocity and acceleration values from kinematic features, spatial distribution entropy and directional autocorrelation values from statistical features, power spectral density and fractal dimension values from time series features, trajectory curvature index and Fitts' law fit value from geometric features, and user environment data and page context data from the multi-dimensional dataset as auxiliary inputs. The server can calculate the ternary mutual information between instantaneous velocity values, trajectory curvature values, and time series data to obtain three-dimensional coupling features of velocity, curvature, and time, capturing the essential differences between humans and machines in velocity, curvature, and time series relationships; it can use a long short-term memory network model to predict the next position, calculate the entropy value of the prediction error distribution, obtain a trajectory predictability index, and distinguish between the low predictability of human behavior and the high determinism or abnormal randomness of machine behavior; it can calculate the Mahalanobis distance of the four-dimensional feature vectors (kinematic, statistical, time series, and geometric) to obtain a multi-dimensional consistency index, detecting the consistency within each dimension of features and identifying contradictory behaviors between dimensions; and it can combine the interval time in the page context data and the browser identifier and gaze density in the user environment data to construct a cognitive load proxy index. This system can reflect the differences between human cognitive processing and machine simulation; it can compare the device type declared by the browser in user environment data with the smoothness of trajectory features, calculate the device behavior matching degree, and detect user agent forgery or abnormal behavior patterns; it can calculate the cosine similarity of feature vectors of multiple operations by the same user to obtain a trajectory fingerprint stability index, and use stable human biological behavioral characteristics to distinguish the inconsistent behavior generated independently by the machine each time; it can use the isolated forest algorithm to detect outliers, analyze the spatial distribution and clustering of outliers, and identify the differences between randomly distributed human outliers and machine-clustered outliers at specific stages; it can integrate a weighted comprehensive index of multi-scale entropy, spectral entropy, and spatial entropy to obtain a comprehensive index of information theory complexity, providing a single-dimensional overall complexity assessment.
[0094] It should be noted that the calculation method for interaction relationships is not limited to mutual information, Mahalanobis distance, or cosine similarity. Correlation analysis, principal component analysis, or neural network encoding can also be used, as long as they can represent the dependencies between features.
[0095] Step S5440: Integrate kinematic features, statistical features, time series features, geometric features, and composite features to obtain multidimensional behavioral features.
[0096] The server integrates kinematic features, statistical features, time series features, geometric features, and composite features to obtain multidimensional behavioral features, which represent the comprehensive characteristics of the target user's interactive behavior.
[0097] The fusion operation can standardize and combine the five types of features to generate feature vectors of a unified dimension.
[0098] The server can read kinematic feature vectors, statistical feature vectors, time-series feature vectors, geometric feature vectors, and composite feature vectors. It then standardizes each feature vector using Z-score standardization or Min-Max normalization to map the feature values to a uniform range. The server can arrange the standardized feature vectors in a preset order, such as kinematic features, statistical features, time-series features, geometric features, and composite features. The server concatenates these vectors to generate an initial multidimensional behavioral feature vector. Missing values are handled using fill-in methods such as mean imputation, zero imputation, or interpolation. Finally, the server stores the initial multidimensional behavioral feature vector as a multidimensional behavioral feature. For example, the feature vector dimension could be 127, encompassing all generated feature items, such as a 32-dimensional kinematic feature vector, a 28-dimensional statistical feature vector, a 35-dimensional time-series feature vector, a 24-dimensional geometric feature vector, and an 8-dimensional composite feature vector.
[0099] It should be noted that the fusion method is not limited to vector concatenation. Weighted summation, attention mechanism fusion, or principal component analysis dimensionality reduction fusion can also be used, as long as the feature information can be preserved.
[0100] As can be seen from the above embodiments, this application achieves the effect of ensuring that feature calculation is based on spatiotemporal correlation data by obtaining trajectory point sequences from mouse behavior data of multi-dimensional datasets; by obtaining kinematic features, statistical features, time series features, and geometric features based on trajectory point coordinates and trajectory point event times, it achieves the effect of introducing page context semantics and enhancing the ability of features to represent specific interactive intentions; by obtaining composite features based on the interaction relationship between kinematic features, statistical features, time series features, and geometric features, it achieves the effect of capturing the essential differences in human-computer behavior in high-order coupling modes and improving feature discrimination; by fusing kinematic features, statistical features, time series features, geometric features, and composite features to obtain multi-dimensional behavioral features, it achieves the effect of constructing a unified feature vector that comprehensively represents user behavior characteristics; on this basis, risk identification can be based on multi-dimensional and deep-level behavioral information, thereby further significantly improving the accuracy and reliability of user interaction behavior risk identification.
[0101] Based on any embodiment of the method in this application, kinematic features, statistical features, time series features, and geometric features are obtained based on the coordinates of the trajectory points and the event time of the trajectory points, including: Step S5510: Calculate instantaneous velocity, acceleration, and jerk based on the coordinate difference and time difference between adjacent trajectory points to obtain kinematic characteristics.
[0102] The server calculates instantaneous velocity, acceleration, and jerk based on the coordinate and time differences between adjacent trajectory points, thus obtaining kinematic characteristics. The calculation process includes traversing the sequence of trajectory points, extracting the coordinate and time values of adjacent trajectory points, performing difference operations to obtain a velocity sequence, performing difference operations again on the velocity sequence to obtain an acceleration sequence, and performing difference operations on the acceleration sequence to obtain a jerk sequence.
[0103] Kinematic characteristics characterize the physical motion of the mouse pointer in page space. Instantaneous velocity reflects the rate of change of displacement per unit time. Acceleration reflects the rate of change of velocity. Jerk reflects the rate of change of acceleration and is used to quantify the smoothness of the trajectory. Human hand movements follow the principle of minimum jerk, resulting in a smooth trajectory and continuous changes in jerk. Automated scripts often ignore higher-order derivative characteristics, exhibiting constant velocity or abrupt acceleration. By calculating third-order kinematic indices, the essential differences between human biomechanical characteristics and machine-simulated behavior can be captured.
[0104] The server can iterate through the sequence of trajectory points, select the i-th trajectory point and the (i+1)-th trajectory point as adjacent trajectory point pairs, calculate the Euclidean distance between adjacent trajectory points to obtain the displacement value, and calculate the time difference between adjacent trajectory points to obtain the time interval value. The server can then divide the displacement value by the time interval value to obtain the instantaneous velocity value. The server can calculate statistics for the velocity sequence, including the mean. Standard deviation Maximum value Minimum value , median skewness kurtosis and coefficient of variation The server can calculate velocity entropy. The formula is as follows: in, For the interval of velocity distribution, This represents the probability that the velocity value falls within the specified interval.
[0105] The server can calculate tangential acceleration, and since human acceleration changes according to the minimum jerk constraint, it can also calculate normal acceleration. ,in Let be the radius of curvature of the trajectory at point i. Humans naturally slow down when turning. The server can calculate statistics for the acceleration sequence, including the mean, standard deviation, and heavy-tailed distribution characteristics.
[0106] The server can calculate accelerometer. ,in This refers to the acceleration value. Humans minimize jerk, while machines generate discontinuous jerk. The server can calculate the smoothness index. Human values are smaller and more stable, while machine values are abnormal. The server can calculate the zero-crossing rate of accelerometers, which is the ratio of the number of accelerometer sign changes to the total number of points. Humans make frequent fine adjustments, while machines make monotonous adjustments.
[0107] The server can calculate coupling characteristics and velocity-curvature coupling. ,in Let the covariance be the velocity versus the curvature. For the speed standard deviation, This represents the standard deviation of curvature. Humans exhibit a strong negative correlation (i.e., deceleration during turns), while machines show a weak correlation. Servers can calculate the time correlation of speed, i.e., the autocorrelation function of a speed sequence. Human speed patterns have short-term memory, while machines either lack it or exhibit abnormalities. Servers can also calculate the rate of change of direction. ,in Let be the direction angle of motion at point i. Humans turn smoothly, while machines may turn abruptly.
[0108] The server can store the calculated instantaneous velocity, acceleration, jerk, statistics, entropy, smoothness index, zero crossing rate, and coupling features as kinematic features.
[0109] Step S5520: Calculate the spatial distribution entropy, time interval entropy, directional autocorrelation, and fluctuation characteristics of the trajectory points to obtain statistical characteristics.
[0110] The server calculates the spatial distribution entropy, time interval entropy, directional autocorrelation, and fluctuation characteristics of the trajectory points to obtain statistical features.
[0111] To calculate the spatial distribution entropy, the server can obtain the width and height of the visible area from the page context data, divide the visible area into m spatial grids, and calculate the probability that the coordinates fall into the j-th spatial grid. ,in , The number of trajectory points falling into the j-th grid. Calculate the spatial distribution entropy value for the total number of trajectory points. Humans have a complex, multi-peaked distribution with high entropy, while machines have a simpler distribution with lower entropy. Servers can also calculate the spatial clustering of trajectory points as a supplementary indicator of spatial distribution, using Ripley's K function or the DBSCAN clustering algorithm to calculate the clustering coefficient of the point set. Humans exhibit clear target clustering, while machines may be scattered or overly uniform. Servers can also calculate the bounding box fill rate. ,in, This is the actual trajectory length. and Given the width and height of the minimum bounding box, the ratio is higher for human trajectories that are more circuitous, and lower for machines that tend to follow the shortest path.
[0112] For calculating the time interval entropy value, the server can calculate the time interval between adjacent trajectory points. probability distribution of statistical time intervals Calculate the entropy value of the time interval. Human time intervals are irregular, resulting in higher entropy values, while machine sampling is uniform, leading to lower entropy values. The server can also detect pause behavior as a supplementary indicator of time distribution, identifying continuous time periods with speeds below a preset threshold, and statistically analyzing the number of pauses, total duration, average duration, and maximum duration. Humans have thought-action cycles, while machines execute continuously. The server can also calculate the complexity of operation rhythms, using the Lempel-Ziv complexity algorithm to measure the compressibility of time series. Human sequences have high complexity, while machine sequences have strong compressibility.
[0113] For calculating directional autocorrelation, the server can calculate the angle between adjacent line segments. ,in, Let be the motion direction angle of the i-th segment. The server calculates statistics (mean, variance, etc.) of the angle distribution. Human angles are diverse, while machine angles follow a pattern. The server can also calculate directional persistence as a supplementary indicator of directional autocorrelation, and calculate the Hurst exponent H of the direction sequence. For humans, H ≈ 0.5 (random walk), while for machines, H → 1 (deterministic) or H → 0 (completely random). The server can also calculate the autocorrelation coefficient of the direction sequence. ,in, The value represents the mean of the direction angles. Human movement directions have short-term memory and a high autocorrelation coefficient.
[0114] For calculating wave characteristics, the server can calculate microtremor detection indicators, perform bandpass filtering on the trajectory coordinate sequence (8-12Hz), and calculate the energy of the filtered signal. Humans exhibit physiological tremors with specific energy values, which machines either lack or simulate unnaturally. The server can also calculate noise spectrum characteristics, perform spectral analysis on the coordinate residuals (after detrending), and calculate the logarithmic power spectral density. With logarithmic frequency The slope of the linear regression is obtained. value, For frequency, humans ,machine or The server can also calculate spectral entropy, based on the power spectral density distribution, with human noise being 1 / Pink noise has a specific entropy value, while machine noise is either white noise or Brownian noise, which have different entropy values.
[0115] The server can store the calculated spatial distribution entropy, time interval entropy, directional autocorrelation, and fluctuation characteristics as statistical features.
[0116] Step S5530: Perform frequency domain transformation and fractal dimension calculation on the instantaneous velocity, and calculate the time domain structure characteristics to obtain the time series characteristics.
[0117] The server performs frequency domain transformation and fractal dimension calculation on instantaneous velocities, and then calculates temporal structural features to obtain time series characteristics. The frequency domain transformation operation is used to analyze the frequency distribution characteristics of velocity fluctuations. The fractal dimension calculation operation is used to quantify the self-similarity and space-filling ability of the velocity sequence. The temporal structural feature calculation operation is used to evaluate the consistency and recursive structure characteristics of trajectory segments. The time series characteristics characterize the dynamic patterns of behavior evolution over time and its long-range correlations.
[0118] Time series characteristics characterize the dynamic patterns of behavior as it evolves over time. Mouse trajectories are treated as multidimensional time series, and time series analysis methods are used to capture these dynamic patterns. Human movement is controlled by the biological nervous system, and speed fluctuations exhibit specific frequency domain distribution characteristics and fractal geometric features.
[0119] For time series features with frequency domain transformation, the server can perform detrending processing on the instantaneous velocity sequence v. Where N is the total number of trajectory points, the velocity sequence is transformed in the frequency domain using the Fast Fourier Transform algorithm, and the power spectral density is calculated. The system calculates the dominant frequency components, detects the peak frequency of the power spectral density and its energy proportion. Humans have characteristic control frequencies (typically concentrated in the 1-3 Hz neural control bandwidth), while machines lack these frequencies or deviate from them. The server can also calculate the high-frequency energy ratio. Humans use moderate high-frequency energy (for hand tremors), while machines use excessively high (noise) or excessively low (for smoothing); the server can also calculate the spectral centroid. The centroid of the human spectrum is stable, while that of the machine is abnormally offset.
[0120] For time series features calculated using fractal dimension, the server can use a side length of... The grid covers the phase space of the velocity sequence, and the number of grids containing the sequence points is counted. Calculate the box count dimension Human beings Satisfy 1.2 < <1.5, while the machine's ≈1 (linear) or ≈2 (random). The server can also calculate the information dimension. ,in, For scale Information entropy: Human information dimension reflects the fractal characteristics of probability distribution, while machine information dimension is abnormal; the server can also calculate the correlation dimension, and calculate the correlation integral based on the Grassberger-Procaccia algorithm. Human correlation dimension captures the correlation structure of trajectory, while machine correlation dimension is abnormal.
[0121] For temporal structural features, the server can divide the trajectory into k segments and calculate the cosine similarity matrix of the feature vectors between segments. Human segments exhibit self-similarity, while machine segments show large differences or excessive uniformity. The server can perform variational mode decomposition, decomposing the trajectory into K intrinsic mode functions (IMFs) and calculating the energy ratio of each mode. Human mode energy distribution is complex, while machine modes concentrate on low frequencies. The server can also perform recursive graph analysis to construct the recursive matrix of the trajectory. ,in, and This represents the state vector after reconstructing the phase space. The distance between vectors For step functions, calculate recursion rate, determinism, and entropy. Humans have rich recursive structures, while machines have simple or random ones.
[0122] The server can also use complexity analysis and complexity features as additional time series features.
[0123] For complexity features, the server can calculate approximate entropy. Construct an m-dimensional vector sequence and calculate the distance between vectors if it is less than the tolerance. Calculate the probability. ,in The server calculates the logarithmic average of the matching probabilities of m-dimensional vectors. Human trajectories are highly complex, while machine trajectories are less complex. The server can also calculate sample entropy SampEn, using an improved approximate entropy algorithm that excludes self-matching terms, is insensitive to data length, and is more robust to human noise. The server can also calculate fuzzy entropy FuzzyEn, using a fuzzy membership function to replace a hard threshold for judging vector similarity, which better handles trajectory noise. The server can also calculate multi-scale entropy, coarsely processing instantaneous velocity sequences and calculating sample entropy at multiple time scales. Human cross-scale entropy is stable, while machine entropy is highly scale-dependent.
[0124] The server can integrate and store the calculated frequency domain features, fractal features, complexity features, and time domain structure features into time series features.
[0125] Step S5540: Calculate the trajectory curvature index, target approach behavior index, gaze simulation index, and trajectory morphology index to obtain geometric features.
[0126] The server calculates trajectory curvature indices, target proximity behavior indices, gaze simulation indices, and trajectory morphology indices to obtain geometric features. The calculation operations include calculating curvature-related indices based on trajectory coordinate derivatives, calculating target proximity behavior indices based on the relationship between task difficulty and motion time, simulating gaze points based on dwell behavior and calculating gaze simulation indices, and calculating trajectory morphology indices based on spatial distribution. These geometric features characterize the spatial geometric properties and visual-motor coordination patterns of the target user's interactive behavior.
[0127] Geometric features characterize the spatial geometric properties and visual-motor coordination patterns of target user interaction behavior. The human visual-motor system follows hand-eye coordination and Fitts's Law, and the trajectory geometry reflects the allocation of visual attention. Geometric features include four independent dimensions: curvature index reflecting the degree of trajectory curvature, target proximity behavior index reflecting directional movement patterns, gaze simulation index reflecting visual attention allocation, and trajectory morphology index reflecting spatial filling characteristics.
[0128] Regarding the geometric characteristics of the trajectory curvature index, the server can smooth the coordinate sequence and calculate the first derivative. and and second derivative and Calculate curvature Human curvature changes continuously, while machine curvature is either piecewise constant or abrupt; the server can also calculate the curvature integral. ,in, As the differential of the arc length between adjacent trajectory points, humans have a larger total turning amount, while machines tend to follow the path with the minimum curvature; the server can calculate the coefficient of curvature variation. ,in, For the standard deviation of curvature, The curvature is the mean value. Human curvature fluctuates greatly, while machine curvature is monotonous. The server can also perform curvature spectrum analysis, perform Fourier transform on the curvature sequence, and analyze the frequency components. Human curvature has specific spectral characteristics.
[0129] For the geometric characteristics of the target proximity behavior indicators, the server can perform a linear regression fitting model. ,in, The distance of the operation. Calculate the coefficient of determination for the target width. human users slope ,intercept machine behavior ,or ,or The residuals exhibit a systematic pattern; the server can also calculate the velocity decay rate during the final approach phase. ,in, Speed required to enter the target area At the moment of clicking, humans decelerate significantly, while machines may enter at a constant speed or stop abruptly; the server can also calculate the efficiency of the approach path. ,in, The Euclidean distance from the starting point to the target is [value], and the actual trajectory length is [value]. Human efficiency is moderate (0.8-0.95), while machine efficiency is either too high or too low. Servers can classify approach strategies based on curvature distance relationships, categorizing them into direct, roundabout, probing, and corrective approaches. Human strategies are diverse, while machine strategies are singular.
[0130] Regarding the geometric characteristics of gaze simulation metrics, the server can count the number of virtual gaze points and identify those that meet the speed requirements. And the segment lasting longer than 100ms. Typically set to 5 pixels per second, this allows humans to focus on key locations while the machine moves at a constant speed; the server can also calculate the spatial distribution entropy of the focus points. ,in, The probability of the gaze point falling into the j-th spatial grid is given. Human gaze points are concentrated in the information area, while machine gaze points are abnormally distributed. The server can also calculate the gaze action interval and record the time interval between the end of the last gaze and the click action. Humans have a typical reaction time (200-400ms), while machine intervals are abnormal.
[0131] Based on the geometric characteristics of the trajectory morphology indicators, the server can calculate the trajectory turning radius. ,in, , The trajectory is the average of the coordinates of the points on the trajectory. Human trajectories have a clear target, and the turning radius is related to the target position. The server can also calculate the ratio of the convex hull area of the trajectory. ,in, The trajectory convex hull area represents the area of the trajectory. Human trajectories fill the convex hull relatively fully, while machine trajectories exhibit abnormally circuitous or overly direct movements. The server can also calculate the fractal dimension of the trajectory using box counting. Fractal dimension is calculated based on spatial coordinate paths, unlike calculations based on velocity sequences. Humans... The machine is close to 1 or 2.
[0132] As can be seen from the above embodiments, this application obtains kinematic features by calculating instantaneous velocity, acceleration, and jerk, achieving the effect of capturing human biomechanical characteristics and minimum jerk constraints, and distinguishing machine interpolation trajectories; it obtains statistical features by calculating spatial distribution entropy, time interval entropy, directional autocorrelation, and fluctuation characteristics, achieving the effect of quantifying the inherent variability of behavior and noise spectrum characteristics, and distinguishing human randomness from machine determinism; it obtains time series features by performing frequency domain transformation, fractal dimension calculation, and time domain structure feature calculation, achieving the effect of capturing dynamic patterns and multi-scale complexity, and distinguishing neural control bandwidth from machine patterns; it obtains geometric features by calculating trajectory curvature indicators, target approach behavior indicators, gaze simulation indicators, and trajectory morphology indicators, achieving the effect of reflecting visual-motor coordination laws and attention allocation, and distinguishing human hand-eye coordination from machine path planning; on this basis, risk identification is based on multi-dimensional and deep-level behavioral features, thereby further improving the accuracy of user interaction behavior risk identification and resistance to simulation by automated scripts.
[0133] Based on any embodiment of the method in this application, before identifying the risk of a target user's interaction behavior according to multidimensional behavioral characteristics to determine the corresponding risk type of the target user, the method includes: Step S5610: Perform a distribution stability test on the multidimensional behavioral features and remove features whose feature distribution offset index exceeds a preset threshold.
[0134] The server performs a distribution stability test on multidimensional behavioral features, removing features whose distribution offset exceeds a preset threshold. The test includes calculating the distribution differences of the multidimensional behavioral features across different data subsets. The removal operation aims to retain highly stable features and eliminate the influence of environmental noise.
[0135] Distribution stability refers to the consistency of feature values across different sample groups or time segments. Feature distribution shift quantifies the degree of distribution variation and is used to assess the sensitivity of features to environmental factors. Multidimensional behavioral features may experience distribution drift due to differences in client environments, device types, or time variations. Unstable feature distributions can cause risk identification models to fail in specific environments. Stability tests are used to screen out robust multidimensional behavioral features.
[0136] The server can acquire a set of multi-dimensional behavioral features and divide the dataset into a baseline group and a test group. The baseline group might be historical normal user data, while the test group might be current user data to be evaluated or data from different time windows. The server can perform binning on each feature, dividing continuous values into several intervals and calculating the distribution difference of each feature between the baseline and test groups.
[0137] Distribution stability tests can be performed using the population stability index PSI, or using KL divergence, JS divergence, Kolmogorov-Smirnov test, or chi-square test, as long as the distribution differences can be quantified.
[0138] Step S5620: Calculate the correlation coefficient between the remaining multidimensional behavioral features and the risk label, and retain the features whose correlation coefficient exceeds the first preset threshold.
[0139] The server calculates the correlation coefficients between the remaining multidimensional behavioral features and the risk labels, retaining features whose correlation coefficients exceed a first preset threshold. The correlation coefficient characterizes the strength of the association between the feature value and the risk label. The retention operation aims to filter out features with significant predictive power for risk assessment.
[0140] The correlation coefficient is a statistical indicator that measures the strength of a linear or non-linear relationship between two variables. A risk label refers to a confirmed risk status identifier in historical data; for example, 0 represents a normal user, and 1 represents a risky user. Even after stability testing, multidimensional behavioral features may still contain redundant features unrelated to the risk status. Retaining irrelevant features introduces noise. The correlation coefficient is calculated to evaluate the explanatory power of each feature for the risk label. A first preset threshold is used to define the minimum standard for feature validity. Retaining highly correlated features reduces data dimensionality and improves the convergence speed and classification accuracy of the risk identification model.
[0141] The correlation coefficient can be calculated using Pearson correlation coefficient or Spearman rank correlation coefficient, or Kendall rank correlation coefficient, information value (IV), mutual information or chi-square test value, as long as the correlation strength between the quantifiable feature and the label can be used.
[0142] The first preset threshold can be adjusted according to the risk tolerance of the specific business scenario. For example, a higher threshold can be used in high-risk scenarios to filter stronger features.
[0143] Risk labels can be binary labels, multi-category labels, or continuous risk scores, as long as they can represent the degree of risk.
[0144] Step S5630: Calculate the multicollinearity among the retained multidimensional behavioral features, and remove features whose multicollinearity exceeds the second preset threshold.
[0145] The server calculates the multicollinearity among the retained multidimensional behavioral features and removes features whose multicollinearity exceeds a second preset threshold. The multicollinearity calculation operation includes evaluating the degree of linear correlation between feature variables. The removal operation aims to eliminate redundant features and reduce the coupling between feature dimensions.
[0146] Multicollinearity refers to the phenomenon where two or more feature variables in a multidimensional behavioral feature set exhibit a high degree of linear correlation. Severe multicollinearity among features can lead to unstable parameter estimation, variance inflation, and increased risk of overfitting in risk identification models. A second preset threshold is used to define the acceptable upper limit of redundancy among features.
[0147] Multicollinearity can be calculated using the variance inflation factor (VIF), the condition number (ConditionNumber), the correlation coefficient matrix between features, or the tolerance factor (Tolerance), as long as it can quantify the degree of linear dependence between features.
[0148] The second preset threshold can be adjusted according to the requirements of the specific processing flow. For example, a stricter threshold can be used for scenarios with high independence requirements.
[0149] As can be seen from the above embodiments, this application achieves the effect of eliminating the influence of environmental noise and ensuring the cross-domain stability of features by performing distribution stability tests on multidimensional behavioral features and removing features whose distribution offset index exceeds a preset threshold; by calculating the correlation coefficient between the remaining features and the risk label and retaining features that exceed a first preset threshold, it achieves the effect of screening features with high discriminative power and enhancing risk prediction capabilities; by calculating the multicollinearity among the retained features and removing features that exceed a second preset threshold, it achieves the effect of eliminating feature redundancy and improving the independence and stability of data processing; on this basis, risk identification is based on a high-quality, low-redundancy feature subset, thereby significantly improving the accuracy, generalization ability and computational efficiency of user interaction behavior risk identification.
[0150] Based on any embodiment of the method in this application, risk identification is performed on the interactive behavior of the target user according to multi-dimensional behavioral characteristics to determine the corresponding risk type of the target user, including: Step S5710: Based on the values of each feature in the multidimensional behavioral features, calculate the matching degree between the target user's interactive behavior and the preset risk behavior pattern.
[0151] The server calculates the matching degree between the target user's interactive behavior and a preset risk behavior pattern based on the numerical values of each feature in the multidimensional behavioral features. The matching degree calculation includes constructing a feature vector from the multidimensional behavioral features, constructing a reference vector or distribution model from the preset risk behavior pattern, and measuring the similarity or distance between the feature vector and the reference vector or distribution model. The matching measure quantifies the degree to which the target user's behavior deviates from the normal pattern or approaches the risk pattern.
[0152] Matching degree refers to the numerical value of the degree of similarity between the target user's behavioral characteristics and known risk characteristic patterns. Preset risk behavior patterns refer to the feature distribution patterns, cluster centers, or rule sets extracted based on historical risk sample data, representing typical risk behavior characteristics.
[0153] Matching degree can be calculated using distance metrics or weighted scoring, or it can use similarity cosine value, kernel function mapping distance, or rule matching score, as long as the degree of conformity between the behavior and the pattern can be quantified.
[0154] Preset risk behavior patterns can be achieved using feature center vectors, risk sample distribution boundaries, cluster centers, or expert rule bases, as long as they can represent risk characteristics.
[0155] Step S5720: Determine the risk probability value of the interaction behavior based on the matching degree. The risk probability value represents the possibility that the interaction behavior belongs to a risky behavior.
[0156] The server determines the risk probability value of the interaction behavior based on the matching degree. The risk probability value represents the likelihood that the interaction behavior is a risky behavior. The operation of determining the risk probability value includes mapping the matching degree value to a probability range and generating a standardized risk quantification index.
[0157] The risk probability value refers to the likelihood that a target user's interaction behavior is judged as a risky behavior, typically ranging from 0 to 1. Matching degree is a similarity measure in the feature space, which may be affected by its dimensions or distribution and is not directly equivalent to probability. Through mapping operations, the matching degree is converted into a probabilistic value, facilitating a unified risk decision-making standard. The higher the risk probability value, the greater the likelihood that the interaction behavior is a risky behavior.
[0158] Step S5730: Determine the risk type corresponding to the target user based on the probability interval to which the risk probability value belongs, where different probability intervals correspond to different risk types.
[0159] The server determines the risk type corresponding to the target user based on the probability interval to which the risk probability value belongs. Different probability intervals correspond to different risk types. The risk type determination operation includes comparing the risk probability value with the preset interval boundaries and matching the corresponding risk category identifier.
[0160] Risk probability values are continuous numerical values, representing the likelihood that an interaction is a risky behavior. Risk types are discrete categories, representing the specific degree or nature of the risk. Probability intervals refer to dividing the range of risk probability values into several continuous sub-ranges. Through interval mapping, continuous probability quantification results are transformed into specific business decision categories, facilitating the subsequent implementation of differentiated risk control measures. Different probability intervals correspond to different risk types; for example, low-risk intervals correspond to normal types, medium-risk intervals to suspicious types, and high-risk intervals to high-risk types. This mapping relationship establishes a logical connection between numerical scoring and business classification, ensuring the standardization and executability of risk assessment.
[0161] The server can be configured with thresholds. and ,in . For example, set it to 0.3. For example, set it to 0.7. The probability interval is divided into the first interval. Second interval and the third interval The server can determine the range in which the risk probability value falls. If it falls into the first range, the risk type is determined to be normal; if it falls into the second range, the risk type is determined to be suspicious; and if it falls into the third range, the risk type is determined to be high-risk.
[0162] The server can perform subsequent operations based on the risk type. For example, if the risk type is normal, the server allows the interaction to continue; if the risk type is suspicious, the server triggers secondary verification; if the risk type is high-risk, the server blocks the interaction.
[0163] As can be seen from the above embodiments, this application achieves the effect of quantifying the degree of conformity between behavioral features and risk patterns and transforming abstract features into measurable indicators by calculating the matching degree between the target user's interactive behavior and the preset risk behavior pattern based on the numerical calculation of each feature in the multidimensional behavioral features; by determining the risk probability value of the interactive behavior based on the matching degree, the matching result is standardized into a unified probability indicator and the difference in units is eliminated; by determining the risk type corresponding to the target user based on the probability interval to which the risk probability value belongs, the continuous probability is transformed into a discrete risk category to support differentiated risk control; on this basis, risk judgment is based on the quantified matching degree and the standardized probability value, thereby improving the interpretability of user interactive behavior risk identification, the degree of decision standardization, and the execution efficiency of risk control strategies.
[0164] like Figure 3 As shown, a user behavior risk identification device according to one aspect of this application includes: a trajectory acquisition module 5100, a parsing and association module 5200, a feature extraction module 5300, and a risk identification module 5400. The trajectory acquisition module 5100 acquires mouse trajectory data generated by a target user during webpage interaction. The parsing and association module 5200 performs structured parsing of the mouse trajectory data to obtain a multi-dimensional data set with a standardized structure, representing the association between mouse behavior data, page context data, and user environment data. The feature extraction module 5300 extracts multi-dimensional behavioral features from the multi-dimensional data set, representing the target user's behavioral characteristics in different dimensions during webpage interaction. The risk identification module 5400 identifies risks in the target user's interaction behavior based on the multi-dimensional behavioral features to determine the corresponding risk type for the target user.
[0165] Based on any embodiment of the device in this application, the parsing and association module 5200 includes: a standardized structure parsing and hierarchy determination module, used to identify the standardized structural framework of mouse trajectory data and determine the nesting hierarchy of mouse trajectory data; a nested data extraction module, used to extract mouse behavior data, page context data, and user environment data from mouse trajectory data according to the nesting hierarchy; and an association binding and integration module, used to associate and bind the mouse behavior data, page context data, and user environment data and integrate them to obtain a multi-dimensional data set.
[0166] Based on any embodiment of the device in this application, the association binding and integration module includes: a time alignment and association module, used to perform time alignment of mouse behavior data and user environment data based on a preset reference timestamp to establish a time association relationship between mouse behavior data and user environment data; a spatial mapping and association module, used to extract the visible area size and interface element coordinates from the page context data, and map the mouse behavior data to the coordinate space defined by the visible area size to establish a spatial association relationship between mouse behavior data and interface element coordinates; and a spatiotemporal integration module, used to integrate mouse behavior data, page context data, and user environment data based on the time association relationship and the spatial association relationship to obtain a multi-dimensional data set.
[0167] Based on any embodiment of the device in this application, the feature extraction module 5300 includes: a trajectory point sequence extraction module, used to obtain a trajectory point sequence from mouse behavior data in a multi-dimensional dataset, the trajectory point sequence including trajectory point coordinates and trajectory point event time; a basic behavior feature extraction module, used to obtain kinematic features, statistical features, time series features, and geometric features based on trajectory point coordinates and trajectory point event time; a composite feature generation module, used to obtain composite features based on the interaction relationship between kinematic features, statistical features, time series features, and geometric features; and a multi-dimensional behavior feature fusion module, used to fuse kinematic features, statistical features, time series features, geometric features, and composite features to obtain multi-dimensional behavior features.
[0168] Based on any embodiment of the device in this application, the basic behavioral feature extraction module includes: a kinematic feature calculation module, used to calculate instantaneous velocity, acceleration, and jerk based on the coordinate difference and time difference between adjacent trajectory points to obtain kinematic features; a statistical feature calculation module, used to calculate the spatial distribution entropy value, time interval entropy value, directional autocorrelation, and fluctuation characteristics of trajectory points to obtain statistical features; a time series feature extraction module, used to perform frequency domain transformation and fractal dimension calculation on instantaneous velocity, and calculate time domain structural features to obtain time series features; and a geometric feature calculation module, used to calculate trajectory curvature index, target approach behavior index, gaze simulation index, and trajectory morphology index to obtain geometric features.
[0169] Based on any embodiment of the device in this application, before the risk identification module 5400, the device includes: a distribution stability test and feature removal module, used to perform a distribution stability test on the multidimensional behavioral features and remove features whose feature distribution offset index exceeds a preset threshold; a correlation screening module, used to calculate the correlation coefficient between the remaining multidimensional behavioral features and the risk label, and retain features whose correlation coefficient exceeds a first preset threshold; and a multicollinearity removal module, used to calculate the multicollinearity among the retained multidimensional behavioral features and remove features whose multicollinearity exceeds a second preset threshold.
[0170] Based on any embodiment of the device in this application, the risk identification module 5400 includes: a matching degree calculation module, used to calculate the matching degree between the interactive behavior of the target user and the preset risk behavior pattern based on the values of each feature in the multi-dimensional behavioral features; a risk probability assessment module, used to determine the risk probability value of the interactive behavior according to the matching degree, wherein the risk probability value represents the possibility that the interactive behavior belongs to a risky behavior; and a risk type determination module, used to determine the risk type corresponding to the target user according to the probability interval to which the risk probability value belongs, wherein different probability intervals correspond to different risk types.
[0171] like Figure 4As shown in the diagram, another embodiment of this application also provides an electronic device, wherein the internal structure of the electronic device is illustrated. The electronic device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement the user behavior risk identification method described in this application. The processor of the electronic device provides computing and control capabilities to support the operation of the entire electronic device. The memory of the electronic device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the user behavior risk identification method described in this application. The network interface of the electronic device is used for communication with a terminal. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0172] In this execution method, the processor is used for execution. Figure 3 The system contains the specific functions of each module and its sub-modules. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this execution method, the memory stores the program code and data required to execute all modules / sub-modules in the user behavior risk identification device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0173] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the method described in any embodiment of this application.
[0174] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM).
[0175] The above description is only a partial implementation method of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for identifying user behavior risks, characterized in that, include: Acquire mouse trajectory data generated by the target user during webpage interaction; The mouse trajectory data is subjected to structured parsing to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the relationship between mouse behavior data, page context data, and user environment data. Multidimensional behavioral features are extracted from the multidimensional dataset, and the multidimensional behavioral features characterize the target user's behavioral characteristics in different dimensions during web page interaction; Based on the multidimensional behavioral characteristics, risk identification is performed on the interaction behavior of the target user to determine the corresponding risk type of the target user.
2. The user behavior risk identification method according to claim 1, characterized in that, The step involves performing structured parsing on the mouse trajectory data to obtain a multi-dimensional data set with a standardized structure. This multi-dimensional data set represents the relationships between mouse behavior data, page context data, and user environment data, including: Identify the standardized structural framework of the mouse trajectory data and determine the nesting level of the mouse trajectory data; Based on the nesting hierarchy, mouse behavior data, page context data, and user environment data are extracted from the mouse trajectory data; The mouse behavior data, the page context data, and the user environment data are associated, bound, and integrated to obtain the multi-dimensional data set.
3. The user behavior risk identification method according to claim 2, characterized in that, The process of associating, binding, and integrating the mouse behavior data, the page context data, and the user environment data to obtain the multi-dimensional data set includes: The mouse behavior data and the user environment data are time-aligned based on a preset reference timestamp to establish a temporal correlation between the mouse behavior data and the user environment data. The visible area size and interface element coordinates are extracted from the page context data, and the mouse behavior data is mapped to the coordinate space defined by the visible area size to establish a spatial association between the mouse behavior data and the interface element coordinates. Based on temporal and spatial relationships, the mouse behavior data, page context data, and user environment data are integrated to obtain a multi-dimensional data set.
4. The user behavior risk identification method according to claim 3, characterized in that, The step of extracting multidimensional behavioral features from the multidimensional data set, wherein the multidimensional behavioral features characterize the target user's behavioral characteristics in different dimensions during webpage interaction, includes: A trajectory point sequence is obtained from the mouse behavior data of the multi-dimensional dataset, the trajectory point sequence including trajectory point coordinates and trajectory point event time; Based on the coordinates of the trajectory points and the event times of the trajectory points, kinematic features, statistical features, time series features, and geometric features are obtained respectively; Based on the interaction relationships among the kinematic features, statistical features, time series features, and geometric features, composite features are obtained; The kinematic features, statistical features, time series features, geometric features, and composite features are fused to obtain the multidimensional behavioral features.
5. The user behavior risk identification method according to claim 4, characterized in that, The process of acquiring kinematic features, statistical features, time series features, and geometric features based on the trajectory point coordinates and trajectory point event time includes: Based on the coordinate difference and time difference between adjacent trajectory points, the instantaneous velocity, acceleration, and jerk are calculated to obtain the kinematic characteristics; The spatial distribution entropy, time interval entropy, directional autocorrelation, and fluctuation characteristics of the trajectory points are calculated to obtain the statistical characteristics. The instantaneous velocity is subjected to frequency domain transformation and fractal dimension calculation, and the time domain structure characteristics are calculated to obtain the time series characteristics; The geometric features are obtained by calculating the trajectory curvature index, target approach behavior index, gaze simulation index, and trajectory morphology index.
6. The user behavior risk identification method according to claim 1, characterized in that, Before determining the corresponding risk type of the target user by identifying the target user's interaction behavior based on the multidimensional behavioral characteristics, the process includes: The distribution stability of the multidimensional behavioral features is tested, and features whose distribution offset index exceeds a preset threshold are removed. Calculate the correlation coefficient between the remaining multidimensional behavioral features and the risk label, and retain features whose correlation coefficient exceeds the first preset threshold; Calculate the multicollinearity among the retained multidimensional behavioral features, and remove features whose multicollinearity exceeds a second preset threshold.
7. The user behavior risk identification method according to claim 1, characterized in that, The step of identifying risks in the target user's interaction behavior based on the multidimensional behavioral characteristics to determine the corresponding risk type for the target user includes: Based on the values of each feature in the multidimensional behavioral features, the matching degree between the target user's interactive behavior and the preset risk behavior pattern is calculated; The risk probability value of the interaction behavior is determined based on the matching degree, and the risk probability value represents the likelihood that the interaction behavior is a risky behavior. Based on the probability range to which the risk probability value belongs, the risk type corresponding to the target user is determined, wherein different probability ranges correspond to different risk types.
8. A user behavior risk identification device, characterized in that, include: The trajectory acquisition module is used to acquire mouse trajectory data generated by the target user during web page interaction; The parsing and association module is used to perform structured parsing on the mouse trajectory data to obtain a multi-dimensional data set with a standardized structure. The multi-dimensional data set represents the association between mouse behavior data, page context data, and user environment data. The feature extraction module is used to extract multi-dimensional behavioral features from the multi-dimensional dataset, wherein the multi-dimensional behavioral features characterize the target user's behavioral features in different dimensions during webpage interaction; The risk identification module is used to identify risks in the interaction behavior of the target user based on the multidimensional behavioral characteristics, so as to determine the risk type corresponding to the target user.
9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.
10. A non-volatile storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.