Intention recognition method, device, equipment, storage medium and program product

By extracting features and recognizing intents from sequences of intelligent agent interaction events, the problem of traditional human-machine recognition mechanisms struggling to identify intelligent agent intents is solved, achieving accurate recognition of intelligent agent intents and ensuring the security and reliability of online services.

CN122366449APending Publication Date: 2026-07-10CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIONPAY
Filing Date
2026-04-08
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional human-machine recognition mechanisms struggle to effectively identify the behavioral intentions of intelligent agents based on large models, threatening the security and reliability of online services.

Method used

By acquiring the event types and attributes in the interaction event sequence, feature extraction is performed to identify the agent's sub-intent sequence and determine its target intent. Accurate identification is then achieved using a pre-trained agent recognition model and intent database.

Benefits of technology

It enables accurate identification of agent intent, provides reliable evidence to distinguish between benevolent and malicious agents, and ensures the security and reliability of online services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122366449A_ABST
    Figure CN122366449A_ABST
Patent Text Reader

Abstract

The application discloses an intention recognition method, device, equipment, storage medium and program product. The method comprises the following steps: obtaining interaction data generated in an interaction process between an operation subject and an operated object, the interaction data comprising an interaction event sequence, the interaction event sequence comprising a plurality of interaction events arranged in time sequence, the interaction event comprising an event type and an event attribute; in the case that the operation subject is an intelligent agent, performing first feature extraction on the interaction events in the interaction event sequence to obtain a first feature sequence; identifying a sub-intention sequence of the intelligent agent based on the first feature sequence; and determining a target intention of the intelligent agent based on the sub-intention sequence. According to the embodiment of the application, the intention of the intelligent agent can be accurately recognized, which provides a reliable basis for subsequent differentiation between a benign intelligent agent and a malicious intelligent agent, thereby finally guaranteeing the safety and reliability of online services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to an intent recognition method, apparatus, device, storage medium and program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, various intelligent agents are gradually becoming important users of online services. These intelligent agents can simulate human operations and autonomously complete complex tasks such as booking flights and comparing prices, greatly improving service efficiency and user experience.

[0003] However, the widespread application of intelligent agents has also brought new security challenges. Malicious intelligent agents may impersonate ordinary users to carry out risky behaviors such as data scraping, fraud attacks, brute-force attacks, and service abuse. Traditional human-machine recognition mechanisms (such as the Turing test and behavioral verification) are mainly designed to distinguish between humans and simple scripts, and are difficult to effectively identify the behavioral intentions of intelligent agents based on large models.

[0004] Therefore, there is an urgent need for an intent recognition method to accurately identify the target intent of an intelligent agent during task execution, providing a reliable basis for distinguishing between benevolent and malicious intelligent agents, thereby ultimately ensuring the security and reliability of online services. Summary of the Invention

[0005] This application provides an intent recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can accurately identify the intent of an intelligent agent, providing a reliable basis for subsequently distinguishing between benevolent and malicious intelligent agents, thereby ultimately ensuring the security and reliability of online services.

[0006] In a first aspect, embodiments of this application provide an intent recognition method, the method comprising: During the interaction between the operating subject and the operated object, the interaction data generated by the interaction process is acquired. The interaction data includes a sequence of interaction events, which includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. When the operating subject is an intelligent agent, the first feature is extracted from the interaction events in the interaction event sequence to obtain the first feature sequence; Based on the first feature sequence, identify the agent's sub-intention sequence; Based on the sub-intent sequence, the target intent of the agent is determined.

[0007] Secondly, embodiments of this application provide an intent recognition device, the device comprising: The acquisition module is used to acquire interaction data generated during the interaction between the operating subject and the operated object. The interaction data includes a sequence of interaction events, which includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. The extraction module is used to extract first features from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, so as to obtain a first feature sequence; The identification module is used to identify the sub-intent sequence of the agent based on the first feature sequence; A determination module is used to determine the target intent of the agent based on the sub-intent sequence.

[0008] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements any of the possible implementations of the first aspect described above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method in any of the possible implementations of the first aspect described above.

[0010] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0011] In this embodiment, since the interaction event sequence includes multiple interaction events arranged in chronological order, and each interaction event includes an event type and an event attribute, by extracting a first feature from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, a first feature sequence is obtained. This first feature sequence preserves the event type information and event attribute information of the intelligent agent during the interaction process, as well as the temporal relationship between events. Based on this, by identifying the sub-intent sequence of the intelligent agent based on the first feature sequence, continuous interaction events can be converted into a sub-intent sequence with specific semantics. By determining the target intent of the intelligent agent based on this sub-intent sequence with specific semantics, the target intent of the intelligent agent during task execution can be accurately identified. Thus, through this embodiment, the intent of the intelligent agent can be accurately identified, providing a reliable basis for subsequently distinguishing between benevolent and malicious intelligent agents, thereby ultimately ensuring the security and reliability of online services. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart illustrating an intent recognition method provided in one embodiment of this application; Figure 2 This is a flowchart illustrating an intent recognition method provided in another embodiment of this application; Figure 3 This is a schematic diagram of the structure of an intent recognition device provided in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0014] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0015] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0016] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0017] Furthermore, the acquisition, storage, use, and processing of data in this application's technical solution all comply with relevant national laws and regulations.

[0018] To address the related technical problems, embodiments of this application provide an intent recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product. The intent recognition method can be applied to scenarios involving the recognition of the behavioral intent of an intelligent agent. The intelligent agent can be, for example, a graphical user interface intelligent agent (GUI agent) or an artificial intelligence (AI) assistant on a smart terminal.

[0019] The intent recognition method provided in the embodiments of this application is described below.

[0020] Figure 1 A flowchart illustrating an embodiment of the intent recognition method provided in this application is shown. This intent recognition method can be executed by an intent recognition system. Figure 1 As shown, the intent recognition method provided in this application includes the following steps: S110. During the interaction between the operating subject and the operated object, acquire the interaction data generated during the interaction process. The interaction data includes a sequence of interaction events. The sequence of interaction events includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. S120. When the operating subject is an intelligent agent, the first feature is extracted from the interaction events in the interaction event sequence to obtain the first feature sequence. S130. Based on the first feature sequence, identify the agent's sub-intention sequence; S140. Determine the target intent of the agent based on the sub-intent sequence.

[0021] In this embodiment, since the interaction event sequence includes multiple interaction events arranged in chronological order, and each interaction event includes an event type and an event attribute, by extracting a first feature from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, a first feature sequence is obtained. This first feature sequence preserves the event type information and event attribute information of the intelligent agent during the interaction process, as well as the temporal relationship between events. Based on this, by identifying the sub-intent sequence of the intelligent agent based on the first feature sequence, continuous interaction events can be converted into a sub-intent sequence with specific semantics. By determining the target intent of the intelligent agent based on this sub-intent sequence with specific semantics, the target intent of the intelligent agent during task execution can be accurately identified. Thus, through this embodiment, the intent of the intelligent agent can be accurately identified, providing a reliable basis for subsequently distinguishing between benevolent and malicious intelligent agents, thereby ultimately ensuring the security and reliability of online services.

[0022] The specific implementation methods for each of the above steps are described below.

[0023] In some embodiments, in S110, the operating entity is an entity that initiates and executes a series of operations during interaction with the electronic device. The operating entity can include intelligent agents and humans. Additionally, the electronic device includes the object being operated on. Specifically, the object being operated on can be a front-end interface element or a back-end service interface of the electronic device operated by the operating entity during the interaction. The front-end interface elements can include interactive elements (such as buttons, input boxes, checkboxes, etc.), navigation elements (such as menu items, links, tabs, etc.), form elements (such as text input fields, file upload controls, etc.), content areas (such as list items, cards, etc.), and container elements (such as pop-ups, drawers, modals, etc.). The back-end service interface can include data resources (such as database tables, caches, file storage, etc.), functional modules (such as authentication modules, payment modules, etc.), and service interfaces (such as microservice interfaces, third-party tool interfaces, native interfaces, etc.).

[0024] Furthermore, upon detecting any interactive behavior on the electronic device (such as front-end access behavior), the system can begin acquiring interaction data and continuously acquire the interaction data generated during the interaction between the operator and the operated object. This interaction data can include a sequence of interaction events. This sequence of interaction events completely depicts the operator's behavioral trajectory from initiating the interaction to the current moment, providing a real-time and complete data foundation for subsequent human-computer recognition and intent recognition.

[0025] As an example, interactive behaviors can be captured in real time using front-end probes (such as JavaScript) or back-end interceptors, recording each operation as an interactive event and caching these events. In front-end scenarios, interactive events include, but are not limited to, click events, mouse movement events, keyboard events, page scrolling events, wait events, and form events. In back-end scenarios, interactive events include, but are not limited to, Application Programming Interface (API) calls and service requests. Interactive events can include event types and event attributes. Event types can include cursor movement, cursor clicks (such as single or double clicks), input, scrolling, form filling, dragging, etc. Event attributes can include at least one of static and dynamic attributes. Static attributes can include coordinate position, timestamp, time interval, whether it is an interactive element, whether it is a visible element, etc. If a static attribute indicates that the manipulated object is an interactive element, the static attribute can also include the element semantic information of the interactive element, which can include the functional semantics of the interactive element and the hierarchical path of the interactive element in the interactive interface. Additionally, dynamic attributes can include instantaneous speed, acceleration, jerk, trajectory curvature, scroll length, etc. By sorting the collected interaction events according to their chronological order, an interaction event sequence can be obtained. Every preset period (e.g., per second), the intent recognition system can retrieve the interaction event sequence from the cache and determine the interaction data based on the sequence. Additionally, for scenarios with high cursor movement frequency, sampling can be configured based on time or distance.

[0026] After acquiring the interaction data, the system can first perform human-machine recognition based on the data. If the operator is identified as an intelligent agent, the system can then proceed with intent recognition. If the operator is identified as a human, the tracking and recognition of the interaction can be stopped.

[0027] In some embodiments, the interaction data may or may not include interaction protocol data. This is not a limitation. Interaction protocol data refers to data carried during the interaction process that reflects characteristics at the interaction protocol level. Examples of interaction protocol data include request headers, authentication fields, and call parameters in API calls. The interaction protocol data may or may not include agent identifiers; this is not a limitation.

[0028] As an example, upon receiving a user instruction, an intelligent agent can plan and reason about actions based on the response, first generating an action chain and then executing it. This action chain may include user interface operations, native API calls, and tool calls. In API authorization mode, the intelligent agent can carry a clear agent identifier in its interaction protocol data. Benign agents are more likely to carry an agent identifier during interactions, while malicious agents, aiming to mimic human actions to avoid detection, typically do not carry an agent identifier.

[0029] Based on this, in order to quickly identify the operating entity at the interaction entry point and avoid redundant human-machine identification processes for accesses that have been clearly identified as intelligent agents, in some embodiments, before S120, the method may further include: If the interaction protocol data includes an agent identifier, the operator is identified as an agent.

[0030] Here, after acquiring the interaction data, we can first parse the interaction protocol data to determine whether it includes an agent identifier. If the interaction protocol data includes an agent identifier, the operator can be directly identified as an agent, thus quickly identifying the operator at the interaction entry point. This avoids redundant human-machine recognition processes for accesses that are already clearly identified as agents, improving the efficiency of subsequent intent recognition.

[0031] Furthermore, in cases where the interaction data does not include interaction protocol data, or the interaction protocol data does not include an agent identifier, in order to accurately identify the operating subject, in some embodiments, the method may further include the following before S120: The second feature sequence is obtained by extracting the second feature from the interaction events in the interaction event sequence. The operator is identified based on the second feature sequence.

[0032] Here, the second feature extraction can be the conversion of interaction events into second feature vectors. The second feature sequence can include multiple second feature vectors obtained from the interaction events, arranged in chronological order.

[0033] In the process of human-machine recognition, considering that the key difference between human and intelligent agent operations lies in some dynamic operational features, in order to improve the accuracy of human-machine recognition, the event attributes in this embodiment can include static attributes and dynamic attributes. Static attributes can be basic static attributes. Examples of basic static attributes include coordinate position, timestamp, time interval, whether it is an interactive element, and whether it is a visible element. Dynamic attributes can include instantaneous velocity, acceleration, jerk, trajectory curvature, and scroll length.

[0034] Based on this, in some embodiments, the above-mentioned extraction of second features from the interaction events in the interaction event sequence to obtain a second feature sequence may specifically include: Encode the event type of the interactive event to obtain the event type code value; Encode the basic static properties of interactive events to obtain the basic static property values; Based on the coordinates and timestamps of interactive events, as well as the coordinates and timestamps of adjacent interactive events, dynamic attribute values ​​are calculated. The dynamic attribute values ​​include at least one of instantaneous velocity, acceleration, jerk, trajectory curvature, and scroll length. The event type encoding value, basic static attribute value, and dynamic attribute value are concatenated to obtain the second feature vector corresponding to the interactive event; The second feature sequence is obtained by sorting the second feature vectors corresponding to the multiple interaction events according to their arrangement in the interaction event sequence.

[0035] Here, event types can include cursor movement, cursor clicks (such as single or double clicks), input, scrolling, form filling, dragging, etc. The encoding of event types can be one-hot encoding. This encoding method avoids the spurious ordering problem introduced by directly assigning numerical values ​​to different event types, thus enabling the model to correctly understand the equal relationships between event types.

[0036] Furthermore, the basic static attributes can be either continuous or discrete variables. If the basic static attributes are continuous variables, normalization can be used to map them to the [0,1] interval to obtain their values. If the basic static attributes are discrete variables, Boolean encoding or one-hot encoding can be used to encode them to obtain their values.

[0037] Furthermore, for any interaction event in the sequence of interaction events, based on the coordinate position and timestamp in the basic static attributes of that interaction event, as well as the coordinate positions and timestamps of adjacent interaction events, the dynamic attribute values ​​corresponding to multiple dynamic attributes can be calculated. The specific method for calculating dynamic attribute values ​​based on coordinate position and timestamp can be found in existing methods and will not be elaborated upon here.

[0038] By performing the above encoding operation on each interactive event in the interactive event sequence, the interactive event can be converted into a second feature vector. By sorting the second feature vectors corresponding to the multiple interactive events according to their order in the interactive event sequence, the second feature sequence can be obtained.

[0039] If the event type encoding value is denoted as The basic static property value is denoted as Record the dynamic attribute value as Then the second feature vector can be The second characteristic sequence can be denoted as

[0040] This application's embodiments effectively capture the operational characteristics of the operating subject by utilizing features such as instantaneous velocity, acceleration, jerk, trajectory curvature, and roll length in dynamic attributes. By encoding event types and basic static attributes and concatenating them with dynamic attribute values ​​to obtain a second feature vector, and then arranging the second feature vector in chronological order to form a second feature sequence, the temporal and feature information of the operating subject's behavioral trajectory is fully preserved. This provides an accurate data foundation for human-machine recognition, thereby improving the accuracy of human-machine recognition.

[0041] After converting the event interaction sequence into a second feature sequence, human-machine recognition can be performed based on the second feature sequence to identify the operating entity.

[0042] Therefore, in order to improve the accuracy and reliability of human-machine recognition, in some embodiments, the above-mentioned identification of the operating subject based on the second feature sequence may specifically include: The subject recognition model identifies the operator based on the second feature sequence and outputs a confidence score that the operator is an agent. The confidence score ranges from 0 to 1, and the higher the score, the greater the probability that the operator is an agent. The subject recognition model is trained based on the agent feature sequence and its corresponding agent label, and the human feature sequence and its corresponding human label. If the confidence score is greater than the first threshold, the operator is determined to be an intelligent agent; If the confidence score is less than the second threshold, the operator is determined to be human. If the confidence score is between the first threshold and the second threshold, return to the process of interaction between the operator and the operated object, and obtain the interaction data generated during the interaction process until the confidence score is greater than the first threshold or less than the second threshold.

[0043] Here, the second feature sequence is a multidimensional time series. The subject recognition model can be a pre-trained time series classification model used for human-machine recognition. This subject recognition model can be used to classify multidimensional time series (such as the second feature sequence).

[0044] As an example, this subject recognition model can be trained based on agent feature sequences and their corresponding agent labels, and human feature sequences and their corresponding human labels. The agent feature sequences and human feature sequences can be pre-collected. Specifically, for each pre-set interaction goal (such as navigating and locating a specific position on a webpage, filling out and editing a form, etc.), agent feature sequences and human feature sequences can be collected separately.

[0045] For human feature sequences, interaction event sequences generated by real human users interacting with the manipulated object are collected by deploying event tracking points (such as JavaScript) on the front end. These sequences are then processed to form human feature sequences. This collection process occurs silently in the background and does not affect the user's normal experience.

[0046] For the agent feature sequence, natural language instructions covering various interaction goals are given to the agent, and the sequence of interaction events generated during the agent's task execution is collected in the cloud. Then, the interaction event sequence is processed for a second feature to form the agent feature sequence.

[0047] To enrich the diversity of training samples, data augmentation techniques can be used to process agent feature sequences and / or human feature sequences. These data augmentation techniques include, but are not limited to, adding noise, modifying timing sequences, and simulating network jitter, among one or more, to generate more variant sample data.

[0048] Additionally, the agent label can be, for example, 1, and the human label can be, for example, 0.

[0049] After training the subject recognition model, the second feature sequence is input into the subject recognition model, which then classifies the second feature sequence to obtain a confidence score. A higher confidence score indicates a greater probability that the subject is an intelligent agent.

[0050] Specifically, if the confidence score is greater than the first threshold, the operator can be identified as an intelligent agent, and the process proceeds to the intent recognition stage. If the confidence score is less than the second threshold, the operator can be identified as a human, and the tracking and recognition of the interaction behavior ceases. If the confidence score is between the first and second thresholds, a definitive judgment cannot be made, and the interaction event sequence can continue to be retrieved from the cache, and human-machine recognition can continue until the current operator is determined to be a real human or an intelligent agent. The first and second thresholds can be pre-determined based on preset evaluation metrics. Specifically, the first and second thresholds can be formulated based on the actual business scenario and the Receiver Operating Characteristic (ROC) curve and Precision-Recall (PR) curve of the dataset. The first threshold is set to a high confidence score. Only when the model output score exceeds this threshold is the operator identified as an intelligent agent. This ensures that while pursuing high accuracy, the false positive rate is kept at a low level, avoiding misjudging real human users. The second threshold is set to a low confidence score. Only when the model output score is below this threshold is the operator identified as a human. This ensures that while pursuing high recall, intelligent agents are identified as much as possible, reducing the risk of false negatives.

[0051] This application embodiment uses a subject recognition model for human-machine recognition and sets a first threshold and a second threshold. It can ensure that high-confidence samples are directly judged, while continuously collecting behavioral data and iteratively recognizing samples with insufficient confidence until a clear judgment can be made. This can maximize the recognition of intelligent agents without harming real users, and significantly improve the accuracy and reliability of human-machine recognition.

[0052] In addition, to efficiently and accurately identify the identity of the operating entity, in some embodiments, the entity recognition model can be, for example, an improved Inception-time neural network model. The overall structure of the entity recognition model includes, in sequence: a bottleneck layer, multiple feature extraction layers of different lengths, a feature fusion layer, a temporal attention layer, and a fully connected layer. The temporal attention layer may include a feedforward neural network layer and a weighted pooling layer.

[0053] Based on this, the above-mentioned subject recognition model identifies the operator based on the second feature sequence and outputs a confidence score that the operator is an agent. Specifically, this can include: The second feature sequence is reduced in dimension by using the bottleneck layer to obtain the reduced feature sequence. By using multiple feature extraction layers of different lengths, depthwise separable convolutions are performed on the dimensionality-reduced feature sequences to extract temporal features at different scales, resulting in multiple first feature maps. Through the feature fusion layer, multiple first feature maps are fused and transformed to obtain a second feature map. The second feature map includes T time steps, each time step corresponding to a fused feature vector, where T is a positive integer. The attention weights are calculated for the feature vectors at each time step through the feedforward neural network layer of the attention layer, resulting in T attention weights. The third feature map is obtained by weighted summation of the feature vectors at T time steps through the weighted pooling layer of the attention layer, based on T attention weights. The third feature map is processed by a linear transformation and activation function through a fully connected layer, and the output is a confidence score of the agent.

[0054] Here, after inputting the second feature sequence into the subject recognition model, the model first performs dimensionality reduction on the second feature sequence using a bottleneck layer. Multiple convolutional filters of length 1 are used to reduce the dimensionality of the input sequence, resulting in a dimensionality-reduced feature sequence that reduces computational complexity while preserving key information. Subsequently, multiple feature extraction layers of different lengths are used to extract multi-scale temporal features from the dimensionality-reduced feature sequence using depthwise separable convolutions. For example, convolutional kernels of lengths 40, 20, and 10 are used to independently extract behavioral features at different time scales, resulting in multiple first feature maps. The feature fusion layer uses 1×1 convolutional kernels to perform channel fusion and transformation on the multiple first feature maps, effectively combining the features extracted at different scales to obtain a second feature map containing T time steps. The attention layer calculates the attention weight for each time step using a feedforward neural network, and then uses a weighted pooling layer to sum the feature vectors of the T time steps, allowing the model to focus on key operational moments in the behavioral trajectory (i.e., the sequence of interaction events). Finally, a fully connected layer is used for linear transformation and activation function processing to output a confidence score for the agent.

[0055] Based on the above structure, this application embodiment reduces computational overhead through a bottleneck layer, captures both short-term instantaneous features and long-term behavioral patterns simultaneously through multi-scale convolutional kernels, reduces the number of parameters while maintaining feature extraction capabilities through depthwise separable convolution, and automatically focuses on operations at key moments through a time attention mechanism, thereby achieving efficient and accurate identification of the identity of the operating subject.

[0056] In some embodiments, in S120, upon identifying the operator as an intelligent agent, the second stage of intent recognition and tracking is automatically triggered. This stage continuously analyzes the agent's subsequent behavior to identify its intent. Specifically, firstly, by performing first feature extraction on the interaction events in the interaction event sequence, the interaction events can be converted into a first feature vector, and the interaction event sequence can be converted into a first feature sequence, providing a data foundation for recognizing the agent's intent. Unlike the aforementioned second feature extraction for human-machine recognition, the first feature extraction focuses on extracting semantic features related to intent recognition. Therefore, event attributes can include extended static attributes without including dynamic attributes, thus reducing data processing complexity while ensuring the accuracy of intent recognition.

[0057] Based on this, in order to reduce data processing complexity while ensuring the accuracy of intent recognition, in some embodiments, the first feature extraction of the interaction events in the interaction event sequence to obtain the first feature sequence may specifically include: Encode the event type of the interactive event to obtain the event type code value; Encode the extended static properties of the interactive event to obtain the extended static property values; The event type encoding value and the extended static attribute value are concatenated to obtain the first feature vector corresponding to the interactive event; The first feature sequence is obtained by sorting the first feature vectors corresponding to the multiple interaction events according to their arrangement order in the interaction event sequence.

[0058] Here, event types can include cursor movement, cursor clicks (such as single or double clicks), input, scrolling, form filling, dragging, etc. The encoding of event types can be one-hot encoding. This encoding method avoids the spurious ordering problem introduced by directly assigning numerical values ​​to different event types, thus enabling the model to correctly understand the equal relationships between event types.

[0059] Extended static properties can include basic static properties, as well as element semantic information of interactive elements corresponding to interactive events. Element semantic information refers to the functional meaning of the interactive element at the business level and its position in the interface structure. In a browser environment, element semantic information can be obtained through the Document Object Model (DOM) tree. Specifically, element semantic information can include the functional semantics of the interactive element and its hierarchical path in the interactive interface. Functional semantics can be the specific function implemented by the interactive element in the business scenario, such as a "login button" for submitting a login request, a "shopping cart icon" for viewing selected products, and a "search input box" for entering search keywords. The hierarchical path can represent the position of the interactive element in the interactive interface structure, such as its depth in the DOM tree, parent-child relationship, and whether it belongs to a form or container. The hierarchical path reflects the contextual relationship of elements in the interface layout; for example, whether a button is located within a "login form" or a "navigation bar" is of significant reference value for understanding the user's intent.

[0060] Based on this, the above-mentioned encoding of extended static attributes of interactive events yields extended static attribute values, which may specifically include: Encode the basic static properties to obtain their values; Encode the semantic information of the element to obtain the semantic encoded value of the element; By concatenating the basic static attribute value and the element semantic encoding value, the extended static attribute value is obtained.

[0061] Here, the basic static attributes can be either continuous or discrete variables. If the basic static attributes are continuous variables, normalization can be used to map them to the [0,1] interval to obtain their values. If the basic static attributes are discrete variables, Boolean encoding or one-hot encoding can be used to encode them to obtain their values.

[0062] Furthermore, the purpose of encoding element semantic information is to transform the aforementioned unstructured raw DOM information into a numerical representation that the model can understand. In some embodiments, the encoding process specifically includes: processing the functional description text of the element using a pre-trained lightweight natural language processing model, converting it into a semantic vector; simultaneously, performing positional encoding on the depth information of the DOM tree to preserve the hierarchical relationship of elements in the interface structure. The element semantic encoding value is obtained by fusing the functional semantic vector and the positional encoding. In this way, the model can not only identify what type of element the operator is acting on, but also understand the specific meaning of the element in the business scenario and its contextual position in the interface, thereby providing rich semantic support for subsequent intent recognition.

[0063] Based on this, extended static attribute values ​​can be obtained by concatenating basic static attribute values ​​and element semantic encoding values. By performing the above encoding operation on each interactive event in the interactive event sequence, the interactive event can be converted into a first feature vector. By sorting the first feature vectors corresponding to multiple interactive events according to their order in the interactive event sequence, the first feature sequence can be obtained.

[0064] If the event type encoding value is denoted as The basic static property value is denoted as The semantic encoding value of the element is denoted as Then the first feature vector can be The first characteristic sequence can be denoted as

[0065] This application embodiment encodes the semantic information of elements, transforming the functional meaning and hierarchical path of interactive elements into numerical representations that the model can understand. This enables the model to understand the specific meaning of the elements operated by the subject in the business scenario and their interface context, thereby providing rich semantic support for intelligent agent intent recognition. This allows subsequent intent recognition to focus on business semantic information related to the operation target, without including dynamic attributes, thus reducing data processing complexity while ensuring the accuracy of intent recognition.

[0066] In some embodiments, in S130, before performing intent recognition, an intent library corresponding to the agent's behavior can be established based on the functions and risk events of the electronic device (including the front-end interface and the back-end system). This intent library can include high-level intents and sub-intents, typical behaviors and risk levels. This intent library is applicable not only to front-end graphical user interface interaction scenarios but also to back-end service interface call scenarios. High-level intents refer to the overall behavioral purpose manifested by the agent when completing a certain business objective. High-level intents can include major categories such as reconnaissance and information collection, account and identity operations, transaction and fund operations, and marketing activities. Each major category can contain several specific high-level intents and corresponding typical malicious behaviors, as shown in Table 1 below: Table 1

[0067] In addition, sub-intents refer to the sub-goals or sub-stages required to complete a higher-level intent; they are more granular action types. Each sub-intent can correspond to one or more interaction events. Sub-intents can include broad categories such as navigation and positioning, information retrieval and understanding, and form filling and editing, with each category containing several specific sub-intents, as shown in Table 2 below: Table 2

[0068] Each high-level intent can include multiple sub-intents arranged in sequence; that is, each high-level intent can correspond to a sequence of sub-intents. Taking the high-level intent "Account Registration" as an example, its corresponding sub-intent sequence could be: {Enter registration portal, fill in form, select, submit registration, verify}.

[0069] Therefore, to accurately identify the agent's high-level intent, we can first identify the sub-intent sequence, and then identify the high-level intent based on the sub-intent sequence to obtain the agent's target intent. Since each sub-intent can correspond to one or more interaction events, after converting the interaction event sequence into a first feature sequence, we can identify the agent's sub-intent sequence based on the first feature sequence. The first feature sequence can include multiple feature sequence segments, and each feature sequence segment can include at least one first feature vector.

[0070] Based on this, in order to accurately locate and identify the implicit sub-intent sequences in the operation behavior and provide structured input for subsequent high-level intent recognition, in some embodiments, the above-mentioned S130 may specifically include: Obtain a library of preset sub-intents, which includes multiple preset sub-intents. Each preset sub-intent corresponds to a preset feature sequence segment. The preset feature sequence segment includes at least one preset feature vector. The preset feature vector is determined based on the event type of the interaction event and the extended static attributes. If a feature sequence segment matches a preset feature sequence segment, the preset sub-intention corresponding to the preset feature sequence segment will be identified as the sub-intention corresponding to the feature sequence segment. The identified sub-intents are arranged according to the order of the feature sequence segments in the first feature sequence to obtain the sub-intent sequence.

[0071] Here, the preset sub-intent library can be the intent library corresponding to the specific sub-intent in the aforementioned intent library. The preset sub-intent can be the specific sub-intent as shown in Table 2. In addition, the preset feature sequence segment lengths corresponding to different preset sub-intents may be different. For example, the "click" sub-intent may correspond to only one click event, while the "fill in form" sub-intent may correspond to multiple input events.

[0072] As an example, after obtaining the first feature sequence, the feature sequence segment can be matched with a preset feature sequence segment using a sliding window approach. Specifically, starting from the initial position i=1, the preset window length L increases from 1 to a preset maximum length threshold L. maxFor each L, extract a feature sequence segment of length L starting from position i in the first feature sequence, and match it with all preset feature sequence segments of length L. If a match is successful, the feature sequence segment is identified as the corresponding sub-intention, the starting position i is updated to i+L, L=1 is reset, and matching continues from the new starting position; if L reaches L... max If no match is found, abandon the current starting position, move the starting position i one position to the right (i.e., i = i + 1), reset L = 1, and continue trying to match. Repeat the above process until the starting position i exceeds the length of the first feature sequence.

[0073] Let's take the first feature sequence as [1,2,3,4,5] (each number represents a first feature vector) as an example. Assume L max =3. Starting from i=1, L=1, match [1]. If the match fails, L=2, match [1,2]. If the match still fails, L=3, match [1,2,3]. If the match succeeds, it is identified as the corresponding sub-intention, i is updated to 4, and L=1 is reset. Starting from i=4, L=1, match [4]. If the match succeeds, it is identified as the corresponding sub-intention, i is updated to 5, and L=1 is reset. Starting from i=5, L=1, match [5]. If the match succeeds, it is identified; if the match fails and L=1 has reached L... max (Because the sequence is not long enough to reach L) max If the actual length is used as the reference, then i is updated to 6, which exceeds the sequence length, and the matching ends.

[0074] If the match [1,2,3] fails, and [2,3,4] is a preset feature sequence segment of a certain sub-intention, then when i=1 and L=3 fails to match, i is updated to 2, L=1 is reset, and starting from position 2, L=1 is used to try matching [2], L=2 to match [2,3], and L=3 to match [2,3,4] in sequence. At this time, [2,3,4] can be successfully matched, thus identifying the sub-intention. Through the above mechanism, it is ensured that no matter where the sub-intention sequence starts or what its length is, it can be accurately identified without missing any possible matching cases.

[0075] Furthermore, when matching a feature sequence segment with a preset feature sequence segment, various methods can be used to calculate the similarity or distance between them to determine whether a match is successful. These methods include, but are not limited to, calculating Euclidean distance, calculating cosine similarity, dynamic time warping algorithms, deep learning matching, and rule-based direct matching (such as whether the feature sequence segment matches at least one feature vector in the preset feature sequence segment). In practical applications, one or more methods can be selected and combined based on the length of the preset feature sequence segment, the dimension of the feature vectors, and the business scenario requirements, and an appropriate similarity threshold can be set to determine whether the feature sequence segment matches the preset feature sequence segment successfully.

[0076] Thus, through the embodiments of this application, it is possible to accurately locate and identify the sub-intents hidden in the operation behavior, and arrange them into a sub-intent sequence according to the order in which the behavior occurs, providing structured input for subsequent high-level intent recognition.

[0077] In some embodiments, in S140, after recognizing the sub-intention sequence, the higher-level intention can be recognized based on the sub-intention sequence to obtain the target intention of the agent. The sub-intention sequence may include one or more sub-intentions.

[0078] In some embodiments, the above-mentioned S140 may specifically include: Obtain the preset intent library, which includes multiple preset intents and their corresponding standard sub-intent sequences; Based on the sequence matching algorithm, the matching degree between the sub-intent sequence and each standard sub-intent sequence is determined; When the matching degree characterization sub-intent sequence matches the standard sub-intent sequence, the preset intent corresponding to the standard sub-intent sequence is determined as the target intent.

[0079] Here, each standard sub-intent sequence may include one or more sub-intents. The lengths of different standard sub-intent sequences may vary. Since the sub-intent sequences obtained in real-time recognition are dynamically growing, and the lengths of different standard sub-intent sequences may vary, a combination of sliding window matching and delayed acknowledgment can be used for high-level intent recognition to balance real-time performance and accuracy.

[0080] Specifically, during real-time sub-intent recognition, a cache of currently recognized sub-intent sequences can be maintained. Whenever a new sub-intent is recognized, it is appended to the end of the cache, resulting in an updated complete sub-intent sequence. For each standard sub-intent sequence, a sequence matching algorithm is used to calculate the matching degree between the current complete sub-intent sequence and that standard sub-intent sequence. The sequence matching algorithm can include at least one of the following: calculating Euclidean distance, calculating cosine similarity, dynamic time warping algorithm, deep learning matching, and rule-based direct matching (e.g., whether the sub-intent sequence matches at least one sub-intent in the standard sub-intent sequence). Since the same sub-intent sequence may match multiple standard sub-intent sequences simultaneously, a multi-candidate parallel matching mechanism is used to ensure correct matches are not missed. The top K standard sub-intent sequences with the highest matching degrees are selected as the current candidate intents, and their respective matching degrees are recorded, where K is a positive integer.

[0081] To avoid mismatches due to incomplete factor intent sequences, a delayed confirmation mechanism is introduced: when the matching degree of a candidate intent reaches a preset initial matching threshold, the high-level intent is not immediately confirmed, but rather marked as an observation state. During the observation period, subsequent sub-intents continue to be identified. Each newly identified sub-intent is appended to the end of the cache, the complete sub-intent sequence is updated, and the matching degree between the candidate intent and the updated complete sub-intent sequence is recalculated, dynamically updating the matching degree. When the matching degree of the candidate intent reaches a preset matching threshold (e.g., a matching degree exceeding 0.8, or a complete match of the entire standard sequence), the candidate intent is officially confirmed as the target intent. If, during the observation period, the matching degree of the candidate intent continues to decrease or deviates significantly, the candidate intent is removed from the candidate list and no longer tracked. If all candidate intents are removed, subsequent sub-intents continue to be identified, forming new candidate intents.

[0082] By employing the above methods, misjudgments caused by incomplete factor intent sequences are avoided, and subsequent handling strategies can be triggered in real time after confirmation, thus achieving a balance between accuracy, real-time performance, and robustness in high-level intent recognition.

[0083] Furthermore, in order to ensure system security while minimizing interference with normal users and achieving a balance between security and user experience, in some embodiments, after determining the preset intent corresponding to the standard sub-intent sequence as the target intent, the method may further include: When the target intent is risky, obtain the risk assessment threshold and its corresponding handling strategy; If the matching degree of the target intent is greater than the risk assessment threshold, the disposal strategy shall be implemented.

[0084] Here, the preset intents in the intent library can all be risky intents, that is, only risky behaviors are modeled and identified; or, the preset intents in the intent library can include both normal intents and risky intents, which are distinguished by intent tags (such as "normal" or "risky"), so that when a normal intent is identified, the process can continue to be tracked or terminated, and when a risky intent is identified, the corresponding handling strategy is triggered.

[0085] It should be noted that different risk intentions may correspond to different risk levels. Therefore, different risk assessment thresholds and corresponding handling strategies can be preset for different risk intentions. The setting of the risk assessment threshold can be adjusted according to the severity of the handling measures: if the handling measures are relatively strong (such as direct blocking, refusing subsequent operations, banning accounts, etc.), the risk assessment threshold can be set higher to prioritize accuracy and avoid harming normal users; if the handling measures are relatively mild (such as popping up CAPTCHAs, increasing access costs, entering observation mode, etc.), the risk assessment threshold can be set lower to prioritize recall and discover as many potential risk intentions as possible.

[0086] In actual business operations, risk assessment thresholds can be dynamically adjusted according to business scenarios. For example, starting from observation mode, continuous iteration and optimization can be performed based on actual feedback. Different risk assessment thresholds correspond to different handling strategies. Specifically: when the matching degree exceeds a higher risk assessment threshold, it indicates that the current behavior highly matches a high-risk intent, and a strong handling strategy can be adopted, such as directly refusing subsequent operations, blocking the session, or reporting to the security center; when the matching degree exceeds a lower risk assessment threshold but does not reach a higher threshold, it indicates that the current behavior is suspected of having a risky intent but the confidence level is insufficient, and a mild handling strategy can be adopted, such as displaying a verification code, increasing operational costs, or entering a manual review queue.

[0087] The aforementioned tiered handling mechanism enables differentiated and refined handling, ensuring system security while minimizing interference with normal users, thus achieving a balance between security and user experience.

[0088] To better understand the above solutions, some specific examples are given based on the above embodiments.

[0089] For example, such as Figure 2 As shown, an intent recognition method provided in this application embodiment may include the following steps: S21. Collect interactive events according to a preset cycle to obtain an interactive event sequence; S22. Cache the interaction event sequence, and perform first feature extraction and second feature extraction on the interaction events in the interaction event sequence to obtain the first feature sequence and the second feature sequence; S23. Extract the first feature sequence of a preset length; S24. Perform human-machine recognition based on the first feature sequence to obtain a confidence score; S25. Determine the relationship between the confidence score and the first threshold and the second threshold; if the confidence score is less than the second threshold, proceed to S26; if the confidence score is greater than the first threshold, proceed to S27; if the confidence score is between the first threshold and the second threshold, proceed to S23. S26. Determine that the operator is a human; S27. Determine that the operating subject is an intelligent agent; S28. Intent recognition and tracking based on the second feature sequence; S29. If a risk intent is identified, risk management shall be carried out.

[0090] In this embodiment, in S23, the preset length is a fixed window size pre-set by the system, representing the number of events participating in recognition each time. If the number of events currently cached is greater than or equal to the preset length, the most recent interaction event of the preset length is taken, and its corresponding first feature vector is arranged in chronological order to form a first feature sequence. If the number of events currently cached is less than the preset length, interaction events are collected until the preset length of interaction events is reached before the first feature sequence is extracted. In addition, after extracting the first feature sequence of the preset length, the first feature sequence can be deleted from the cache to avoid reuse. When S23 is executed again later (for example, when the confidence score is in an ambiguous range and re-recognition is required), the system obtains the latest first feature sequence of the preset length extracted again based on the remaining interaction events in the cache and after collecting new events, thereby ensuring that the data participating in recognition each time is the latest unused event, avoiding the lack of incremental information in the recognition result due to the reuse of the same batch of data, and improving the accuracy of human-machine recognition.

[0091] This application embodiment establishes a more intelligent and secure identity and intent recognition mechanism at the AI ​​interaction entry point, which can promptly intercept malicious AI-initiated false registrations and attempts to invade user accounts, thereby ensuring the security and reliability of online services.

[0092] Based on the intent recognition method provided in the above embodiments, this application also provides specific implementations of the intent recognition device. Please refer to the following embodiments.

[0093] like Figure 3 As shown, an intent recognition device 300 provided in one embodiment of this application includes the following modules: The acquisition module 310 is used to acquire interaction data generated during the interaction between the operating subject and the operated object. The interaction data includes a sequence of interaction events, which includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. The extraction module 320 is used to extract the first feature from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, so as to obtain the first feature sequence. The recognition module 330 is used to recognize the sub-intent sequence of the intelligent agent based on the first feature sequence; The determination module 340 is used to determine the target intent of the agent based on the sub-intent sequence.

[0094] The intent recognition device 300 described above will be explained in detail below: In some embodiments, the event attribute is an extended static attribute. Based on this, the extraction module 320 is specifically used for: Encode the event type of the interactive event to obtain the event type code value; Encode the extended static properties of the interactive event to obtain the extended static property values; The event type encoding value and the extended static attribute value are concatenated to obtain the first feature vector corresponding to the interactive event; The first feature sequence is obtained by sorting the first feature vectors corresponding to the multiple interaction events according to their arrangement order in the interaction event sequence.

[0095] In some embodiments, the extended static attributes include basic static attributes and element semantic information of the interactive element corresponding to the interaction event. The element semantic information includes the functional semantics of the interactive element and the hierarchical path of the interactive element in the interactive interface. Based on this, the extraction module 320 is specifically used for: Encode the basic static properties to obtain their values; Encode the semantic information of the element to obtain the semantic encoded value of the element; By concatenating the basic static attribute value and the element semantic encoding value, the extended static attribute value is obtained.

[0096] In some embodiments, the first feature sequence includes multiple feature sequence segments, each feature sequence segment including at least one first feature vector. Based on this, the recognition module 330 is specifically used for: Obtain a library of preset sub-intents, which includes multiple preset sub-intents. Each preset sub-intent corresponds to a preset feature sequence segment. The preset feature sequence segment includes at least one preset feature vector. The preset feature vector is determined based on the event type of the interaction event and the extended static attributes. If a feature sequence segment matches a preset feature sequence segment, the preset sub-intention corresponding to the preset feature sequence segment will be identified as the sub-intention corresponding to the feature sequence segment. The identified sub-intents are arranged according to the order of the feature sequence segments in the first feature sequence to obtain the sub-intent sequence.

[0097] In some embodiments, the determining module 340 is specifically used for: Obtain the preset intent library, which includes multiple preset intents and their corresponding standard sub-intent sequences; Based on the sequence matching algorithm, the matching degree between the sub-intent sequence and each standard sub-intent sequence is determined; When the matching degree characterization sub-intent sequence matches the standard sub-intent sequence, the preset intent corresponding to the standard sub-intent sequence is determined as the target intent.

[0098] In some embodiments, the acquisition module 310 is further configured to, after determining the preset intent corresponding to the standard sub-intent sequence as the target intent, acquire the risk judgment threshold and its corresponding handling strategy when the target intent is a risky intent; The intent recognition device 300 may further include an execution module, used to execute a handling strategy when the matching degree corresponding to the target intent is greater than the risk judgment threshold.

[0099] In some embodiments, the interaction data also includes interaction protocol data. Based on this, the determining module 340 is further configured to determine that the operating subject is an intelligent agent before performing first feature extraction on the interaction events in the interaction event sequence, provided that the interaction protocol data includes an intelligent agent identifier, when the operating subject is an intelligent agent.

[0100] In some embodiments, the extraction module 320 is further configured to perform a first feature extraction on the interactive events in the interactive event sequence to obtain a second feature sequence before obtaining a first feature sequence, when the operating subject is an intelligent agent; The identification module 330 is also used to identify the operating subject based on the second feature sequence.

[0101] In some embodiments, the event attributes include basic static attributes, such as coordinates and a timestamp. Based on this, the extraction module 320 is specifically used for: Encode the event type of the interactive event to obtain the event type code value; Encode the basic static properties of interactive events to obtain the basic static property values; Based on the coordinates and timestamps of interactive events, as well as the coordinates and timestamps of adjacent interactive events, dynamic attribute values ​​are calculated. The dynamic attribute values ​​include at least one of instantaneous velocity, acceleration, jerk, trajectory curvature, and scroll length. The event type encoding value, basic static attribute value, and dynamic attribute value are concatenated to obtain the second feature vector corresponding to the interactive event; The second feature sequence is obtained by sorting the second feature vectors corresponding to the multiple interaction events according to their arrangement in the interaction event sequence.

[0102] In some embodiments, the operating entities include intelligent agents and humans. Based on this, the identification module 330 is specifically used for: The subject recognition model identifies the operator based on the second feature sequence and outputs a confidence score that the operator is an agent. The confidence score ranges from 0 to 1, and the higher the score, the greater the probability that the operator is an agent. The subject recognition model is trained based on the agent feature sequence and its corresponding agent label, and the human feature sequence and its corresponding human label. If the confidence score is greater than the first threshold, the operator is determined to be an intelligent agent; If the confidence score is less than the second threshold, the operator is determined to be human. If the confidence score is between the first threshold and the second threshold, return to the process of interaction between the operator and the operated object, and obtain the interaction data generated during the interaction process until the confidence score is greater than the first threshold or less than the second threshold.

[0103] In some embodiments, the subject recognition model includes a bottleneck layer, multiple feature extraction layers of different lengths, a feature fusion layer, a temporal attention layer, and a fully connected layer. The temporal attention layer includes a feedforward neural network layer and a weighted pooling layer. Based on this, the recognition module 330 is specifically used for: The second feature sequence is reduced in dimension by using the bottleneck layer to obtain the reduced feature sequence. By using multiple feature extraction layers of different lengths, depthwise separable convolutions are performed on the dimensionality-reduced feature sequences to extract temporal features at different scales, resulting in multiple first feature maps. Through the feature fusion layer, multiple first feature maps are fused and transformed to obtain a second feature map. The second feature map includes T time steps, each time step corresponding to a fused feature vector, where T is a positive integer. The attention weights are calculated for the feature vectors at each time step through the feedforward neural network layer of the attention layer, resulting in T attention weights. The third feature map is obtained by weighted summation of the feature vectors at T time steps through the weighted pooling layer of the attention layer, based on T attention weights. The third feature map is processed by a linear transformation and activation function through a fully connected layer, and the output is a confidence score of the agent.

[0104] In this embodiment, since the interaction event sequence includes multiple interaction events arranged in chronological order, and each interaction event includes an event type and an event attribute, by extracting a first feature from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, a first feature sequence is obtained. This first feature sequence preserves the event type information and event attribute information of the intelligent agent during the interaction process, as well as the temporal relationship between events. Based on this, by identifying the sub-intent sequence of the intelligent agent based on the first feature sequence, continuous interaction events can be converted into a sub-intent sequence with specific semantics. By determining the target intent of the intelligent agent based on this sub-intent sequence with specific semantics, the target intent of the intelligent agent during task execution can be accurately identified. Thus, through this embodiment, the intent of the intelligent agent can be accurately identified, providing a reliable basis for subsequently distinguishing between benevolent and malicious intelligent agents, thereby ultimately ensuring the security and reliability of online services.

[0105] Based on the intent recognition method provided in the above embodiments, this application also provides specific implementation methods for electronic devices. Figure 4 A schematic diagram of the structure of an electronic device provided in one embodiment of this application is shown.

[0106] like Figure 4 As shown, the electronic device 400 may include a processor 410 and a memory 420 storing computer program instructions.

[0107] Specifically, the processor 410 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0108] Memory 420 may include mass storage for data or instructions. For example, and not limitingly, memory 420 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where suitable, memory 420 may include removable or non-removable (or fixed) media. Where suitable, memory 420 may be internal or external to electronic device 400. In a particular embodiment, memory 420 is a non-volatile solid-state memory.

[0109] In specific embodiments, the memory 420 may be implemented as a read-only memory (ROM), random access memory (RAM), static storage device, dynamic storage device, etc. The memory 420 may store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 420 and executed by the processor 410. The processor 410 implements any of the intent recognition methods in the above embodiments by reading and executing the computer program instructions stored in the memory 420.

[0110] The processor 410 implements any of the intent recognition methods described in the above embodiments by reading and executing computer program instructions stored in the memory 420.

[0111] In one example, the electronic device 400 may also include a communication interface 430 and a bus 440. For example, Figure 4 As shown, the processor 410, memory 420, and communication interface 430 are connected through bus 440 and complete communication with each other.

[0112] The communication interface 430 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0113] Bus 440 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 440 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0114] For example, the electronic device 400 can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.

[0115] The electronic device can execute the intent recognition method in the embodiments of this application, thereby achieving the combination Figures 1 to 2 The intent recognition method described, and the beneficial effects of the corresponding method embodiments, will not be elaborated further here.

[0116] Furthermore, in conjunction with the intent recognition methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the intent recognition methods in the above embodiments. Examples of such computer-readable storage media include non-transitory computer-readable storage media, such as read-only memory (ROM).

[0117] The computer program instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the intent recognition method as shown in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0118] Based on the intent recognition methods described in the above embodiments, this application provides a computer program product for implementation. When the instructions in this computer program product are executed by the processor of an electronic device, they implement any of the intent recognition methods described in the above embodiments.

[0119] The computer program products of the above embodiments are used to implement the intent recognition method shown in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0120] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0121] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0122] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0123] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0124] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. An intent recognition method, characterized in that, include: During the interaction between the operating subject and the operated object, the interaction data generated by the interaction process is acquired. The interaction data includes a sequence of interaction events, which includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. When the operating subject is an intelligent agent, the first feature is extracted from the interaction events in the interaction event sequence to obtain the first feature sequence; Based on the first feature sequence, identify the agent's sub-intention sequence; Based on the sub-intent sequence, the target intent of the agent is determined.

2. The method according to claim 1, characterized in that, The event attribute is an extended static attribute. The step of extracting a first feature sequence from the interactive events in the interactive event sequence includes: The event type of the interaction event is encoded to obtain the event type encoding value; The extended static attributes of the interactive event are encoded to obtain the extended static attribute values; The event type encoding value and the extended static attribute value are concatenated to obtain the first feature vector corresponding to the interactive event; The first feature sequence is obtained by sorting the first feature vectors corresponding to the multiple interactive events according to their arrangement order in the interactive event sequence.

3. The method according to claim 2, characterized in that, The extended static attributes include basic static attributes and element semantic information of the interactive element corresponding to the interactive event. The element semantic information includes the functional semantics of the interactive element and the hierarchical path of the interactive element in the interactive interface. Encoding the extended static attributes of the interactive event to obtain extended static attribute values ​​includes: The basic static attributes are encoded to obtain the basic static attribute values; The semantic information of the element is encoded to obtain the element semantic encoding value; The extended static attribute value is obtained by concatenating the basic static attribute value and the element semantic encoding value.

4. The method according to claim 1, characterized in that, The first feature sequence includes multiple feature sequence segments, each of which includes at least one first feature vector. The step of identifying the agent's sub-intent sequence based on the first feature sequence includes: Obtain a preset sub-intent library, which includes multiple preset sub-intents. Each preset sub-intent corresponds to a preset feature sequence segment. The preset feature sequence segment includes at least one preset feature vector. The preset feature vector is determined based on the event type of the interaction event and the extended static attributes. If the feature sequence segment matches the preset feature sequence segment, the preset sub-intention corresponding to the preset feature sequence segment is identified as the sub-intention corresponding to the feature sequence segment; The identified sub-intents are arranged according to the order of the feature sequence segments in the first feature sequence to obtain a sub-intent sequence.

5. The method according to claim 4, characterized in that, Determining the target intent of the agent based on the sub-intent sequence includes: Obtain a preset intent library, which includes multiple preset intents and their corresponding standard sub-intent sequences; Based on a sequence matching algorithm, the matching degree between the sub-intent sequence and each of the standard sub-intent sequences is determined; When the matching degree indicates that the sub-intent sequence matches the standard sub-intent sequence, the preset intent corresponding to the standard sub-intent sequence is determined as the target intent.

6. The method according to claim 5, characterized in that, After determining the preset intent corresponding to the standard sub-intent sequence as the target intent, the method further includes: If the target intent is a risk intent, obtain the risk assessment threshold and its corresponding handling strategy; If the matching degree corresponding to the target intent is greater than the risk assessment threshold, the handling strategy is executed.

7. The method according to any one of claims 1-6, characterized in that, The interaction data also includes interaction protocol data. When the operating entity is an intelligent agent, before performing the first feature extraction on the interaction events in the interaction event sequence, the method further includes: If the interaction protocol data includes an agent identifier, the operating entity is identified as an agent.

8. The method according to any one of claims 1-6, characterized in that, Before extracting the first feature sequence from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, the method further includes: The second feature sequence is obtained by extracting a second feature from the interaction events in the interaction event sequence; The operating entity is identified based on the second feature sequence.

9. The method according to claim 8, characterized in that, The event attributes include basic static attributes, which include coordinate position and timestamp. The second feature extraction from the interactive events in the interactive event sequence to obtain a second feature sequence includes: The event type of the interaction event is encoded to obtain the event type encoding value; The basic static attributes of the interactive event are encoded to obtain the basic static attribute values; Based on the coordinate position and timestamp of the interactive event, as well as the coordinate position and timestamp of adjacent interactive events, dynamic attribute values ​​are calculated. The dynamic attribute values ​​include at least one of instantaneous velocity, acceleration, jerk, trajectory curvature, and scroll length. The event type encoding value, the basic static attribute value, and the dynamic attribute value are concatenated to obtain the second feature vector corresponding to the interactive event; The second feature sequence is obtained by sorting the second feature vectors corresponding to the multiple interactive events according to their arrangement order in the interactive event sequence.

10. The method according to claim 8, characterized in that, The operating subject includes intelligent agents and humans, and the identification of the operating subject based on the second feature sequence includes: The subject recognition model identifies the operating subject based on the second feature sequence and outputs a confidence score that the operating subject is an intelligent agent. The confidence score ranges from 0 to 1, and the higher the score, the greater the probability that the operating subject is an intelligent agent. The subject recognition model is trained based on the intelligent agent feature sequence and its corresponding intelligent agent label, and the human feature sequence and its corresponding human label. If the confidence score is greater than the first threshold, the operator is determined to be an intelligent agent; If the confidence score is less than the second threshold, the operator is determined to be human. If the confidence score is between the first threshold and the second threshold, the process of obtaining the interaction data generated during the interaction between the operating subject and the operated object is returned until the confidence score is greater than the first threshold or less than the second threshold.

11. The method according to claim 10, characterized in that, The subject recognition model includes a bottleneck layer, multiple feature extraction layers of different lengths, a feature fusion layer, a temporal attention layer, and a fully connected layer. The temporal attention layer includes a feedforward neural network layer and a weighted pooling layer. The subject recognition model identifies the operating subject based on the second feature sequence and outputs a confidence score indicating that the operating subject is an agent, including: The second feature sequence is reduced in dimension by the bottleneck layer to obtain the reduced feature sequence. By using the multiple feature extraction layers of different lengths, depthwise separable convolutions are performed on the dimensionality-reduced feature sequences to extract temporal features at different scales, resulting in multiple first feature maps. The feature fusion layer performs channel fusion and transformation on the multiple first feature maps to obtain a second feature map. The second feature map includes T time steps, and each time step corresponds to a fused feature vector, where T is a positive integer. The attention weights are calculated through the feedforward neural network layer of the attention layer for each time step of the feature vector, resulting in T attention weights. The third feature map is obtained by weighted summing the feature vectors at the T time steps through the weighted pooling layer of the attention layer, based on the T attention weights. The fully connected layer performs linear transformation and activation function processing on the third feature map to output the confidence score of the operator as an agent.

12. An intent recognition device, characterized in that, The device includes: The acquisition module is used to acquire interaction data generated during the interaction between the operating subject and the operated object. The interaction data includes a sequence of interaction events, which includes multiple interaction events arranged in chronological order. Each interaction event includes an event type and an event attribute. The extraction module is used to extract first features from the interaction events in the interaction event sequence when the operating subject is an intelligent agent, so as to obtain a first feature sequence; The identification module is used to identify the sub-intent sequence of the agent based on the first feature sequence; A determination module is used to determine the target intent of the agent based on the sub-intent sequence.

13. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the intent recognition method as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the intent recognition method as described in any one of claims 1-11.

15. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the intent recognition method as described in any one of claims 1-11.