A data processing method, device and equipment
By vectorizing and clustering user behavior sequence data, and combining it with language model analysis, the problem of inconvenient understanding of user behavior in existing technologies is solved, and efficient category label generation and optimization of risk prevention and control systems are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-05
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to effectively understand and label experience issues within user behavior sequence data, leading to inconvenience in the user deregulation process within risk control systems.
By vectorizing and clustering user behavior sequence data, and using a pre-trained language model to understand the clusters, similarities in operational behaviors are extracted, generating high-quality category label information.
It improves the efficiency of obtaining category label information, reduces the acquisition cost, and enhances the accuracy and efficiency of understanding user behavior in the risk prevention and control system.
Smart Images

Figure CN117290735B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a data processing method, apparatus, and device. Background Technology
[0002] As people become increasingly concerned about their privacy, risk control is becoming more and more important. Typically, risk control systems determine whether a user poses a risk based on user profiles and operational behavior, and issue different types of penalties to users based on the identified risk type. Penalized users may experience restrictions on one or more functions in certain scenarios (such as payments, receiving payments, and transfers). Users can choose automatic unblocking upon expiration or check the unblocking methods for self-unblocking. To understand whether users' actual behavioral sequence data meets unblocking expectations, identify experience issues in behavioral sequence data, and obtain corresponding tags, a better technical solution is needed for acquiring tags and understanding behavioral sequence data. Summary of the Invention
[0003] The purpose of the embodiments in this specification is to provide a technical solution for a better understanding mechanism for acquiring tag and behavioral sequence data.
[0004] To achieve the above technical solution, the embodiments in this specification are implemented as follows:
[0005] This specification provides a data processing method, comprising: acquiring behavioral sequence data generated during the execution of a target service by multiple different users within a preset time period; determining representation information capable of characterizing each behavioral sequence data; clustering the behavioral sequence data based on the determined representation information to obtain one or more different clusters; using operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information; inputting the prompt information and the obtained clusters into a pre-trained language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters; and determining category label information corresponding to different clusters based on the understanding information and / or intent information of the operation behavior corresponding to different clusters.
[0006] This specification provides a data processing apparatus, comprising: a behavior data acquisition module for acquiring behavior sequence data generated during the execution of a target service by multiple different users within a preset time period; a clustering module for determining representation information that can characterize each behavior sequence data, and performing clustering processing on the behavior sequence data based on the determined representation information to obtain one or more different clusters; an understanding module for using operation purpose information and operation intent information corresponding to behavior sequence data with a similarity greater than a preset similarity threshold, as well as behavior sequence data belonging to the same cluster, as prompt information, and inputting the prompt information and the obtained clusters into a pre-trained language model to obtain understanding information and / or intent information of operation behaviors corresponding to different clusters; and a label determination module for determining category label information corresponding to different clusters based on the understanding information and / or intent information of operation behaviors corresponding to different clusters.
[0007] This specification provides a data processing device comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, cause the processor to: acquire behavioral sequence data generated during the execution of a target service triggered by multiple different users within a preset time period; determine representation information capable of characterizing each behavioral sequence data; perform clustering processing on the behavioral sequence data based on the determined representation information to obtain one or more different clusters; use operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information; input the prompt information and the obtained clusters into a pre-trained language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters; and determine category label information corresponding to different clusters based on the understanding information and / or intent information of the operation behavior corresponding to different clusters.
[0008] This specification also provides a storage medium for storing computer-executable instructions. When executed by a processor, these instructions implement the following process: acquiring behavioral sequence data generated during the execution of a target service triggered by multiple different users within a preset time period; determining representational information capable of characterizing each behavioral sequence data; clustering the behavioral sequence data based on the determined representational information to obtain one or more different clusters; using the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information; inputting the prompt information and the obtained clusters into a pre-trained language model to obtain understanding information and / or intent information of the operational behavior corresponding to different clusters; and determining category label information corresponding to different clusters based on the understanding information and / or intent information of the operational behavior corresponding to different clusters. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is an embodiment of a data processing method described in this specification;
[0011] Figure 2 This is another embodiment of a data processing method described in this specification;
[0012] Figure 3 This is a schematic diagram of a data processing procedure described in this specification;
[0013] Figure 4 This is yet another embodiment of a data processing method described in this specification;
[0014] Figure 5 This is a schematic diagram illustrating another data processing procedure described in this specification;
[0015] Figure 6 This is yet another embodiment of a data processing method described in this specification;
[0016] Figure 7 This specification provides an embodiment of a data processing apparatus.
[0017] Figure 8 This is an embodiment of a data processing device described in this specification. Detailed Implementation
[0018] This specification provides a data processing method, apparatus, and device through its embodiments.
[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0020] This specification provides a mechanism for understanding and perceiving behavioral sequence data. This mechanism can be applied to risk control systems, which determine whether a user poses a risk based on user profiles and operational behaviors. Based on the determined risk type, different types of penalties are issued to the corresponding user. Penalized users will be subject to restrictions on one or more functions in certain scenarios (such as payment, receipt, and transfer). Users can choose automatic unrestriction upon expiration or view unrestriction methods for self-service unrestriction. The self-service unrestriction process may include verifying identity information and uploading unrestriction certificates. As unrestriction methods become more diverse and unrestriction processes are adjusted, the types of nodes involved in the process may increase or decrease accordingly. Ultimately, the user's behavioral sequence data consists of a discrete sequence of nodes traversed by the user's actions at that time.
[0021] To understand whether users' actual behavioral sequence data meets the expected outcome of the penalty process and to identify experience issues within the behavioral sequence data, a behavioral sequence data perception scheme was designed based on the aforementioned objectives. First, the behavioral sequence data of users penalized daily is vectorized. Then, the vectorized behavioral sequence data is clustered. Finally, the resulting clusters are analyzed. Typically, images are more intuitive than text, and text is more sequential. The clusters in the clustering results are in sequence form, which greatly complicates cluster analysis. Furthermore, the amount of data within the same cluster varies, making analysis of larger clusters even more difficult. How to more effectively describe clusters and extract key information from them is crucial for understanding behavioral sequence data and identifying experience issues in behavioral sequence data design. Therefore, the above processing method incorporates a cluster understanding process. The resulting clusters are input as prompts into a large language model, indicating that behaviors within clusters are similar, and the similarities in behavior are extracted. The model extracts and summarizes the textual meaning of similar behavioral sequence data within the same cluster, presenting the content of the clusters in a more intuitive way. Considering the data security issues associated with using external large language models to extract text, model distillation can be used to deploy the distilled model internally, thus avoiding the risk of data leakage. Specific details can be found in the following embodiments.
[0022] Example 1
[0023] like Figure 1 As shown in the embodiments of this specification, a data processing method is provided. The execution subject of this method can be a terminal device or a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or tablet computer, or a computer device such as a laptop or desktop computer, or an IoT device (specifically, a smartwatch, in-vehicle device, etc.). The server can be a single server or a server cluster composed of multiple servers, etc. The server can be a backend server for financial business or online shopping business, or a backend server for an application, etc. This embodiment uses a server as the execution subject for detailed description. For the case where the execution subject is a terminal device, please refer to the following server case processing, which will not be repeated here. The method may specifically include the following steps:
[0024] In step S102, behavioral sequence data generated during the execution of target services by multiple different users within a preset time period are obtained.
[0025] The preset duration can be set according to actual conditions, such as 1 day (i.e., 24 hours) or 7 days. The user can be any user capable of triggering the execution of the target service. The target service can include various types, such as lifting payment restrictions, unfreezing frozen accounts, or reporting risks, etc., which can be set according to actual conditions. This specification does not limit this. Behavioral sequence data can be data arranged in chronological order of various behaviors generated during the execution of the target service. Behavioral sequence data can be presented in the form of text strings or strings, for example: pushing a penalty notification to the user—the user views the lifting plan—the user uploads the lifting certificate—lifting successful. In addition, behavioral sequence data can also be constructed in other ways. For example, to more conveniently and intuitively display the behavioral sequence data, it can be presented through behavioral flow. The behavioral flow can refer to users who are penalized by the risk control system and are therefore unable to make payments, receive payments, transfer funds, etc. The sequence of action nodes during the unblocking process includes actions such as viewing the unblocking plan, uploading unblocking credentials, requesting manual assistance, and successful unblocking. Specifically, it can be represented as B1>C1>C3>F10>F6>E7>C1, or B1>C1>C4>F6>F10>E7>C1, where B1 indicates pushing a penalty notification to the user, C1 indicates viewing the unblocking plan, C3 indicates self-service unblocking, C4 indicates requesting manual assistance, F6 indicates the user uploading unblocking credentials, F10 indicates the unblocking request, and E7 indicates successful unblocking.
[0026] In implementation, behavioral sequence data generated during the execution of target services by multiple different users can be obtained in various ways. For example, the server of the target service can record various operational behaviors generated during the execution of target services by different users, and can arrange each operational behavior in chronological order to obtain the corresponding behavioral sequence data. In addition, in order to simplify and intuitively display the above behavioral sequence data, the above behavioral sequence data can be recorded in the form of behavioral movement lines. For this purpose, various operational behaviors can be encoded according to the actual situation, and each code corresponds to an action node. For example, code B1 can be used to represent the action node corresponding to pushing a penalty notification to the user, etc. The specific settings can be set according to the actual situation, and this embodiment of the specification does not limit this. When it is necessary to understand, judge the intent of, or label the behavioral sequence data, behavioral sequence data generated during the execution of target services by multiple different users within a preset time period can be obtained from the data recorded above. For example, behavioral sequence data generated within 24 hours before the current time can be obtained. For example, the obtained behavioral sequence data (represented by behavioral movement lines) may include B1>C1>C3>F10>F6>E7>C1, B1>C1>C4>F6>F10>E7>C1...
[0027] In addition, behavioral sequence data generated during the execution of target services by multiple different users within a preset time period can be obtained from a specified database. For example, behavioral sequence data generated within 7 days prior to the current moment can be obtained. The specific settings can be configured according to the actual situation.
[0028] In step S104, characterization information that can characterize each behavioral sequence data is determined, and the behavioral sequence data is clustered based on the determined characterization information to obtain one or more different clusters.
[0029] The representation information can include various types, such as vector representation or matrix representation, and can be set according to the actual situation. This specification does not limit this in the embodiments.
[0030] In implementation, to facilitate subsequent processing of the behavioral sequence data, each behavioral sequence data can be transformed into representational information that can characterize the corresponding behavioral sequence data. For example, each behavioral sequence data can be vectorized to convert it into a corresponding vector; or, each behavioral sequence data can be converted into a matrix of a preset order to obtain a matrix of a preset order corresponding to each behavioral sequence data; or, each behavioral sequence data can be converted into a number string of a specified structure to obtain a number string of a specified structure corresponding to each behavioral sequence data, etc. The specific settings can be determined according to the actual situation, and this specification does not limit this aspect in the embodiments.
[0031] To uncover potential experience problems in behavioral sequence data, clustering can be performed on the data based on multiple defined representations. Specifically, a clustering algorithm can be pre-defined according to the actual situation. This clustering algorithm is an unsupervised learning algorithm. The goal of the clustering algorithm is to group the data in the dataset into several clusters according to a certain similarity metric. Data within the same cluster have high similarity (i.e., the similarity of data within the same cluster is higher than a first preset threshold), while data between different clusters have low similarity (i.e., the similarity of data within the same cluster is lower than a second preset threshold, which can be the same as or different from the first preset threshold). Various clustering algorithms can be used, such as the Birch algorithm, Chameleon algorithm, DBSCAN algorithm, and OPTICS algorithm. This clustering algorithm can be used to perform clustering calculations on multiple defined representations, thereby grouping similar or identical behavioral sequence data together to obtain one or more different clusters.
[0032] In step S106, the operation purpose information and operation intention information corresponding to the behavior sequence data with a similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0033] The language model can include various types, such as the Large Language Model (LLM), the BERT model, or models based on the Transformer architecture. The LLM can be a natural language processing model composed of hundreds of millions of parameters, based on deep learning algorithms, and pre-trained using large-scale corpora for text understanding, generation, and conversion. LLM models can include GPT-3, ERNIE Bot, etc., and the specific model can be set according to the actual situation. The similarity threshold can also be set according to the actual situation, such as 80% or 95%. The prompt information can be a prompt, used to explain and interpret the data input to the model to indicate that the model should output the correct result.
[0034] In implementation, a corresponding algorithm can be acquired, and a language model can be built based on this algorithm. For example, a large language model can be built based on the Transformer architecture. The input data of this language model can be prompt information and cluster information. The cluster information can include relevant information about each member (such as identifier, data content, data volume, and interrelationship information). The output data can be understanding information and / or intent information of the operation behavior corresponding to different clusters. Then, training samples (i.e., information about clusters or groups with certain relationships, and corresponding prompt information) can be acquired for training the language model. The language model can be trained using these training samples. During the training process, a target function can be pre-defined, and the parameters in the language model can be optimized based on this target function to finally obtain the trained language model.
[0035] After obtaining clusters through the above method, corresponding prompts can be obtained, which may include operation purpose information and operation intent information corresponding to behavior sequence data with similarity greater than a preset similarity threshold, as well as behavior sequence data belonging to the same cluster. Then, the obtained clusters and the above prompts can be input into the trained language model. By analyzing and processing the above information through the trained language model, the understanding information and / or intent information of the operation behavior corresponding to different clusters can be obtained.
[0036] It should be noted that the prompt information may include the operation purpose information and operation intention information corresponding to the behavior sequence data with a similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster. In practical applications, the prompt information may include the operation purpose information and operation intention information corresponding to the behavior sequence data with a similarity greater than a preset similarity threshold, as well as one or two of the behavior sequence data belonging to the same cluster. The specific settings can be made according to the actual situation, and this specification does not limit this.
[0037] In step S108, based on the understanding information and / or intent information of the operation behavior corresponding to different clusters, the category label information corresponding to different clusters is determined.
[0038] The category label information can be a label used to indicate the category to which the cluster belongs. The category can include a variety of categories, such as a category with fraud risk, a category with illegal financial activity risk, or a category with no risk. The specific category can be set according to the actual situation, and the embodiments in this specification do not limit it.
[0039] In practice, the understanding information and / or intent information of the different clusters corresponding to the operation behaviors obtained above can be analyzed. Based on the analysis results, the category label information corresponding to the different clusters can be determined.
[0040] This specification provides a data processing method. It involves acquiring behavioral sequence data generated during the execution of a target service by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on this determined representational information, the behavioral sequence data is clustered to obtain one or more different clusters. Subsequently, the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, along with behavioral sequence data belonging to the same cluster, are used as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it allows for the determination of category label information corresponding to different clusters. In this way, behavioral sequence data is first processed using representational methods, and then clustered using clustering to group the behavioral sequence data with representational information. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0041] Example 2
[0042] like Figure 2As shown in the embodiments of this specification, a data processing method is provided. The execution subject of this method can be a terminal device or a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or tablet computer, or a computer device such as a laptop or desktop computer, or an IoT device (specifically, a smartwatch, in-vehicle device, etc.). The server can be a single server or a server cluster composed of multiple servers, etc. The server can be a backend server for financial business or online shopping business, or a backend server for an application, etc. This embodiment uses a server as the execution subject for detailed description. For the case where the execution subject is a terminal device, please refer to the following server case processing, which will not be repeated here. The method may specifically include the following steps:
[0043] In step S202, second behavior sequence data samples generated during the execution of target services by multiple different users are obtained.
[0044] The processing in step S202 described above can be obtained in a variety of different ways. For details, please refer to the relevant content in Embodiment 1 above. That is, the second row of sequence data samples can be obtained from the data recorded in the specified server, or the second row of sequence data samples can be obtained from the specified database, etc.
[0045] In step S204, second sample representation information that can characterize each second row sequence data sample is determined, and the second row sequence data samples are clustered based on the determined second sample representation information to obtain one or more different second sample clusters.
[0046] The processing of step S204 described above can be found in the relevant content of Embodiment 1 above, and will not be repeated here.
[0047] In practical applications, each second row of sequence data samples can be vectorized based on a pre-trained self-supervised model to obtain a vector corresponding to each second row of sequence data samples. The vector corresponding to each second row of sequence data samples is used as the second sample representation information corresponding to each second row of sequence data samples. The self-supervised model is obtained after training the model through self-supervised contrastive learning and / or the self-supervised model is constructed through a graph representation model.
[0048] Furthermore, based on a preset density-based clustering algorithm or a preset hierarchical clustering algorithm, the second row of sequential data samples can be clustered using determined second sample representation information to obtain one or more different second sample clusters. The density-based clustering algorithm can include DBSCAN, OPTICS, DENCLUE, etc., while the hierarchical clustering algorithm can include DIANA, BIRCH, Chameleon, etc. The specific algorithm can be chosen according to actual conditions, and this specification does not limit this. The specific processing procedure described above can be executed based on the specific algorithm used, and will not be elaborated further here.
[0049] In this embodiment, the language model can be a model obtained by performing model distillation on a generative large model, specifically as shown in step S206.
[0050] In step S206, the second sample category label information corresponding to each second sample cluster is obtained. The operation purpose information and operation intention information corresponding to the second line sequence data samples with similarity greater than a preset similarity threshold, as well as the second line sequence data samples belonging to the same second sample cluster, are used as sample prompt information. Based on the sample prompt information, the second sample cluster and the corresponding second sample category label information, the generative large model is used as the teacher model and the language model is used as the student model. The student model is trained by knowledge distillation through the teacher model to obtain the trained language model.
[0051] Generative large models can include GPT models or LLaMa models, etc.
[0052] In implementation, such as Figure 3 As shown, the second sample category label information corresponding to each second sample cluster can be obtained through expert experience and other methods. Then, the second sample cluster, the sample hint information corresponding to the second sample cluster, and the corresponding second sample category label information can be input into the language model as the student model. The output of the language model can be adjusted by the generative large model as the teacher model, and the model parameters of the language model can be adjusted by a preset loss function to train the language model until the preset loss function converges, thus obtaining the trained language model. The specific model training process described above is only one optional process. The language model can also be trained in many different ways, which can be set according to the actual situation. This specification does not limit this embodiment.
[0053] In step S208, behavioral sequence data generated during the execution of target services by multiple different users within a preset time period is obtained.
[0054] In step S210, each behavior sequence data is vectorized based on the pre-trained self-supervised model to obtain the vector corresponding to each behavior sequence data. The vector corresponding to each behavior sequence data is used as the representation information corresponding to each behavior sequence data. The self-supervised model is trained by self-supervised contrastive learning and / or the self-supervised model is constructed by graph representation model.
[0055] In implementation, a self-supervised model can be trained using self-supervised contrastive learning to obtain the trained self-supervised model. Alternatively, a self-supervised model can be constructed using graph representation models. In practical applications, besides the methods mentioned above, self-supervised models can also be constructed using various other methods, which can be set according to the actual situation. This specification does not limit the specific methods used in the embodiments. Figure 3 As shown, after constructing a self-supervised model in the above manner, the trained self-supervised model can be used to vectorize each action sequence data to obtain the vector corresponding to each action sequence data (e.g., ...). Figure 3 In this context, behavioral sequence data (represented by behavioral movement lines) are B1>C1>C3>F10>F6>E7>C1 and B1>C1>C4>F6>F10>E7>C1, etc., with corresponding vectors such as [0.58,0.05,-1.51,0.11,-0.24,0.28] and [0.62,-0.04,-1.41,0.09,-0.22,0.29], etc. The vector corresponding to a certain behavioral sequence data can be used as the representation information corresponding to that behavioral sequence data, thereby obtaining the representation information (i.e., the vector corresponding to the corresponding behavioral sequence data) for each behavioral sequence data.
[0056] In step S212, the behavioral sequence data is clustered based on the determined multiple characterization information to obtain one or more different clusters.
[0057] In implementation, such as Figure 3 As shown, multiple different clusters can be obtained through clustering.
[0058] In step S214, the operation purpose information and operation intention information corresponding to the behavior sequence data with similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0059] Based on the processing in step S214, such as Figure 3As shown, by performing cluster understanding on different clusters (i.e., by using a pre-trained language model and performing cluster understanding based on the obtained clusters and corresponding prompts), we can obtain the understanding information and / or intent information of the operation behavior corresponding to different clusters.
[0060] In step S216, the understanding information and / or intent information of the operation behavior corresponding to different clusters are provided to the target terminal.
[0061] The target terminal can be a terminal device used by experts or managers, such as a mobile terminal device like a mobile phone or tablet, a computer device like a laptop or desktop computer, or an IoT device (such as a smartwatch or in-vehicle device).
[0062] In implementation, after the target terminal receives the understanding information and / or intent information of the operation behavior corresponding to different clusters, the relevant experts or managers can analyze the understanding information and / or intent information of the operation behavior corresponding to different clusters. Through the above simple analysis, the category label information corresponding to different clusters can be obtained.
[0063] In step S218, the target terminal sends category label information corresponding to different clusters.
[0064] This specification provides a data processing method. It involves acquiring behavioral sequence data generated during the execution of a target service by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on this determined representational information, the behavioral sequence data is clustered to obtain one or more different clusters. Subsequently, the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, along with behavioral sequence data belonging to the same cluster, are used as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it allows for the determination of category label information corresponding to different clusters. In this way, behavioral sequence data is first processed using representational methods, and then clustered using clustering to group the behavioral sequence data with representational information. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0065] To understand whether users' actual behavioral patterns align with the product's expectations and to identify experience issues within these patterns, the behavioral patterns of penalized users each day (i.e., behavioral sequence data) are first vectorized. Then, the vectorized behavioral patterns are clustered, and finally, the resulting clusters are analyzed. To more effectively describe the clusters and extract key information from them—crucial for understanding behavioral patterns and identifying experience problems in their design—a cluster understanding process is added. This involves extracting similarities in operational behaviors within the clusters generated by clustering. Considering the potential data security risks associated with using external large language models, this embodiment provides a model distillation mechanism or uses a pre-trained model. The distilled or finely tuned model is deployed internally, avoiding the risk of data leakage.
[0066] Example 3
[0067] like Figure 4 As shown in the embodiments of this specification, a data processing method is provided. The execution subject of this method can be a terminal device or a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or tablet computer, or a computer device such as a laptop or desktop computer, or an IoT device (specifically, a smartwatch, in-vehicle device, etc.). The server can be a single server or a server cluster composed of multiple servers, etc. The server can be a backend server for financial business or online shopping business, or a backend server for an application, etc. This embodiment uses a server as the execution subject for detailed description. For the case where the execution subject is a terminal device, please refer to the following server case processing, which will not be repeated here. The method may specifically include the following steps:
[0068] In step S402, the first behavior sequence data samples generated during the execution of the target service triggered by multiple different users are obtained.
[0069] The processing in step S402 described above can be obtained in a variety of different ways. For details, please refer to the relevant content in Embodiment 1 above. That is, the first row of sequence data samples can be obtained from the data recorded in the specified server, or the first row of sequence data samples can be obtained from the specified database, etc.
[0070] In step S404, first sample representation information that can characterize each first row of sequence data samples is determined, and the first row of sequence data samples are clustered based on the determined first sample representation information to obtain one or more different first sample clusters.
[0071] The processing of step S404 described above can be found in the relevant content of Embodiment 1 above, and will not be repeated here.
[0072] In practical applications, each first row of sequence data samples can be vectorized based on a pre-trained self-supervised model to obtain a vector corresponding to each first row of sequence data samples. The vector corresponding to each first row of sequence data samples is used as the first sample representation information corresponding to each first row of sequence data samples. The self-supervised model is obtained after training the model through self-supervised contrastive learning and / or the self-supervised model is constructed through a graph representation model.
[0073] Furthermore, based on a preset density-based clustering algorithm or a preset hierarchical clustering algorithm, the first row of sequential data samples can be clustered using determined first sample representation information to obtain one or more different first sample clusters. The density-based clustering algorithm may include DBSCAN, OPTICS, DENCLUE, etc., while the hierarchical clustering algorithm may include DIANA, BIRCH, Chameleon, etc. The specific algorithm can be chosen according to actual conditions, and this specification does not limit this. The specific processing described above can be executed based on the specific algorithm used, and will not be elaborated further here.
[0074] In step S406, the first sample category label information corresponding to each first sample cluster is obtained, and the pre-trained language model is fine-tuned based on the first sample cluster and the corresponding first sample category label information to obtain the trained language model.
[0075] The language model can include generative large models, which can include GPT models or LLaMa models, etc.
[0076] In implementation, such as Figure 5 As shown, the first sample category label information corresponding to each first sample cluster can be obtained through expert experience or other means. Then, the first sample cluster, the sample prompt information corresponding to the first sample cluster, and the corresponding first sample category label information can be input into the pre-trained language model to obtain the corresponding output results. Based on the output results and the first sample category label information, the corresponding loss information can be determined through a preset loss function. Based on the loss information, the model parameters of the pre-trained language model can be adjusted to train the pre-trained language model until the preset loss function converges, resulting in the trained language model. The above model training method is only one optional processing method. Various other methods can also be used to train the pre-trained language model. The specific method can be set according to the actual situation, and the embodiments in this specification do not limit this.
[0077] In step S408, behavioral sequence data generated during the execution of target services by multiple different users within a preset time period is obtained.
[0078] In step S410, each behavior sequence data is vectorized based on the pre-trained self-supervised model to obtain the vector corresponding to each behavior sequence data. The vector corresponding to each behavior sequence data is used as the representation information corresponding to each behavior sequence data. The self-supervised model is obtained after training the model through self-supervised contrastive learning and / or the self-supervised model is constructed through the graph representation model.
[0079] In implementation, such as Figure 5 As shown, after constructing a self-supervised model in the above manner, the trained self-supervised model can be used to vectorize each action sequence data to obtain the vector corresponding to each action sequence data (e.g., ...). Figure 5 In this context, behavioral sequence data (represented by behavioral movement lines) are B1>C1>C3>F10>F6>E7>C1 and B1>C1>C4>F6>F10>E7>C1, etc., with corresponding vectors such as [0.58,0.05,-1.51,0.11,-0.24,0.28] and [0.62,-0.04,-1.41,0.09,-0.22,0.29], etc. The vector corresponding to a certain behavioral sequence data can be used as the representation information corresponding to that behavioral sequence data, thereby obtaining the representation information (i.e., the vector corresponding to the corresponding behavioral sequence data) for each behavioral sequence data.
[0080] In step S412, the behavioral sequence data is clustered based on the determined multiple characterization information to obtain one or more different clusters.
[0081] In implementation, such as Figure 5 As shown, multiple different clusters can be obtained through clustering.
[0082] In step S414, the operation purpose information and operation intention information corresponding to the behavior sequence data with similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0083] Based on the processing in step S414, such as Figure 5 As shown, by performing cluster understanding on different clusters, we can obtain understanding information and / or intent information of the corresponding operational behaviors of different clusters.
[0084] In step S416, the understanding information and / or intent information of the operation behavior corresponding to different clusters are provided to the target terminal.
[0085] In step S418, the category label information corresponding to different clusters sent by the target terminal is received.
[0086] This specification provides a data processing method. It involves acquiring behavioral sequence data generated during the execution of a target service by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on this determined representational information, the behavioral sequence data is clustered to obtain one or more different clusters. Subsequently, the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, along with behavioral sequence data belonging to the same cluster, are used as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it allows for the determination of category label information corresponding to different clusters. In this way, behavioral sequence data is first processed using representational methods, and then clustered using clustering to group the behavioral sequence data with representational information. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0087] To understand whether users' actual behavioral patterns align with the product's expectations and to identify experience issues within these patterns, the behavioral patterns of penalized users each day (i.e., behavioral sequence data) are first vectorized. Then, the vectorized behavioral patterns are clustered, and finally, the resulting clusters are analyzed. To more effectively describe the clusters and extract key information from them—crucial for understanding behavioral patterns and identifying experience problems in their design—a cluster understanding process is added. This involves extracting similarities in operational behaviors within the clusters generated by clustering. Considering the potential data security risks associated with using external large language models, this embodiment provides a model distillation mechanism or uses a pre-trained model. The distilled or finely tuned model is deployed internally, avoiding the risk of data leakage.
[0088] Example 4
[0089] The following describes in detail a data processing method provided by the embodiments of this specification, in conjunction with specific application scenarios. The target business can be a business set up to lift restrictions when a user's preset permissions are restricted in risk control. The behavior sequence data can include one or more of the following: pushing a penalty notification to the user, viewing the lifting method, uploading the lifting certificate, manual assistance, and successful lifting. For ease of subsequent description, the target business can be described as the target lifting business, which can be any lifting business. The behavior sequence data is described as the lifting behavior flow.
[0090] like Figure 6 As shown in the embodiments of this specification, a data processing method is provided. The execution subject of this method can be a terminal device or a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or tablet computer, or a computer device such as a laptop or desktop computer, or an IoT device (specifically, a smartwatch, in-vehicle device, etc.). The server can be a single server or a server cluster composed of multiple servers, etc. The server can be a backend server for financial business or online shopping business, or a backend server for an application, etc. This embodiment uses a server as the execution subject for detailed description. For the case where the execution subject is a terminal device, please refer to the following server case processing, which will not be repeated here. The method may specifically include the following steps:
[0091] In step S602, data samples of the third unlocking behavior flow generated during the execution of the target unlocking service triggered by multiple different users are obtained.
[0092] The processing in step S602 above can be obtained in a variety of different ways. For details, please refer to the relevant content in Embodiment 1 above. That is, the data sample of the third delimitation behavior flow can be obtained from the data recorded in the specified server, or the data sample of the third delimitation behavior flow can be obtained from the specified database, etc.
[0093] In step S604, the data samples of each third unbound behavior movement line are vectorized based on the pre-trained self-supervised model to obtain the vector corresponding to the data sample of each third unbound behavior movement line. The vector corresponding to the data sample of each third unbound behavior movement line is used as the third sample representation information corresponding to the data sample of each third unbound behavior movement line. The self-supervised model is obtained after training the model through self-supervised contrastive learning and / or the self-supervised model is constructed through the graph representation model.
[0094] In step S606, based on a preset density-based clustering algorithm or a preset hierarchical clustering algorithm, the data samples of the third unbound behavior movement are clustered using the determined third sample characterization information to obtain one or more different third sample clusters.
[0095] Density-based clustering algorithms may include DBSCAN, OPTICS, and DENCLUE, while hierarchical clustering algorithms may include DIANA, BIRCH, and Chameleon. The specific algorithms can be chosen based on actual needs, and this specification does not limit their implementation. The detailed processing described above can be executed based on the specific algorithm used, and will not be elaborated further here.
[0096] In this embodiment, the language model can be illustrated by taking the model obtained by model distillation of a pre-deployed generative large model as an example. Specifically, it can be processed in step S608. In the case where the generative large model is an external large model, the model parameters of the pre-trained language model can be fine-tuned using the information such as the third sample cluster to obtain the trained language model. For details, please refer to the aforementioned related content, which will not be repeated here.
[0097] In step S608, the third sample category label information corresponding to each third sample cluster is obtained. The operation purpose information and operation intention information corresponding to the data samples of the third unrestricted behavior movement with similarity greater than a preset similarity threshold, as well as the data samples of the third unrestricted behavior movement belonging to the same third sample cluster, are used as sample prompt information. Based on the sample prompt information, the third sample cluster and the corresponding third sample category label information, the generative large model is used as the teacher model and the language model is used as the student model. The student model is trained by knowledge distillation through the teacher model to obtain the trained language model.
[0098] Generative large models can include GPT models or LLaMa models, etc.
[0099] In step S610, data on the unrestriction behavior flow generated during the execution of the target unrestriction service by multiple different users within a preset time period is obtained.
[0100] In step S612, the data of each unbound behavior movement line is vectorized based on the pre-trained self-supervised model to obtain the vector corresponding to the data of each unbound behavior movement line. The vector corresponding to the data of each unbound behavior movement line is used as the representation information corresponding to the data of each unbound behavior movement line. The self-supervised model is obtained after training the model through self-supervised contrastive learning and / or the self-supervised model is constructed through the graph representation model.
[0101] In step S614, the data of the unrestricted behavior flow are clustered based on the determined multiple characterization information to obtain one or more different clusters.
[0102] In step S616, the operation purpose information and operation intention information corresponding to the data of the unrestricted behavior movement lines with a similarity greater than a preset similarity threshold, as well as the data of the unrestricted behavior movement lines belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into the pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0103] In step S618, the understanding information and / or intent information of the operation behavior corresponding to different clusters are provided to the target terminal.
[0104] In step S620, the category label information corresponding to different clusters sent by the target terminal is received.
[0105] This specification provides a data processing method. It involves acquiring behavioral sequence data generated during the execution of a target service by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on this determined representational information, the behavioral sequence data is clustered to obtain one or more different clusters. Subsequently, the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, along with behavioral sequence data belonging to the same cluster, are used as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it allows for the determination of category label information corresponding to different clusters. In this way, behavioral sequence data is first processed using representational methods, and then clustered using clustering to group the behavioral sequence data with representational information. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0106] To understand whether users' actual behavioral patterns align with the product's expectations and to identify experience issues within these patterns, the behavioral patterns of penalized users each day (i.e., behavioral sequence data) are first vectorized. Then, the vectorized behavioral patterns are clustered, and finally, the resulting clusters are analyzed. To more effectively describe the clusters and extract key information from them—crucial for understanding behavioral patterns and identifying experience problems in their design—a cluster understanding process is added. This involves extracting similarities in operational behaviors within the clusters generated by clustering. Considering the potential data security risks associated with using external large language models, this embodiment provides a model distillation mechanism or uses a pre-trained model. The distilled or finely tuned model is deployed internally, avoiding the risk of data leakage.
[0107] Example 5
[0108] The above describes the data processing method provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a data processing apparatus, such as... Figure 7 As shown.
[0109] The data processing device includes: a behavioral data acquisition module 701, a clustering module 702, an understanding module 703, and a label determination module 704, wherein:
[0110] The behavior data acquisition module 701 acquires behavior sequence data generated during the execution of target services by multiple different users within a preset time period.
[0111] Clustering module 702 determines representation information that can characterize each behavioral sequence data, and performs clustering processing on the behavioral sequence data based on the determined multiple representation information to obtain one or more different clusters;
[0112] The understanding module 703 takes the operation purpose information and operation intention information corresponding to the behavior sequence data with similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0113] The label determination module 704 determines the category label information corresponding to different clusters based on the understanding information and / or intent information of the operation behavior corresponding to different clusters.
[0114] In this embodiment of the specification, the clustering module 702 performs vectorization processing on each behavior sequence data based on a pre-trained self-supervised model to obtain a vector corresponding to each behavior sequence data. The vector corresponding to each behavior sequence data is used as the representation information corresponding to each behavior sequence data. The self-supervised model is obtained after model training through self-supervised contrastive learning and / or the self-supervised model is constructed through a graph representation model.
[0115] In the embodiments of this specification, the clustering module 702 performs clustering processing on the behavioral sequence data based on a preset density-based clustering algorithm or a preset hierarchical clustering algorithm, using a plurality of determined representation information to obtain one or more different clusters.
[0116] In the embodiments of this specification, the language model includes a generative large model, which includes a GPT model or an LLaMa model, or the language model is a model obtained by performing model distillation on the generative large model.
[0117] In the embodiments described in this specification, the device further includes:
[0118] The first sample acquisition module acquires first action sequence data samples generated during the execution of target business by multiple different users.
[0119] The first sample clustering module determines first sample representation information that can characterize each first behavioral sequence data sample, and performs clustering processing on the first behavioral sequence data samples based on the determined first sample representation information to obtain one or more different first sample clusters.
[0120] The first training module obtains the first sample category label information corresponding to each first sample cluster, and fine-tunes the pre-trained language model based on the first sample cluster and the corresponding first sample category label information to obtain the trained language model.
[0121] In the embodiments described in this specification, the device further includes:
[0122] The second sample acquisition module acquires second behavior sequence data samples generated during the execution of target business by multiple different users.
[0123] The second sample clustering module determines second sample representation information that can characterize each second behavioral sequence data sample, and performs clustering processing on the second behavioral sequence data samples based on the determined second sample representation information to obtain one or more different second sample clusters.
[0124] The second training module acquires the second sample category label information corresponding to each second sample cluster. It uses the operation purpose information and operation intention information corresponding to the second action sequence data samples with similarity greater than a preset similarity threshold, as well as the second action sequence data samples belonging to the same second sample cluster, as sample prompt information. Based on the sample prompt information, the second sample cluster, and the corresponding second sample category label information, and using a generative large model as the teacher model and the language model as the student model, it performs knowledge distillation training on the student model through the teacher model to obtain the trained language model.
[0125] In this embodiment of the specification, the label determination module 704 includes:
[0126] The information sending unit provides the target terminal with the understanding information and / or intent information of the operation behavior corresponding to the different clusters;
[0127] The tag receiving unit receives category tag information corresponding to different clusters sent by the target terminal.
[0128] In the embodiments of this specification, the target service is a service set up to unblock a user's preset permissions when they are restricted in risk prevention and control. The behavior sequence data includes one or more of the following: pushing a penalty notification to the user, viewing the unblocking method, uploading unblocking credentials, manual assistance, and successful unblocking.
[0129] This specification provides a data processing apparatus that acquires behavioral sequence data generated during the execution of target services by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on the determined representational information, it performs clustering processing on the behavioral sequence data to obtain one or more different clusters. Subsequently, it uses operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it determines the category label information corresponding to different clusters. In this way, behavioral sequence data is first processed using representation, and then clustering is used to cluster the behavioral sequence data with representational information. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0130] To understand whether users' actual behavioral patterns align with the product's expectations and to identify experience issues within these patterns, the behavioral patterns of penalized users each day (i.e., behavioral sequence data) are first vectorized. Then, the vectorized behavioral patterns are clustered, and finally, the resulting clusters are analyzed. To more effectively describe the clusters and extract key information from them—crucial for understanding behavioral patterns and identifying experience problems in their design—a cluster understanding process is added. This involves extracting similarities in operational behaviors within the clusters generated by clustering. Considering the potential data security risks associated with using external large language models, this embodiment provides a model distillation mechanism or uses a pre-trained model. The distilled or finely tuned model is deployed internally, avoiding the risk of data leakage.
[0131] Example 6
[0132] The above describes the data processing apparatus provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a data processing device, such as... Figure 8 As shown.
[0133] The data processing device can be a terminal device or a server, as described in the above embodiments.
[0134] Data processing devices can vary considerably depending on configuration or performance, and may include one or more processors 801 and memory 802. Memory 802 may store one or more application programs or data. Memory 802 may be temporary or persistent storage. The application programs stored in memory 802 may include one or more modules (not shown), each module including a series of computer-executable instructions for the data processing device. Furthermore, processor 801 may be configured to communicate with memory 802 and execute the series of computer-executable instructions in memory 802 on the data processing device. The data processing device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, and one or more keyboards 806.
[0135] Specifically, in this embodiment, the data processing device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the data processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0136] Acquire behavioral sequence data generated during the execution of target services by multiple different users within a preset time period;
[0137] Determine the characterization information that can characterize each of the behavioral sequence data, and perform clustering processing on the behavioral sequence data based on the determined characterization information to obtain one or more different clusters;
[0138] The operation purpose information and operation intention information corresponding to the behavior sequence data with similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0139] Based on the understanding information and / or intent information of the operation behavior corresponding to different clusters, the category label information corresponding to different clusters is determined.
[0140] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the data processing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0141] This specification provides a data processing device that acquires behavioral sequence data generated during the execution of target services by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on the determined representational information, it performs clustering processing on the behavioral sequence data to obtain one or more different clusters. Subsequently, it uses the operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it determines the category label information corresponding to different clusters. In this way, the behavioral sequence data is first processed using a representational method, and then clustered into clusters using clustering processing. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0142] Example 7
[0143] Furthermore, based on the above Figures 1 to 6 The method shown in this specification, along with one or more embodiments, also provides a storage medium for storing computer-executable instruction information. In one specific embodiment, the storage medium can be a USB flash drive, optical disc, hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, it can achieve the following process:
[0144] Acquire behavioral sequence data generated during the execution of target services by multiple different users within a preset time period;
[0145] Determine the characterization information that can characterize each of the behavioral sequence data, and perform clustering processing on the behavioral sequence data based on the determined characterization information to obtain one or more different clusters;
[0146] The operation purpose information and operation intention information corresponding to the behavior sequence data with similarity greater than a preset similarity threshold, as well as the behavior sequence data belonging to the same cluster, are used as prompt information. The prompt information and the obtained clusters are input into a pre-trained language model to obtain the understanding information and / or intention information of the operation behavior corresponding to different clusters.
[0147] Based on the understanding information and / or intent information of the operation behavior corresponding to different clusters, the category label information corresponding to different clusters is determined.
[0148] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the above-described storage medium embodiment is basically similar to the method embodiment, so the description is relatively simple; relevant parts can be referred to the description of the method embodiment.
[0149] This specification provides a storage medium that acquires behavioral sequence data generated during the execution of target services by multiple different users within a preset time period. Then, it determines representational information that characterizes each behavioral sequence data. Based on the determined representational information, it performs clustering processing on the behavioral sequence data to obtain one or more different clusters. Subsequently, it uses operation purpose information and operation intent information corresponding to behavioral sequence data with a similarity greater than a preset similarity threshold, as well as behavioral sequence data belonging to the same cluster, as prompt information. This prompt information and the obtained clusters are input into a language model to obtain understanding information and / or intent information of the operation behavior corresponding to different clusters. Furthermore, it determines the category label information corresponding to different clusters. In this way, the behavioral sequence data is first processed using a representational method, and then clustered into clusters using clustering processing. This clustering method provides a foundation for discovering unknown experience problems. After clustering, the set prompts are input into the language model. By leveraging the cross-domain knowledge of the language model to identify similarities in operational behaviors within the clusters, the similarities in operational behaviors are extracted. This leads to an understanding of the operational behaviors and / or intentions of similar behavioral sequence data within the clusters, resulting in high-quality category label information. This significantly reduces the cost of acquiring category label information and improves the efficiency of category label information production.
[0150] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0151] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using a hardware physical module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0152] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0153] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0154] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0155] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] Embodiments in this specification are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable parallel device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable parallel device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable fraud device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0159] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0160] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0161] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0162] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0163] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0165] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0166] The above description is merely an embodiment of this specification and is not intended to limit this document. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A data processing method, the method comprising: obtaining behavior sequence data generated in a process of executing a target service triggered by a plurality of different users within a preset time period; determining representation information capable of representing each of the behavior sequence data, performing clustering processing on the behavior sequence data based on a preset clustering algorithm through the determined plurality of representation information, and obtaining one or more different clustering clusters, the representation information being capable of being represented by a vector, represented by a matrix, or represented by a digital string of a specified structure; generating prompt information based on operation purpose information corresponding to behavior sequence data with a similarity greater than a preset similarity threshold and operation intent information corresponding to behavior sequence data with a similarity greater than a preset similarity threshold, and behavior sequence data belonging to the same clustering cluster, inputting the prompt information and the obtained clustering cluster into a pre-trained language model, and obtaining understanding information and / or intent information of operation behavior corresponding to different clustering clusters, input data of the language model being prompt information and clustering cluster information, and output data being understanding information and / or intent information of operation behavior corresponding to different clustering clusters; determining class label information corresponding to different clustering clusters based on understanding information and / or intent information of operation behavior corresponding to different clustering clusters.
2. The method of claim 1, wherein the determining representation information capable of representing each of the behavior sequence data comprises: performing vectorization processing on each of the behavior sequence data based on a pre-trained self-supervised model to obtain a vector corresponding to each of the behavior sequence data, and taking the vector corresponding to each of the behavior sequence data as representation information corresponding to each of the behavior sequence data, the self-supervised model being obtained through model training in a self-supervised contrast learning manner and / or the self-supervised model being constructed through a graph representation model.
3. The method of claim 1, wherein the performing clustering processing on the behavior sequence data based on the determined plurality of representation information to obtain one or more different clustering clusters comprises: performing clustering processing on the behavior sequence data based on a preset density-based clustering algorithm or a preset hierarchical-based clustering algorithm through the determined plurality of representation information to obtain one or more different clustering clusters.
4. The method of claim 1, wherein the language model comprises a generative large model, the generative large model comprising a GPT model or an LLaMa model, or the language model is a model obtained through model distillation processing on the generative large model.
5. The method of claim 4, further comprising: obtaining first behavior sequence data samples generated in a process of executing a target service triggered by a plurality of different users; determining first sample representation information capable of representing each of the first behavior sequence data samples, performing clustering processing on the first behavior sequence data samples based on the determined plurality of first sample representation information, and obtaining one or more different first sample clustering clusters. Obtain the first sample category label information corresponding to each first sample clustering cluster, and fine-tune the pre-trained language model based on the first sample clustering cluster and the corresponding first sample category label information to obtain the trained language model.
6. The method of claim 4, further comprising: Obtaining second behavior sequence data samples generated in the process of executing the target service triggered by different users; Determine the second sample representation information capable of representing each of the second behavior sequence data samples, and perform clustering processing on the second behavior sequence data samples based on the determined plurality of second sample representation information to obtain one or more different second sample clustering clusters; Obtain the second sample category label information corresponding to each second sample clustering cluster, and the operation purpose information and operation intent information of the second behavior sequence data samples with a similarity greater than a preset similarity threshold, and the second behavior sequence data samples belonging to the same second sample clustering cluster as sample prompt information, and based on the sample prompt information, the second sample clustering cluster and the corresponding second sample category label information, and using the generative large model as the teacher model and the language model as the student model, the knowledge distillation training of the student model is performed through the teacher model to obtain the trained language model.
7. The method of claim 1, wherein the category label information corresponding to different clustering clusters is determined based on the understanding information and / or intent information of the operation behavior of the different clustering clusters, comprising: Providing the understanding information and / or intent information of the operation behavior of the different clustering clusters to the target terminal; Receiving the category label information corresponding to the different clustering clusters sent by the target terminal.
8. The method of any one of claims 1-7, wherein the target service is a limit removal service set when the user's preset permission is limited in risk prevention and control, and the behavior sequence data includes one or more of pushing a penalty notice to the user, viewing a limit removal method, uploading a limit removal credential, seeking human assistance, and successfully removing the limit.
9. A data processing apparatus, comprising: A behavior data acquisition module for acquiring behavior sequence data generated in the process of executing the target service triggered by different users within a preset time period; A clustering module for determining representation information capable of representing each of the behavior sequence data, and performing clustering processing on the behavior sequence data based on a pre-set clustering algorithm through the determined plurality of representation information to obtain one or more different clustering clusters, wherein the representation information can be represented by a vector, a matrix or a specified structure digital string. The understanding module generates prompt information based on the operation purpose information corresponding to the behavior sequence data with a similarity greater than a preset similarity threshold and the operation intention information corresponding to the behavior sequence data with a similarity greater than the preset similarity threshold and the behavior sequence data belonging to the same cluster, inputs the prompt information and the obtained cluster into a pre-trained language model, and obtains understanding information and / or intention information of operation behaviors corresponding to different clusters, wherein input data of the language model is prompt information and cluster information, and output data is understanding information and / or intention information of operation behaviors corresponding to different clusters. The label determination module determines category label information corresponding to different clusters based on the understanding information and / or intention information of operation behaviors corresponding to different clusters.
10. A data processing device, comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: obtain behavior sequence data generated in a process of triggering a target service by a plurality of different users within a preset time period; determine representation information capable of representing each of the behavior sequence data, perform clustering processing on the behavior sequence data based on a preset clustering algorithm through the determined representation information, and obtain one or more different clusters, wherein the representation information can be represented by a vector, a matrix, or a digital string with a specified structure; generate prompt information based on operation purpose information corresponding to behavior sequence data with a similarity greater than a preset similarity threshold and operation intention information corresponding to behavior sequence data with a similarity greater than the preset similarity threshold and behavior sequence data belonging to the same cluster, input the prompt information and the obtained cluster into a pre-trained language model, and obtain understanding information and / or intention information of operation behaviors corresponding to different clusters, wherein input data of the language model is prompt information and cluster information, and output data is understanding information and / or intention information of operation behaviors corresponding to different clusters; determine category label information corresponding to different clusters based on the understanding information and / or intention information of operation behaviors corresponding to different clusters.
Citation Information
Patent Citations
A method and apparatus for monitoring information risk
CN109086961A
Aircraft behavior intention recognition method based on multi-modal deep learning
CN114358211A