User Analysis Method and System Based on AI and Streaming Computing

By adopting user analysis methods based on AI and streaming computing in the financial banking field, and using the target event recognition network to process streaming data, the problem of insufficient target event recognition and positioning capabilities in traditional methods is solved, and more accurate user analysis and recognition performance is achieved.

CN115687732BActive Publication Date: 2025-06-24HANGYIN CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211516603.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-06-24
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

In the field of financial banking, traditional streaming data processing methods have weak comprehensive data display capabilities, resulting in poor identification accuracy and positioning capabilities of target events, which in turn affects the accuracy of user analysis.

Method used

Using user analysis methods based on AI and streaming computing, by receiving streaming data uploaded by terminal devices, collecting according to the data generation timing, obtaining streaming data group chains, and using the debugged target event identification network to identify and bucket the streaming data group to generate a target streaming data sequence immediately adjacent to the acquisition time.

Benefits of technology

The recognition accuracy and positioning capabilities of target event data are improved, making user analysis more accurate. At the same time, through joint debugging of the target event recognition network, the dependence on large-scale tag debugging templates is reduced, the tag pressure is alleviated and the recognition performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115687732B_ABST
    Figure CN115687732B_ABST
Patent Text Reader

Abstract

The user analysis method and system based on AI and stream computing provided by the embodiments of the present application perform group-by recognition on the stream data event set according to the target event recognition network. While identifying the target stream data group containing the target event data, the event distribution of the target event data is also determined, making the recognition process of the target event data more accurate. In addition, after obtaining the target stream data group, the target stream data group is divided according to the matching degree between the time sequence of the stream data group and the event distribution of the target event data, generating a target stream data sequence that is adjacent in acquisition time. The time sequence distribution of the target stream data sequence obtained based on this in the stream data event set to be analyzed can reflect the time sequence distribution of the target event data in the stream data event set to be analyzed, obtaining the time sequence distribution of the target event data in the stream data event set and the event distribution of the target event data, improving the accuracy of target event recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically, to a user analysis method and system based on AI and streaming computing. Background Art

[0002] In the financial banking field, the financial behaviors of a large number of users generate a large amount of data every moment within the system, including data generated by the interaction between the system and external business systems. Real-time analysis of these users' financial data, mining the characteristic information contained therein, and identifying target events can help the operation platform make decisions, such as identifying user anomalies and pushing financial products to users. Due to the fast-changing characteristics of data in the financial field, there are high requirements for the real-time processing of data, and streaming computing is usually used for processing. In traditional streaming data processing methods, due to the weak comprehensive data presentation ability, the recognition accuracy and positioning ability of target events are poor, resulting in inaccurate user analysis. Based on this, how to accurately identify target events and accurately locate them is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The purpose of the present invention is to provide a user analysis method and system based on AI and streaming computing to improve the above-mentioned problems.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] In a first aspect, an embodiment of the present application provides a user analysis method based on AI and streaming computing, which is applied to a user analysis AI system, and the method includes:

[0006] Receive the streaming data of the target user uploaded by the terminal device corresponding to the target user, and collect the streaming data according to the generation time sequence of the streaming data to obtain a streaming data event set to be analyzed, wherein the data time span in the streaming data event set to be analyzed is a preset time span;

[0007] Obtaining a streaming data group chain corresponding to the streaming data event set to be analyzed;

[0008] Performing target event identification on each streaming data group in the streaming data group chain in sequence according to the debugged target event identification network, obtaining a target streaming data group having target event data in the streaming data group chain and an event distribution of the target event data in the target streaming data group;

[0009] For the target streaming data group whose streaming data events to be analyzed are concentrated and closely collected in time, bucketing is performed according to the matching of event distribution of the target event data to obtain multiple target streaming data sequences that are closely collected in time;

[0010] Output the time series distribution of the multiple target streaming data sequences adjacent in acquisition time in the to-be-analyzed streaming data event set and the event distribution of the target event data.

[0011] Optionally, the obtaining of the streaming data group chain corresponding to the to-be-analyzed streaming data event set includes:

[0012] Obtain the to-be-analyzed streaming data event set, separate the to-be-analyzed streaming data event set according to the streaming data group capacity of the to-be-analyzed streaming data event set to obtain a plurality of streaming data clusters;

[0013] Perform data sampling according to a preset mining frequency in each of the streaming data clusters to obtain a preset number of streaming data groups;

[0014] Obtain the streaming data group chain according to the preset number of streaming data groups obtained in each streaming data cluster.

[0015] Optionally, the performing of target event recognition on each streaming data group in the streaming data group chain in sequence according to the debug-completed target event recognition network includes:

[0016] Load the multiple streaming data groups in the streaming data group chain into the debug-completed target event recognition network in sequence;

[0017] According to the feature vector mining module of the target event recognition network, mine the feature vector set corresponding to the streaming data group;

[0018] According to the event classification module of the target event recognition network, obtain the type and probability variable of each feature vector in the feature vector set according to the feature vector set of the streaming data group;

[0019] The obtaining of the target streaming data groups with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data groups includes:

[0020] Obtain the type and probability variable of each feature vector in the feature vector set output by the event classification module;

[0021] According to the probability variable of the target event data which is the data interval corresponding to each feature vector in the feature vector set and the estimated event distribution of the estimated temporary window corresponding to each feature vector, determine the target event data recognition result of the streaming data group; wherein, the target event data recognition result includes whether the streaming data group includes target event data and the event distribution of the target event data;

[0022] Based on the recognition results of the target event data in each flow data group in the flow data group chain, obtain the target flow data groups with target event data in the flow data group chain and the event distribution of the target event data in the target flow data groups.

[0023] Optionally, the method further includes:

[0024] Obtain a set of labeled debugging templates for debugging the target event recognition network;

[0025] Based on the labeling indication information of each labeled debugging template in the set of labeled debugging templates, determine the field coverage range of the target event data in the labeled debugging template;

[0026] Bucket the field coverage range of the target event data in the labeled debugging template to obtain multiple bucket centroids;

[0027] Determine the field coverage range represented by the bucket centroids as external reference variables for debugging the target event recognition network, and then perform supervised debugging on the target event recognition network through the labeled debugging templates.

[0028] Optionally, the labeled debugging templates for debugging the target event recognition network are obtained through the following steps:

[0029] Obtain multiple flow data event debugging sets;

[0030] For each flow data event debugging set, start searching from the first flow data group in the flow data event debugging set. When the searched flow data group is different from the adjacent flow data group, add the searched flow data group to the set of proposed labeled debugging templates; when the searched flow data group is similar to the adjacent flow data group, skip the searched flow data group until all the flow data groups in the flow data event debugging set are searched;

[0031] Based on the set of proposed labeled debugging templates obtained after the search of the multiple flow data event debugging sets, determine the set of labeled debugging templates for debugging the target event recognition network.

[0032] Optionally, the labeled debugging templates for debugging the target event recognition network are obtained through the following steps:

[0033] Obtain the target event data missing debugging templates in the set of labeled debugging templates that indicate no target event data;

[0034] According to the pre-set embedded event distribution, perform target event data simulation embedding on the target event data missing debugging templates to obtain simulated target event data debugging templates;

[0035] Determine the embedded event distribution as the marker indication information of the simulation target event data debugging template, and then add the simulation target event data debugging template indicating the target event data to the marker debugging template set;

[0036] Among them, the method of performing target event data simulation embedding on the target event data missing debugging template according to the pre-set embedded event distribution to obtain a simulation target event data debugging template includes: performing target event data simulation embedding on the target event data missing debugging template according to the pre-set embedded event distribution, based on one or more of the target event transaction type, target event transaction object, and target event transaction link, to obtain a simulation target event data debugging template.

[0037] Optionally, the debugging process of the target event recognition network includes:

[0038] Estimate the marker debugging templates in the marker debugging template set through the target event recognition network to obtain the estimation information of each representation vector in the representation vector set of the marker debugging templates; among them, the estimation information of the representation vector includes the estimated event distribution of the estimated temporary window, the estimated probability variable indicating whether the target event data is included in the estimated temporary window, and the estimated probability variable indicating whether the estimated temporary window is the target event data;

[0039] Obtain the first error information, second error information, and third error information of the marker debugging template based on the estimation information of the representation vectors of the representation vector set and the marker indication information of the marker debugging template;

[0040] Among them, the first error information is used to indicate the error between the event distribution of the estimated temporary window and the event distribution of the labeled temporary window;

[0041] The second error information is used to indicate the error between the estimated probability variable indicating the existence of target event data in the data interval corresponding to the representation vector and the labeled probability variable, and to indicate the error between the estimated probability variable indicating the non-existence of target event data in the data interval corresponding to the representation vector and the actual probability variable;

[0042] The third error information is used to indicate the error between the estimated probability variable and the actual probability variable indicating whether the data interval corresponding to the representation vector includes the target event data;

[0043] Optimize the network parameter variables of the target event recognition network according to the first error information, second error information, and third error information of the marker debugging templates in the marker debugging template set, so as to perform supervised debugging on the target event recognition network.

[0044] Optionally, the method further includes:

[0045] Obtain an unlabeled debugging template set, inject noise into the unlabeled debugging templates in the unlabeled debugging template set, and obtain an unlabeled template approximation group based on the unlabeled debugging templates and the debugging templates obtained by adding noise.

[0046] Determine the basic target event recognition network obtained by supervised debugging based on the labeled debugging template set as the basic target event recognition network, and respectively estimate the debugging templates included in the unlabeled template approximation group through the basic target event recognition network to obtain the respective estimation results corresponding to the debugging templates included in the unlabeled template approximation group.

[0047] Determine the common error of the unlabeled template approximation group according to the errors between the respective estimation results corresponding to the debugging templates included in the unlabeled template approximation group.

[0048] Determine the semi-supervised error according to the common error of the unlabeled template approximation group and the labeled debugging error of the labeled debugging template, and optimize the network parameters of the basic target event recognition network through the semi-supervised error to obtain the debugged target event recognition network.

[0049] Among them, the obtaining of the unlabeled debugging template set includes:

[0050] Obtain the original unlabeled debugging template set, estimate each unlabeled debugging template in the original unlabeled debugging template set according to the basic target event recognition network, and determine the semi-supervised label of the unlabeled debugging template according to the estimation result.

[0051] The semi-supervised label includes a first semi-supervised label and a second semi-supervised label.

[0052] If the number of unlabeled debugging templates with the semi-supervised label being the first semi-supervised label characterized by the estimation result is more than the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, extract the unlabeled debugging templates with the semi-supervised label being the first semi-supervised label according to the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, and obtain the unlabeled debugging template set based on the unlabeled debugging templates with the semi-supervised label being the second semi-supervised label and the extracted unlabeled debugging templates with the semi-supervised label being the first semi-supervised label.

[0053] Optionally, the determining of the semi-supervised error according to the common error of the unlabeled template approximation group and the labeled debugging error of the labeled debugging template includes:

[0054] Obtain an estimated probability variable of whether the labeled debugging template includes target event data based on the estimated result of the labeled debugging template by the basic target event recognition network;

[0055] Determine a target debugging template as a labeled debugging template whose estimated probability variable of whether it includes target event data is not greater than a preset probability variable;

[0056] Determine a semi-supervised error based on the common error of the unlabeled template approximation group and the labeled debugging error of the target debugging template;

[0057] The determining the common error of the unlabeled template approximation group according to the errors between the estimated results corresponding to the debugging templates included in the unlabeled template approximation group includes: performing a strengthening operation on the estimated results corresponding to the debugging templates included in the unlabeled template approximation group, and determining the common error of the unlabeled template approximation group based on the estimated results of the strengthening operation;

[0058] Among them, the performing a strengthening operation on the estimated results corresponding to the debugging templates included in the unlabeled template approximation group includes: when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximation group is greater than the preset probability variable, maintaining the unlabeled template approximation group for determining the common error; when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximation group is less than the preset probability variable, cleaning up the unlabeled template approximation group.

[0059] In a second aspect, an embodiment of the present application provides a user analysis AI system, including a processor and a memory, where the memory stores a computer program, and when the computer program is executed by the processor, the above method is implemented.

[0060] The user analysis method and system based on AI and streaming computing provided by the embodiments of the present application identify the streaming data event set in groups according to the target event recognition network that has been debugged. While identifying the target streaming data group containing the target event data, it also determines the event distribution of the target event data in the target streaming data group, making the identification process of the target event data more accurate. In addition, after obtaining the target streaming data group, the target streaming data group is divided according to the matching degree between the time series of the streaming data group and the event distribution of the target event data, generating a target streaming data sequence that is adjacent in acquisition time. The matching degree of the event distribution of the target event data in the same streaming data sequence is greater than a preset value. Based on this, the time series distribution of the target streaming data sequence in the streaming data event set to be analyzed can reflect the time series distribution of the target event data in the streaming data event set to be analyzed, and the event distribution of the target event data in the target streaming data sequence can reflect the event distribution of the target event data in the streaming data event set to be analyzed. In this way, the time series distribution of the target event data and the event distribution of the target event data in the streaming data event set are obtained, improving the accuracy of target event recognition.

[0061] In addition, during the debugging process of the target event recognition network, the target event recognition network is debugged in a supervised manner according to a small number of labeled debugging templates. The basic target event recognition network obtained by debugging with labeled constraints is used to estimate an unlabeled debugging template set, obtaining the common error between the unlabeled debugging template and the corresponding noisy debugging template. Based on the labeled debugging error and the common error of the labeled debugging template, the basic target event recognition network is jointly debugged to obtain the target event recognition network that has been debugged, without relying on a large number of labeled debugging templates, alleviating the labeling pressure and increasing the recognition performance of the target event recognition network.

[0062] In the following description, other features will be partly stated. When examining the following content and the drawings, those skilled in the art will partly discover these features, or may learn about these features through production or application. By practicing or using various aspects of the methods, tools, and combinations listed in the detailed examples described later, the features in the current application can be implemented and obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The drawings herein are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0064] Figure 1 It is a schematic diagram of the application scenario of the user analysis method based on AI and streaming computing provided by the embodiments of the present application.

[0065] Figure 2It is a flowchart of a user analysis method based on AI and stream computing provided by an embodiment of the present application.

[0066] Figure 3 It is a schematic diagram of the functional module architecture of a user analysis device provided by an embodiment of the present application.

[0067] Figure 4 It is a schematic diagram of the composition of a user analysis AI system provided by an embodiment of the present application. Detailed implementation manners

[0068] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0069] In the following descriptions, reference is made to "some embodiments", "as an implementation manner / scheme", "in an implementation manner". These describe subsets of all possible embodiments. However, it can be understood that "some embodiments", "as an implementation manner / scheme", "in an implementation manner" can be the same subset or different subsets of all possible embodiments, and they can be combined with each other without conflict.

[0070] In the following descriptions, the terms "first / second / third" and other similar terms are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0071] The user analysis method based on AI and stream computing provided by the embodiments of the present application can be executed by an electronic device such as a user analysis AI system. The electronic device can be various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), or can also be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0072] Next, an exemplary application will be described when the user analysis AI system is implemented as a server. The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0073] Figure 1 It is a schematic diagram of an application scenario of the user analysis method based on AI and stream computing provided by the embodiments of the present application. Communication connections are established between multiple terminal devices 100 and a user analysis AI system 300 through a network 200. The user analysis AI system 300 is used to execute the method provided by the embodiments of the present application. Specifically, the embodiments of the present application provide a user analysis method based on AI and stream computing. This method is applied to the user analysis AI system 300. As Figure 2 shown, this method includes:

[0074] Step 101: Receive the stream data of the target user uploaded by the terminal device corresponding to the target user, and collect it according to the generation time sequence of the stream data to obtain a set of stream data events to be analyzed.

[0075] In the embodiment of the present application, the data time span in the stream data event set to be analyzed is a preset time span. For example, if the preset time span is 12 hours, then when analyzing the data in the stream data event set, it is the continuous stream data within 12 hours. Since the stream data is the data collected in real time, within the preset time span, the data is also flowing, and the flow rate is, for example, the collection period of the stream data, so as to ensure the real-time nature of data analysis. In the embodiment of the present application, the specific application scenario of the stream data can be the financial field, especially the financial banking field, which generates a large amount of data all the time and flows between various business systems. By analyzing these stream data to obtain implicit characteristics, it helps to complete the analysis of the users corresponding to the data. For example, it analyzes risks such as financial fraud and securities trading fraud to help make decisions. Among them, the terminal device corresponding to the target user can be a client equipped with a financial app, such as a smart phone, a tablet computer, etc., and through the terminal device, data interaction is carried out with other financial business devices to generate the stream data of the target user.

[0076] Step 102: Obtain the stream data group chain corresponding to the stream data event set to be analyzed.

[0077] The stream data event set to be analyzed includes multiple stream data groups, and the data generation links or times corresponding to each stream data group are different. The stream data groups are arranged according to the time sequence corresponding to the stream data groups to obtain the stream data group chain. Among them, the stream data group chain can be constructed based on all the stream data groups included in the stream data event set to be analyzed, or constructed based on some of the stream data groups included in the stream data event set to be analyzed.

[0078] Specifically, as an implementation manner, obtaining a streaming data group chain may include the following steps: obtaining a set of streaming data events to be analyzed, separating the set of streaming data events to be analyzed according to the streaming data group capacity of the set of streaming data events to be analyzed to obtain a plurality of streaming data sub-clusters; performing data sampling on each streaming data sub-cluster according to a preset mining frequency to obtain a preset number of streaming data groups; obtaining a streaming data group chain based on the preset number of streaming data groups obtained in each streaming data sub-cluster. Among them, a streaming data sub-cluster is a data cluster obtained by dividing the set of streaming data events to be analyzed according to the streaming data group capacity of the set of streaming data events to be analyzed. For example, the streaming data group capacity of the set of streaming data events to be analyzed is x, and the number of data groups in the set of streaming data events to be analyzed is y. Starting from the first streaming data group, each x capacity is divided to obtain the corresponding streaming data sub-cluster. Then, according to the preset mining frequency and preset number, a single streaming data sub-cluster is extracted to obtain a preset number of streaming data groups. By extracting the streaming data sub-cluster according to the preset mining frequency and extracting according to the time equalization rule, the streaming data group chain can comprehensively reflect the set of streaming data events to be analyzed. The set of streaming data events to be analyzed is divided according to the streaming data group capacity of the set of streaming data events to be analyzed to obtain a streaming data sub-cluster, which improves the comprehensive reflection degree of the streaming data group chain.

[0079] Step 103: Sequentially perform target event recognition on each streaming data group in the streaming data group chain according to the target event recognition network that has been debugged to obtain the target streaming data groups with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data groups.

[0080] In the embodiments of the present application, the target event recognition network is established based on, for example, a deep learning network architecture. The target event recognition network learns to recognize the feature information of the target event data, such as the field coverage, transaction type, transaction object, transaction link, etc. Based on this, the target event recognition network can have a generalization recognition performance for the changes in the target event data. In order to enable the target event recognition network to learn the above feature information, the debugging template adopted in the debugging includes rich information of the target event data in different aspects such as field coverage, transaction type, transaction object, transaction link, etc. The recognition result output by the target event recognition network for the target event recognition of the streaming data group includes whether the streaming data group includes the target event data, and for the streaming data group with the target event data, the event distribution of the target event data (such as the data location of the event distribution, time distribution situation, link distribution situation). In the present application, the streaming data group with the target event data in the streaming data group chain is determined as the target streaming data group. The target event involved in the embodiments of the present application may be a financial event that needs to be concerned in the analysis, such as an event in anomaly recognition, such as program trading, or in marketing management analysis, the target event is a consumption record of a target type (such as credit card payment). The specific type of the target event in the embodiments of the present application is not limited.

[0081] After obtaining the streaming data group chain, the streaming data groups of the streaming data group chain are loaded into the target event recognition network one by one. The target event recognition network performs target event recognition on each streaming data group one by one, outputs the recognition result corresponding to each streaming data group, and recognizes the target streaming data group with the target event data and the event distribution of the target event data in the target streaming data group.

[0082] Step 104: For the target streaming data groups that are adjacent in the acquisition time in the proposed analysis streaming data event set, perform bucketing according to the matching of the event distribution of the target event data to obtain multiple target streaming data sequences that are adjacent in the acquisition time.

[0083] In the implementation of this application, whether any two target streaming data groups are target streaming data groups that are adjacent in acquisition time can be identified based on whether the time difference between the event times corresponding to the two target streaming data groups is within a preset time difference. The preset time difference can be set according to actual needs, and this application does not limit it. As an implementation method, any two target streaming data groups are obtained, and when the time difference between the event times corresponding to any two target streaming data groups is not greater than the preset time difference, it is determined that any two target streaming data groups are target streaming data groups that are adjacent in acquisition time. For example, after determining each target streaming data group according to the target event recognition network, the target streaming data group and the corresponding event time are saved, and then the target streaming data groups are sorted according to the order of the event times, and it is analyzed whether two adjacent target streaming data groups are target streaming data groups that are adjacent in acquisition time (by obtaining the event times of two adjacent target streaming data groups, and when the time difference between the event times of two adjacent target streaming data groups is not greater than the preset time difference, it is determined that these two target streaming data groups are target streaming data groups that are adjacent in acquisition time), which is convenient and fast. Among them, the matching degree of the event distribution of the target event data represents the matching degree of the target event data in the event distributions of two streaming data groups. The higher the matching degree, the greater the possibility that the target event data can be bucketed (such as mean clustering) in these two streaming data groups. After obtaining the target streaming data groups that are adjacent in acquisition time, the target streaming data groups are bucketed according to the matching degree of the event distribution of the target event data in each target streaming data group, and multiple target streaming data sequences that are adjacent in acquisition time are obtained. In the same target streaming data sequence, adjacent target streaming data groups are continuous. In addition, the matching degree of the event distribution of the target event data in multiple target streaming data groups in the target streaming data sequence is high and the event distribution similarity is high. As an implementation method, the matching degree of the event distribution of the target event data can be evaluated by a vector distance (such as the Jaccard distance, and of course it can also be the cosine distance, Euclidean distance), for example, obtaining the intersection-to-union ratio of the data ranges of the target event data of the target streaming data groups that are adjacent in acquisition time, and determining the intersection-to-union ratio as the matching degree of the event distribution of the target event data in the target streaming data groups that are adjacent in acquisition time.

[0084] For example, if two target streaming data groups that are adjacent in collection time include A and B, then based on the event distribution of the target event data in the target streaming data group A and the event distribution of the target event data in the target streaming data group B, the intersection-union ratio of the data ranges of the target event data in the target streaming data group A and the target streaming data group B is obtained, and the intersection-union ratio of the data ranges is determined as the matching degree of the event distributions of the target event data in the target streaming data group A and the target streaming data group B that are adjacent in collection time. By obtaining the corresponding matching degree based on the intersection-union ratio of the data ranges of the target event data in the target streaming data group, the bucketing characteristics of the target event data in the target streaming data group are fully reflected.

[0085] Step 105: Output the temporal distribution of each of the multiple target streaming data sequences that are adjacent in collection time in the set of streaming data events to be analyzed and the event distribution of the target event data.

[0086] Because the target event data in the set of streaming data events is changing, the existence time and the existence data interval of the target event data in the set of streaming data events may change. Based on this, the present application integrates the target event data in the dimension of the set of streaming data events to obtain multiple target streaming data sequences that are adjacent in collection time. Then, each of the target streaming data sequences that are adjacent in collection time is projected onto the time dimension to obtain the temporal distribution of each of the target streaming data sequences that are adjacent in collection time in the set of streaming data events to be analyzed. At the same time, based on the event distribution of the target event data in each target streaming data group in the target streaming data sequences that are adjacent in collection time, the event distribution of the target event data of the target streaming data sequences that are adjacent in collection time is obtained.

[0087] In the above-mentioned user analysis method based on AI and streaming computing, the set of streaming data events is sequentially identified according to the target event recognition network that has been debugged. At the same time, the target streaming data group in which the target event data exists and the event distribution of the target event data in the target streaming data group are determined. The target event data is recognized more accurately. In addition, after obtaining the target streaming data group, the target streaming data group is sorted according to the continuity of the streaming data group and the matching degree of the event distribution of the target event data, and target streaming data sequences that are adjacent in collection time are obtained. The matching degree of the event distributions of the target event data in the same streaming data sequence is greater than the preset value. Based on this, the temporal distribution of the target streaming data sequences output in the set of streaming data events to be analyzed can reflect the temporal distribution of the target event data in the set of streaming data events to be analyzed, and the event distribution of the target event data in the target streaming data sequences can reflect the event distribution of the target event data in the set of streaming data events to be analyzed, thereby improving the accuracy of target event recognition.

[0088] As an implementation, when identifying target events for a streaming data group based on a target event recognition network, multiple streaming data groups in the streaming data group chain are sequentially loaded into the target event recognition network that has been debugged; according to the feature vector mining module of the target event recognition network, a set of feature vectors corresponding to the streaming data group is mined; according to the event classification module of the target event recognition network, based on the set of feature vectors of the streaming data group, the type and probability variable of each feature vector in the set of feature vectors are obtained.

[0089] As an implementation, the target event recognition network includes a feature vector mining module and an event classification module. The feature vector mining module can be built based on mature architectures such as CNN, RNN, LSTM, etc., which can accommodate sub-units such as convolution, pooling, normalization, activation, etc. Exemplarily, the feature vector mining module can be trained based on the Google machine translation model, and the feature vector is used to represent the feature vector characteristics of the data.

[0090] The composition of the streaming data group can be a third-order tensor. After being processed by the target event recognition network, the obtained set of feature vectors can also be a third-order tensor. The estimated information corresponding to each feature vector includes the estimated event distribution of each estimated temporary window, the estimated probability variable (i.e., the possibility that there is target event data in the estimated temporary window), and the type (used to indicate whether the data in the estimated temporary window is target event data). The role of the window is to select the data corresponding to the target event, and the temporary window is the candidate window.

[0091] As an implementation manner, after obtaining the type and probability variable of each representation vector in the representation vector set output by the event classification module of the target event recognition network, according to the data interval corresponding to each representation vector, determine the probability variable of the target event data and the estimated event distribution of the estimated temporary window corresponding to each representation vector, and obtain the recognition result of the target event data of the streaming data group. The recognition result of the target event data includes whether the streaming data group includes the target event data and the event distribution of the target event data; according to the recognition result of the target event data of each streaming data group in the streaming data group chain, obtain the target streaming data group with the target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data group. Each representation vector of the streaming data group corresponds to a type and a probability variable. According to the type and the probability variable, determine which data intervals corresponding to the representation vectors are the target event data. If the possibility that the data interval corresponding to the representation vector in the streaming data group is the target event data is greater than a preset value, then the streaming data group is the target streaming data group. For the estimated event distribution of the estimated temporary window of the representation vector with the possibility greater than the preset value, determine the estimated event distribution of each estimated temporary window as the event distribution of the target event data, and obtain the recognition result of the target event of the streaming data group, that is, the streaming data group has the target event data and the event distribution of the target event data.

[0092] In the above process, because the event classification module is obtained by debugging based on a large number of templates and has learned the relevant feature information for recognizing the target event data, the confidence level of the recognition result of the target event data of each streaming data group obtained according to the representation vector set output by the event classification module is high, which is beneficial to the accuracy of the recognition result.

[0093] Then, the debugging steps of the target event recognition network will be introduced below. The debugging process of the target event recognition network in the embodiments of the present application can adopt labeled constrained debugging, unlabeled free debugging, or a combined debugging of the two. When performing combined debugging, the target event recognition network is debugged through the labeled debugging template set to obtain the basic target event recognition network, and then the basic target event recognition network is debugged through the labeled debugging template set and the unlabeled debugging template set to obtain the corresponding network, which is determined as the target event recognition network.

[0094] For labeled-constraint debugging, it is a supervised process. As an implementation method, a set of labeled debugging templates for debugging the target event recognition network is obtained; according to the labeling indication information of each labeled debugging template in the set of labeled debugging templates, the field coverage of the target event data in the labeled debugging template is determined; the field coverage of the target event data in the labeled debugging template is bucketed to obtain multiple bucket centroids (i.e., the center points of the buckets around which clustering is performed); after determining the field coverage represented by the bucket centroids as the external reference variables for debugging the target event recognition network, the target event recognition network is debugged based on the labeled debugging templates in a supervised manner.

[0095] Before labeled-constraint debugging, determine the field coverage of the estimated temporary window as the external reference variable. For example, when there is target event data in the labeled debugging template, its labeling indication information includes the field coverage of the labeled temporary window containing the target event data; obtain the field coverage corresponding to the labeled debugging template with target event data in the set of labeled debugging templates, bucket the field coverage (such as using the K-means algorithm), and after bucketing, obtain multiple bucket centroids. Determine the field coverage corresponding to the bucket centroids as the external reference variable. The external reference variable is a hyperparameter, so that during labeled-constraint debugging, the target event recognition network estimates the set of labeled debugging templates based on the external reference variable to obtain the corresponding estimated results. In the above process, before performing supervised debugging, first determine the field coverage of the estimated temporary window as the external reference variable according to the set of labeled debugging templates adopted in labeled-constraint debugging to ensure the learning performance of the target event recognition network. As an implementation method, the labeled debugging templates for debugging the target event recognition network can be obtained based on manual annotation.

[0096] The process of obtaining a labeled debugging template based on annotation may specifically include: obtaining multiple streaming data event debugging sets; for each streaming data event debugging set, starting from the first streaming data group in the streaming data event debugging set, when the searched streaming data group is different from the adjacent streaming data group, adding the searched streaming data group to the set of quasi-labeled debugging templates, and when the searched streaming data group is similar to the adjacent streaming data group, skipping the searched streaming data group until all the streaming data groups in the streaming data event debugging set are searched; determining the set of labeled debugging templates for debugging the target event recognition network based on the set of quasi-labeled debugging templates obtained after the search of multiple streaming data event debugging sets. The streaming data event debugging sets are information obtained within the scope permitted by laws and regulations. The above determination of whether two streaming data groups are different or similar can be based on determining the distance (such as the Minkowski distance) between the hash vectors of the two streaming data groups. For example, a preset distance is set, and when the hash vector distance is less than the preset distance, it is determined that the two are similar, and when it is greater than the preset distance, it is determined that the two are different. In the above process, duplicate search is performed on the streaming data groups in the streaming data event debugging set. When the searched streaming data group is different from the adjacent streaming data group, the searched streaming data group is labeled, and the similar streaming data groups are cleaned to prevent repeated labeling to improve the labeling speed.

[0097] As an implementation, in the obtained streaming data event debugging sets, most of the streaming data groups may have no target event data. Once there are too few streaming data groups with target event data, the difficulty of debugging the target event recognition network will become very large, and the obtained network has weak generalization. To overcome this problem, the embodiment of the present application improves the labeled debugging template obtained based on annotation according to the simulation policy. Then, as an implementation, obtaining the target event data missing debugging template in the set of labeled debugging templates that indicates no target event data; performing target event data simulation embedding on the target event data missing debugging template (that is, the debugging template without target event data) according to the pre-set embedded event distribution to obtain a simulated target event data debugging template; determining the embedded event distribution as the labeled indication information of the simulated target event data debugging template, and then adding the simulated target event data debugging template indicating the existence of target event data to the set of labeled debugging templates.

[0098] In the above process, by debugging the template according to the simulated target event data obtained from the simulation, the annotation cost can be reduced. The simulated scenarios can be freely determined and are more diverse. The obtained debugging template of the simulated target event data can include diverse target event data, enabling the target event recognition network to learn more feature information during debugging and improving the network generalization ability. As an implementation manner, the simulation embedding of the target event data specifically includes: based on the pre-set embedding event distribution, simulating and embedding the target event data into the debugging template with missing target event data according to one or more of the target event transaction type, the target event transaction object, and the target event transaction link to obtain the debugging template of the simulated target event data.

[0099] As an implementation manner, the labeled constraint debugging process of the target event recognition network may specifically include: estimating the labeled debugging templates in the labeled debugging template set through the target event recognition network to obtain the estimation information of each representation vector in the representation vector set of the labeled debugging templates; the estimation information of the representation vector includes: the estimated event distribution of the estimated temporary window, the estimated probability variable of whether the target event data is included in the estimated temporary window, and the estimated probability variable of whether the estimated temporary window is the target event data; according to the estimation information of the representation vectors in the representation vector set and the labeled indication information of the labeled debugging templates, obtaining the first error information, the second error information, and the third error information of the labeled debugging templates; where the first error information represents the error between the event distribution of the estimated temporary window and the event distribution of the annotated temporary window; the second error information is used to indicate the error between the estimated probability variable and the annotated probability variable of the existence of the target event data in the data interval corresponding to the representation vector, and to indicate the error between the estimated probability variable and the actual probability variable of the non-existence of the target event data in the data interval corresponding to the representation vector; the third error information is used to indicate the error between the estimated probability variable and the actual probability variable of whether the data interval corresponding to the representation vector includes the target event data; according to the first error information, the second error information, and the third error information of the labeled debugging templates in the labeled debugging template set, optimizing the network parameter variables of the target event recognition network to perform supervised debugging on the target event recognition network. According to the above annotation and simulation, the labeled debugging template set is obtained, and then the labeled debugging template set is loaded into the target event recognition network for supervised debugging. The error determination algorithm based on the labeled constraint debugging can be configured according to the actual situation, such as cross entropy, log likelihood, etc., and the present application does not limit this. In the above process, when performing supervised debugging, the network parameter variables are optimized by integrating multiple types of errors, improving the recognition ability of the target event recognition network.

[0100] When there is labeled constraint debugging, the cost of obtaining the labeled debugging template based on the annotation is high, and the number of labeled debugging templates is insufficient. Even though the simulated target event data debugging template obtained through the lack of debugging template in the target event data can expand the labeled target event data debugging template obtained based on the annotation, the quantity is limited and cannot bring the recognition ability of the target event recognition network into full play. Therefore, the embodiments of the present application provide joint debugging (i.e., semi-supervised learning method) for the target event recognition network, and jointly debug the target event recognition network through the unlabeled debugging template set and the labeled debugging template set to increase the recognition ability of the target event recognition network and make the generalization of the network stronger.

[0101] As an implementation manner, obtain the unlabeled debugging template set, perform noise injection on the unlabeled debugging templates in the unlabeled debugging template set to enhance the template data, and obtain an unlabeled template approximation group based on the unlabeled debugging templates and the debugging templates obtained by adding noise; determine the basic target event recognition network obtained by supervised debugging based on the labeled debugging template set, and respectively estimate the debugging templates included in the unlabeled template approximation group through the basic target event recognition network to obtain the respective estimation results corresponding to the debugging templates included in the unlabeled template approximation group; determine the common error of the unlabeled template approximation group according to the errors between the respective estimation results corresponding to the debugging templates included in the unlabeled template approximation group; determine the semi-supervised error according to the common error of the unlabeled template approximation group and the labeled debugging error of the labeled debugging template, and optimize the network parameter variables of the basic target event recognition network through the semi-supervised error to obtain the target event recognition network after debugging, where the unlabeled template approximation group includes the unlabeled debugging templates and the debugging templates obtained by adding noise.

[0102] As an implementation, the set of unlabeled debugging templates is obtained through type adjustment, which may specifically include: obtaining the original set of unlabeled debugging templates, estimating each unlabeled debugging template in the original set of unlabeled debugging templates according to the basic target event recognition network, and determining the semi-supervised labels (pseudo labels, Pseudo Labelling) of the unlabeled debugging templates according to the estimation results; the semi-supervised labels include the first semi-supervised label and the second semi-supervised label; if the number of unlabeled debugging templates with the semi-supervised label being the first semi-supervised label is more than the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, sampling extraction is performed on the unlabeled debugging templates with the semi-supervised label being the first semi-supervised label according to the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, and the set of unlabeled debugging templates is obtained based on the unlabeled debugging templates with the semi-supervised label being the second semi-supervised label and the extracted unlabeled debugging templates with the semi-supervised label being the first semi-supervised label. After the basic target event recognition network estimates the original set of unlabeled debugging templates, the semi-supervised labels of each unlabeled debugging template in the original set of unlabeled debugging templates are obtained. If the number of unlabeled debugging templates under the first semi-supervised label is more than the number of unlabeled debugging templates under the second semi-supervised label, the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label is used to extract the unlabeled debugging templates with the semi-supervised label being the first semi-supervised label, so that the number of the extracted unlabeled debugging templates with the first semi-supervised label is equal to the number of unlabeled debugging templates with the second semi-supervised label. Then, the set of unlabeled debugging templates is obtained based on the extracted unlabeled debugging templates with the first semi-supervised label and the unlabeled debugging templates with the second semi-supervised label. The set of unlabeled debugging templates is obtained through the above type adjustment process. The above process performs type adjustment according to the semi-supervised labels obtained from the estimation results of the basic target event recognition network, which can prevent overfitting and enhance the recognition ability of the target event recognition network. After the set of unlabeled debugging templates is generated through type adjustment, the data of the unlabeled debugging templates is strengthened according to the above noise addition process to obtain an approximate group of unlabeled templates, and then the basic target event recognition network estimates each debugging template included in the approximate group of unlabeled templates respectively (such as estimating the commonality).Commonality estimation is a process of mining information in the unlabeled debugging template during joint debugging. By adding commonality estimation to joint debugging, it is hoped that when the graph data is unstable and noisy, the target event recognition network can still accurately estimate it. Commonality estimation is aimed at a large number of unlabeled debugging templates obtained without obstacles and the debugging templates obtained by adding noise to them. Through the configured error algorithm, the target event recognition network performs commonality estimation on the unlabeled debugging template and the debugging template obtained by adding noise. In other words, the estimation results of the target event recognition network for the two need to be the same. Then, commonality estimation gives a constraint on the generalization performance of the target event recognition network, and guides the target event recognition network to extend in the direction of high generalization performance through a relatively large number of unlabeled debugging templates.

[0103] The commonality error of the unlabeled template approximation group is determined by the error between the estimation results corresponding to each debugging template included in the unlabeled template approximation group output by the basic target event recognition network. Then, the semi-supervised error is determined by comprehensively considering the labeled debugging error of the labeled debugging template. The gradient is determined based on the semi-supervised error, and the network parameters of the basic target event recognition network are optimized according to the gradient to implement joint debugging and obtain the target event recognition network after debugging. The process of obtaining the semi-supervised error can be: Y = L1 + α·L2. Where Y is the semi-supervised error, L1 is the labeled debugging error, L2 is the commonality error, and α is the parameter variable for optimizing and adjusting the weights of the labeled debugging error and the commonality error. In the above process, the commonality error is obtained through the unlabeled debugging template set, the labeled debugging error is obtained through the labeled debugging template set, and the network parameters of the target event recognition network are optimized by comprehensively considering the commonality error and the labeled debugging error for joint debugging, which enhances the recognition ability of the target event recognition network and makes the generalization of the network stronger.

[0104] As an implementation method, during joint debugging, due to the insufficient number of labeled debugging templates, overfitting is likely to occur. To overcome this problem, as an implementation method, when determining the semi-supervised error based on the commonality error of the unlabeled template approximation group and the labeled debugging error of the labeled debugging template, the estimation probability variable of whether the labeled debugging template includes target event data can be obtained according to the estimation result of the basic target event recognition network for the labeled debugging template; the labeled debugging template with the estimation probability variable of whether it includes target event data not greater than the preset probability variable is determined as the target debugging template; the semi-supervised error is determined based on the commonality error of the unlabeled template approximation group and the labeled debugging error of the target debugging template. For the labeled debugging template, too high an estimation probability variable indicates that the target event recognition network has too strong an estimation expectation for this debugging template, which is likely to cause overfitting for this debugging template. Therefore, in the embodiment of the present application, the labeled debugging template with the estimation probability variable not greater than the preset probability variable is determined as the target debugging template for error determination, and the labeled debugging template with the estimation probability variable greater than the preset probability variable is removed to prevent overfitting.

[0105] As an implementation manner, an enhancement operation can be performed on the estimated results corresponding to the debugging templates included in the unlabeled template approximation group, and the common error of the unlabeled template approximation group can be determined based on the estimated results of the enhancement operation. If the number of labeled debugging templates is insufficient, the basic target event recognition network does not learn enough about the labeled debugging templates, and the estimated distribution of the estimated results of the unlabeled debugging templates is underfitted, resulting in most of the semi-supervised errors originating from the labeled debugging templates, deviating from the joint debugging through the unlabeled debugging templates. If the estimated result distribution included in the estimated result of the unlabeled debugging template is sufficient, it is convenient for joint debugging. Then, in the embodiment of the present application, an enhancement operation is performed on the estimated results corresponding to the debugging templates included in the unlabeled template approximation group, and the common error of the unlabeled template approximation group is determined based on the estimated results of the enhancement operation to obtain the corresponding semi-supervised error. The above process performs an enhancement operation on the estimated results corresponding to the debugging templates included in the unlabeled template approximation group, preventing most of the semi-supervised errors from originating from the labeled debugging errors and facilitating joint debugging. As an implementation manner, the process of the enhancement operation may specifically include: when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximation group is greater than the preset probability variable, maintaining the unlabeled template approximation group to determine the common error; when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximation group is less than the preset probability variable, clearing the unlabeled template approximation group.

[0106] A low estimated probability variable of the unlabeled debugging template indicates that the basic target event recognition network has a poor estimation effect on the unlabeled debugging template. Then, the unlabeled template approximation group where the unlabeled debugging template is located does not determine the common error. A high estimated probability variable of the unlabeled debugging template indicates that the basic target event recognition network has a good estimation effect on the unlabeled debugging template. Then, the unlabeled template approximation group where the unlabeled debugging template is located determines the common error.

[0107] As an implementation manner, in the embodiment of the present application, the debugging process of the target event recognition network may further include:

[0108] Step 100: Based on the labeled debugging template set, perform supervised debugging on the target event recognition network to obtain a basic target event recognition network.

[0109] The labeled debugging template can be obtained through annotation and simulation. The annotation process can specifically include: obtaining multiple streaming data event debugging sets; for each streaming data event debugging set, starting from the first streaming data group in the streaming data event debugging set, if the found streaming data group is different from the adjacent streaming data group, add the found streaming data group to the set of quasi-labeled debugging templates to be marked. If the found streaming data group is similar to the adjacent streaming data group, skip the found streaming data group until all the streaming data groups in the streaming data event debugging set are searched; according to the set of quasi-labeled debugging templates obtained after the search of multiple streaming data event debugging sets, determine the set of labeled debugging templates for debugging the target event recognition network. The simulation process can specifically include: obtaining the target event data missing debugging template indicating no target event data in the set of labeled debugging templates; according to the pre-set embedded event distribution, perform target event data simulation embedding on the target event data missing debugging template to obtain the simulated target event data debugging template; after determining the embedded event distribution as the marking indication information of the simulated target event data debugging template, add the simulated target event data debugging template indicating the existence of target event data to the set of labeled debugging templates.

[0110] Step 200: Obtain the set of unlabeled debugging templates. Respectively estimate the unlabeled debugging templates and the corresponding noise-added debugging templates in the set of unlabeled debugging templates through the basic target event recognition network, obtain their respective corresponding estimation results, and determine the common error based on the error between the corresponding estimation results of the unlabeled debugging template and the corresponding noise-added debugging template.

[0111] Step 300: Jointly debug the basic target event recognition network according to the labeled debugging error of the labeled debugging template and the common error to obtain the target event recognition network with debugging completed.

[0112] The specific process and principle have been described in the foregoing embodiments and will not be repeated here. The above process first performs supervised debugging on the target event recognition network. The basic target event recognition network obtained by debugging according to the labeled constraints estimates the set of unlabeled debugging templates to obtain the common error between the unlabeled debugging template and the corresponding noise-added debugging template. According to the labeled debugging error and the common error of the labeled debugging template, the basic target event recognition network is comprehensively debugged, which not only reduces the labeling cost of the debugging template but also enhances the recognition ability of the target event recognition network.

[0113] The method provided by the embodiments of the present application can generally be divided into streaming data group mining, target event data identification of the streaming data group, and integration of the identification results. In terms of refined steps, it mainly includes division of the streaming data event set data group, establishment of the target event identification network, selection of the debugging template, generation of the simulated debugging template, generation of the joint debugging architecture, integration and analysis of the identification results, etc. Specifically, the complete process can include the following:

[0114] Obtain the streaming data event set to be analyzed, separate the streaming data event set to be analyzed according to the streaming data group capacity of the streaming data event set to be analyzed to obtain multiple streaming data clusters; perform data sampling at a preset mining frequency in each streaming data cluster to obtain a preset number of streaming data groups; obtain a streaming data group chain based on the preset number of streaming data groups obtained in each streaming data cluster; load each streaming data group in the streaming data group chain into the target event identification network that has been debugged in sequence; mine the representation vector set corresponding to the streaming data group according to the representation vector mining module of the target event identification network; according to the event classification module of the target event identification network, obtain the type and probability variable of each representation vector in the representation vector set based on the representation vector set of the streaming data group; obtain the type and probability variable of each representation vector in the representation vector set output by the event classification module; determine the target event data identification result of the streaming data group according to the probability variable of the target event data corresponding to the data interval of each representation vector in the representation vector set and the estimated event distribution of the estimated temporary window corresponding to each representation vector, and the target event data identification result includes whether the streaming data group includes target event data and the event distribution of the target event data; obtain the target streaming data groups with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data groups according to the target event data identification results of each streaming data group in the streaming data group chain; obtain any two target streaming data groups; when the time difference between the event times corresponding to any two target streaming data groups is not greater than the preset time difference, determine that any two target streaming data groups are target streaming data groups adjacent in the acquisition time; obtain the intersection ratio of the data ranges of the target event data of the target streaming data groups adjacent in the acquisition time; determine the intersection ratio of the data ranges as the matching degree of the event distribution of the target event data in the target streaming data groups adjacent in the acquisition time; for the target streaming data groups adjacent in the acquisition time in the streaming data event set to be analyzed, perform bucketing according to the matching degree of the event distribution of the target event data to obtain multiple target streaming data sequences adjacent in the acquisition time; output the time series distribution of each of the multiple target streaming data sequences adjacent in the acquisition time in the streaming data event set to be analyzed and the event distribution of the target event data.

[0115] In addition, for the debugging of the target event recognition network, the process mainly includes: obtaining a set of labeled debugging templates for debugging the target event recognition network; determining the field coverage of the target event data in the labeled debugging template according to the labeling indication information of each labeled debugging template in the set of labeled debugging templates; binning the field coverage of the target event data in the labeled debugging template to obtain multiple bin centroids; after determining the field coverage represented by the bin centroids as the external reference variables for debugging the target event recognition network, using the labeled debugging template to perform supervised debugging on the target event recognition network.

[0116] The debugging template used for joint debugging is obtained from the set of labeled debugging templates based on labeling and simulation. For labeling acquisition, it includes: obtaining a streaming data event debugging set to get multiple streaming data event debugging sets; obtaining a small-batch streaming data event debugging set from the multiple streaming data event debugging sets to get the remaining streaming data event debugging sets. The small-batch streaming data event debugging set is used to generate labeled debugging templates, and the remaining streaming data event debugging sets are used to generate unlabeled debugging templates; for each streaming data event debugging set, starting from the first streaming data group in the streaming data event debugging set for searching, if the searched streaming data group is different from the adjacent streaming data group, add the searched streaming data group to the set of quasi-labeled debugging templates. If the searched streaming data group is similar to the adjacent streaming data group, skip the searched streaming data group until all the streaming data groups in the streaming data event debugging set are searched; determine the set of labeled debugging templates for debugging the target event recognition network according to the set of quasi-labeled debugging templates obtained after the search of multiple streaming data event debugging sets; this set of labeled debugging templates includes a target event data debugging template and a target event data missing debugging template. The process of obtaining the debugging template according to simulation includes: obtaining the target event data missing debugging template in the set of labeled debugging templates that indicates the absence of target event data; according to the pre-set embedded event distribution, simulating and embedding target event data into the target event data missing debugging template based on one or more of the target event transaction type, target event transaction object, and target event transaction link to obtain a simulated target event data debugging template; after determining the embedded event distribution as the labeling indication information of the simulated target event data debugging template, add the simulated target event data debugging template indicating the presence of target event data to the set of labeled debugging templates.

[0117] The parts for marked constraint debugging mainly include: estimating the marked debugging templates in the marked debugging template set through the target event recognition network to obtain the estimation information of each representation vector in the representation vector set of the marked debugging templates; the estimation information of the representation vector includes: the estimated event distribution of the estimated temporary window, the estimated probability variable of whether the target event data is included in the estimated temporary window, and the estimated probability variable of whether the estimated temporary window is the target event data; according to the estimation information of the representation vectors in the representation vector set and the marked indication information of the marked debugging templates, obtaining the first error information, the second error information and the third error information of the marked debugging templates, where the first error information is used to indicate the error between the event distribution of the estimated temporary window and the event distribution of the labeled temporary window, the second error information is used to indicate the error between the estimated probability variable and the labeled probability variable that the data interval corresponding to the representation vector has the target event data, and the error between the estimated probability variable and the actual probability variable that the data interval corresponding to the representation vector does not have the target event data, and the third error information is used to indicate the error between the estimated probability variable and the actual probability variable of whether the data interval corresponding to the representation vector includes the target event data; optimizing the network parameter variables of the target event recognition network according to the first error information, the second error information and the third error information of the marked debugging templates in the marked debugging template set to perform supervised debugging on the target event recognition network. The network obtained through marked constraint debugging is determined as the basic target event recognition network for joint debugging to improve the recognition ability of the target event recognition network.

[0118] For the joint debugging part, it mainly includes: obtaining the original set of unlabeled debugging templates, estimating each unlabeled debugging template in the original set of unlabeled debugging templates according to the basic target event recognition network, and determining the semi-supervised labels of the unlabeled debugging templates according to the estimation results. The semi-supervised labels include the first semi-supervised label and the second semi-supervised label; if the number of unlabeled debugging templates with the semi-supervised label being the first semi-supervised label in the estimation results is more than the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, then extract the unlabeled debugging templates with the semi-supervised label being the first semi-supervised label according to the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, and obtain the set of unlabeled debugging templates according to the unlabeled debugging templates with the semi-supervised label being the second semi-supervised label and the extracted unlabeled debugging templates with the semi-supervised label being the first semi-supervised label; estimate the unlabeled debugging templates and the corresponding noisy debugging templates in the set of unlabeled debugging templates respectively through the basic target event recognition network, and obtain their respective corresponding estimation results; perform a strengthening operation on the estimation results corresponding to the debugging templates included in the unlabeled template approximation group, and determine the common error of the unlabeled template approximation group according to the estimation results of the strengthening operation; obtain the estimated probability variable of whether the labeled debugging template includes the target event data according to the estimation result of the basic target event recognition network for the labeled debugging template; determine the target debugging template as the labeled debugging template with the estimated probability variable of whether it includes the target event data not greater than the preset probability variable; determine the semi-supervised error according to the common error of the unlabeled template approximation group and the labeled debugging error of the target debugging template; perform joint debugging on the basic target event recognition network according to the labeled debugging error and the common error of the labeled debugging template, and obtain the target event recognition network with debugging completed.

[0119] In summary, the user analysis method and system based on AI and streaming computing provided by the embodiments of the present application identify the streaming data event set in groups according to the target event recognition network that has been debugged. While identifying the target streaming data group containing the target event data, the event distribution of the target event data in the target streaming data group is also determined, making the recognition process of the target event data more accurate. In addition, after obtaining the target streaming data group, the target streaming data group is divided according to the matching degree between the time sequence of the streaming data group and the event distribution of the target event data, generating a target streaming data sequence that is adjacent in acquisition time. The matching degree of the event distribution of the target event data in the same streaming data sequence is greater than a preset value. Based on this, the time sequence distribution of the target streaming data sequence in the streaming data event set to be analyzed can reflect the time sequence distribution of the target event data in the streaming data event set to be analyzed, and the event distribution of the target event data in the target streaming data sequence can reflect the event distribution of the target event data in the streaming data event set to be analyzed. In this way, the time sequence distribution of the target event data in the streaming data event set and the event distribution of the target event data are obtained, improving the accuracy of target event recognition.

[0120] In addition, during the debugging process of the target event recognition network, the target event recognition network is debugged in a supervised manner according to a small number of labeled debugging templates. The basic target event recognition network obtained according to the labeled constraint debugging is used to estimate the unlabeled debugging template set, and the common error between the unlabeled debugging template and the corresponding noisy debugging template is obtained. According to the labeled debugging error and the common error of the labeled debugging template, the basic target event recognition network is jointly debugged to obtain the debugged target event recognition network, without relying on a large number of labeled debugging templates, alleviating the labeling pressure and increasing the recognition performance of the target event recognition network.

[0121] Based on the above embodiments, the embodiments of the present application provide a user analysis device. Figure 3 It is a user analysis device 340 provided by the embodiments of the present application, as Figure 3 shown. The device 340 includes:

[0122] An event data acquisition module 341, configured to receive the streaming data of the target user uploaded by the terminal device corresponding to the target user, and collect it according to the generation time sequence of the streaming data to obtain a streaming data event set to be analyzed. Among them, the time span of the data in the streaming data event set to be analyzed is a preset time span;

[0123] A data group chain acquisition module 342, configured to acquire a streaming data group chain corresponding to the streaming data event set to be analyzed;

[0124] A target event recognition module 343, configured to sequentially perform target event recognition on each streaming data group in the streaming data group chain according to the target event recognition network that has completed debugging, so as to obtain a target streaming data group with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data group;

[0125] A data sequence determination module 344, configured to perform bucketing on the target streaming data groups that are adjacent in acquisition time in the to-be-analyzed streaming data event set according to the matching of the event distribution of the target event data, so as to obtain multiple target streaming data sequences that are adjacent in acquisition time;

[0126] An event information output module 345, configured to output the timing distribution of each of the multiple target streaming data sequences that are adjacent in acquisition time in the to-be-analyzed streaming data event set and the event distribution of the target event data.

[0127] The description of the above device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0128] If the technical solution of the present application involves personal or private information, before the product applying the technical solution of the present application processes personal information, the personal information processing rules have been clearly informed, and the individual's independent consent has been obtained. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, the individual's separate consent has been obtained, and at the same time, the requirements of "express consent" are met, and it is collected within the scope of laws and regulations. For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consent to the collection of their personal information; or on the device for processing personal information, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0129] It should be noted that in the embodiments of the present application, if the above-mentioned alarm processing method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), magnetic disks, or optical discs. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0130] The embodiments of the present application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the computer program, the above-mentioned alarm processing method is implemented.

[0131] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned alarm processing method is implemented. The computer-readable storage medium can be transient or non-transient.

[0132] The embodiments of the present application provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, part or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0133] It should be noted that Figure 4 is a schematic diagram of the hardware entity of a user analysis AI system 300 provided by the embodiments of the present application, as Figure 4As shown, the hardware entities of the user analysis AI system 300 include: a processor 310, a communication interface 320, and a memory 330, where: The processor 310 generally controls the overall operation of the user analysis AI system 300. The communication interface 320 enables the electronic device to communicate with other terminals or servers through a network. The memory 330 is configured to store instructions and applications executable by the processor 310, and can also cache data to be processed or already processed by the processor 310 and each module in the user analysis AI system 300 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 310, the communication interface 320, and the memory 330 through a bus 340. It should be noted here that: The description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0134] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0135] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0136] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings between the various components shown or discussed, or direct couplings, or communication connections can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0137] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0139] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.

[0140] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.

[0141] As described above, it is only the implementation mode of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.

Claims

1. A user analysis method based on AI and streaming computing, characterized in that Applied to a user analysis AI system, the method includes: Receiving the streaming data of the target user uploaded by the terminal device corresponding to the target user, and aggregating it according to the generation time sequence of the streaming data to obtain a set of streaming data events to be analyzed; wherein, the data time span in the set of streaming data events to be analyzed is a preset time span; Obtaining a streaming data group chain corresponding to the set of streaming data events to be analyzed; Sequentially performing target event recognition on each streaming data group in the streaming data group chain according to the target event recognition network that has been debugged, to obtain the target streaming data groups with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data groups; For the target streaming data groups that are adjacent in collection time in the set of streaming data events to be analyzed, performing bucketing according to the matching of the event distribution of the target event data, to obtain multiple target streaming data sequences that are adjacent in collection time; Outputting the time sequence distribution of the multiple target streaming data sequences that are adjacent in collection time in the set of streaming data events to be analyzed and the event distribution of the target event data.

2. The method according to claim 1, wherein The obtaining of the streaming data group chain corresponding to the set of streaming data events to be analyzed includes: Obtaining the set of streaming data events to be analyzed, and separating the set of streaming data events to be analyzed according to the streaming data group capacity of the set of streaming data events to be analyzed, to obtain multiple streaming data sub-groups; Performing data sampling according to a preset mining frequency in each streaming data sub-group to obtain a preset number of streaming data groups; Obtaining the streaming data group chain according to the preset number of streaming data groups obtained in each streaming data sub-group.

3. The method according to claim 1, wherein The sequentially performing target event recognition on each streaming data group in the streaming data group chain according to the target event recognition network that has been debugged includes: Sequentially loading the multiple streaming data groups in the streaming data group chain into the target event recognition network that has been debugged; Mining the set of feature vectors corresponding to the streaming data group according to the feature vector mining module of the target event recognition network; According to the event classification module of the target event recognition network, obtaining the type and probability variable of each feature vector in the set of feature vectors according to the set of feature vectors of the streaming data group; The obtaining of the target streaming data groups with target event data in the streaming data group chain and the event distribution of the target event data in the target streaming data groups includes: Obtaining the type and probability variable of each feature vector in the set of feature vectors output by the event classification module; Determining the recognition result of the target event data of the streaming data group according to the probability variable of the target event data corresponding to the data interval of each feature vector in the set of feature vectors and the estimated event distribution of the estimated temporary window corresponding to each feature vector; wherein, the recognition result of the target event data includes whether the streaming data group includes target event data and the event distribution of the target event data. Based on the recognition results of the target event data in each flow data group in the flow data group chain, obtain the target flow data groups with target event data in the flow data group chain and the event distribution of the target event data in the target flow data groups.

4. The method according to claim 1, characterized in that The method further includes: Obtain a set of labeled debugging templates for debugging the target event recognition network; Based on the labeling indication information of each labeled debugging template in the set of labeled debugging templates, determine the field coverage of the target event data in the labeled debugging template; Bucket the field coverage of the target event data in the labeled debugging template to obtain multiple bucket centroids; Determine the field coverage represented by the bucket centroids as external reference variables for debugging the target event recognition network, and then perform supervised debugging on the target event recognition network through the labeled debugging template.

5. The method according to claim 1, wherein The labeled debugging template for debugging the target event recognition network is obtained through the following steps: Obtain multiple flow data event debugging sets; For each of the flow data event debugging sets, start searching from the first flow data group in the flow data event debugging set. When the searched flow data group is different from the adjacent flow data group, add the searched flow data group to the set of proposed labeled debugging templates; When the searched flow data group is similar to the adjacent flow data group, skip the searched flow data group until all the flow data groups in the flow data event debugging set are searched; Based on the set of proposed labeled debugging templates obtained after the search of the multiple flow data event debugging sets is completed, determine the set of labeled debugging templates for debugging the target event recognition network.

6. The method according to claim 1, wherein The labeled debugging template for debugging the target event recognition network is obtained through the following steps: Obtain the target event data missing debugging template in the set of labeled debugging templates that indicates no target event data; Based on the pre-set embedded event distribution, perform target event data simulation embedding on the target event data missing debugging template to obtain a simulated target event data debugging template; Determine the embedded event distribution as the labeling indication information of the simulated target event data debugging template, and then add the simulated target event data debugging template indicating the presence of target event data to the set of labeled debugging templates; Among them, the step of performing target event data simulation embedding on the target event data missing debugging template based on the pre-set embedded event distribution to obtain a simulated target event data debugging template includes: based on the pre-set embedded event distribution, based on one or more of the target event transaction type, target event transaction object, and target event transaction link, perform target event data simulation embedding on the target event data missing debugging template to obtain a simulated target event data debugging template.

7. The method according to claim 1, characterized in that The debugging process of the target event recognition network includes: Estimate the labeled debugging templates in the labeled debugging template set through the target event recognition network to obtain the estimation information of each representation vector in the representation vector set of the labeled debugging templates; wherein, the estimation information of the representation vector includes the estimated event distribution of the estimated temporary window, the estimated probability variable of whether the target event data is included in the estimated temporary window, and the estimated probability variable of whether the estimated temporary window is the target event data; Obtain the first error information, the second error information, and the third error information of the labeled debugging template according to the estimation information of the representation vectors in the representation vector set and the labeling indication information of the labeled debugging template; Among them, the first error information is used to indicate the error between the event distribution of the estimated temporary window and the event distribution of the labeled temporary window; The second error information is used to indicate the error between the estimated probability variable and the labeled probability variable that the data interval corresponding to the representation vector has the target event data, and to indicate the error between the estimated probability variable and the actual probability variable that the data interval corresponding to the representation vector does not have the target event data; The third error information is used to indicate the error between the estimated probability variable and the actual probability variable of whether the data interval corresponding to the representation vector includes the target event data; Optimize the network parameter variables of the target event recognition network according to the first error information, the second error information, and the third error information of the labeled debugging templates in the labeled debugging template set, so as to perform supervised debugging on the target event recognition network.

8. The method according to claim 1, wherein The method further includes: Obtain an unlabeled debugging template set, inject noise into the unlabeled debugging templates in the unlabeled debugging template set, and obtain an unlabeled template approximation group according to the unlabeled debugging templates and the debug templates obtained by adding noise; Determine the target event recognition network obtained by performing supervised debugging based on the labeled debugging template set as the basic target event recognition network, estimate each of the debug templates included in the unlabeled template approximation group through the basic target event recognition network, and obtain the respective estimation results corresponding to the debug templates included in the unlabeled template approximation group; Determine the common error of the unlabeled template approximation group according to the error between the respective estimation results corresponding to the debug templates included in the unlabeled template approximation group; Determine the semi-supervised error according to the common error of the unlabeled template approximation group and the labeled debugging error of the labeled debugging template, and optimize the network parameter variables of the basic target event recognition network through the semi-supervised error to obtain the target event recognition network with debugging completed; Among them, the obtaining of the unlabeled debugging template set includes: Obtain the original unlabeled debugging template set, estimate each unlabeled debugging template in the original unlabeled debugging template set according to the basic target event recognition network, and determine the semi-supervised label of the unlabeled debugging template according to the estimation result; The semi-supervised label includes the first semi-supervised label and the second semi-supervised label; If the estimated result indicates that the number of unlabeled debugging templates with the semi-supervised label being the first semi-supervised label is greater than the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, then extract the unlabeled debugging templates with the semi-supervised label being the first semi-supervised label according to the number of unlabeled debugging templates with the semi-supervised label being the second semi-supervised label, and obtain an unlabeled debugging template set based on the unlabeled debugging templates with the semi-supervised label being the second semi-supervised label and the extracted unlabeled debugging templates with the semi-supervised label being the first semi-supervised label.

9. The method according to claim 8, wherein The determining of the semi-supervised error according to the common error of the unlabeled template approximate group and the labeled debugging error of the labeled debugging template includes: According to the estimated result of the labeled debugging template by the basic target event recognition network, obtain an estimated probability variable of whether the labeled debugging template includes target event data; Determine the target debugging template as the labeled debugging template whose estimated probability variable of whether it includes target event data is not greater than the preset probability variable; Determine the semi-supervised error according to the common error of the unlabeled template approximate group and the labeled debugging error of the target debugging template; The determining of the common error of the unlabeled template approximate group according to the errors between the estimated results corresponding to the debugging templates included in the unlabeled template approximate group includes: performing a strengthening operation on the estimated results corresponding to the debugging templates included in the unlabeled template approximate group, and determining the common error of the unlabeled template approximate group according to the estimated result of the strengthening operation; Among them, the performing of the strengthening operation on the estimated results corresponding to the debugging templates included in the unlabeled template approximate group includes: when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximate group is greater than the preset probability variable, maintain the unlabeled template approximate group for the determination of the common error; when the estimated probability variable in the estimated result of the debugging template included in the unlabeled template approximate group is less than the preset probability variable, clean up the unlabeled template approximate group.

10. A user analysis AI system, characterized in that, It includes a processor and a memory, and the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Prediction model training method and device, computer equipment and storage medium

    CN113344196A

  • System and method for automatic summarization of content with event based analysis

    US20210256221A1