Abnormal behavior dynamic recognition method and system based on multi-view decoupled representation learning

Through the method based on multi-view decoupled representation learning, and using the combination of pre-trained models, the problem of dynamic characterization of abnormal behavior and redundant information processing of multi-view data in the prior art is solved, and more accurate and efficient abnormal behavior recognition is achieved.

CN119723682BActive Publication Date: 2025-05-20HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510239475.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing abnormal behavior recognition methods are difficult to effectively characterize the dynamics of abnormal behavior, and cannot accurately retain the complementary information in the multi-view data to remove redundant information.

Method used

Using a method based on multi-view decoupled representation learning, the dynamic recognition of abnormal behavior is achieved by pre-training the model including multiple specific view encoders, Transformer encoders, discrete multi-view encoders, view fusion decoders and prediction layers. This method removes redundant information and retains complementary information through the decoupling and adaptive masking mechanism of multi-view embedding vectors.

Benefits of technology

It improves the robustness and generalization ability of the model, can dynamically and accurately identify individual abnormal behaviors, provide more effective and non-redundant representations, and enhances the accuracy and efficiency of abnormal behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723682B_ABST
    Figure CN119723682B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for dynamic identification of abnormal behavior based on multi-view decoupled representation learning, and relates to the field of multi-view data modeling. The mask vector quantization method based on orthogonal constraints designed by the present invention can make the codebook vectors orthogonal to each other, improve the decoupling and generalization ability of vector features, and the adaptive mask mechanism for discrete representation constructed on this basis can effectively remove redundant information in multi-view representation and retain complementary information, improve the robustness and generalization ability of the model, and obtain more effective and non-redundant representations, so as to dynamically and accurately identify individual abnormal behaviors, and provide a more reliable feature basis for downstream tasks. In addition, the multi-period survival probability prediction method designed with time-sample monotonicity joint perception realizes cross-time monotonicity by constructing a risk accumulation function, and at the same time constructs a focal survival loss combined with a multi-period BCE loss to realize cross-sample sorting monotonicity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of multi-view data modeling and dynamic abnormal behavior recognition, and particularly relates to a method and system for dynamic abnormal behavior recognition based on multi-view decoupled representation learning. Background Art

[0002] Dynamic abnormal behavior recognition is an important branch in the fields of machine learning and computer vision, etc. It aims to detect and recognize behaviors or events that significantly deviate from normal patterns by analyzing data, and has been widely applied in fields such as intelligent security, medical monitoring, and traffic control, providing technical support for accident prevention, efficiency improvement, and safety guarantee. However, most of the existing abnormal behavior recognition methods are static, that is, predicting whether an abnormal behavior will occur at a certain future time point or period, and it is difficult to effectively depict the dynamics of abnormal behaviors, such as the time when an abnormal behavior occurs. The existing methods that can be used for dynamic abnormal behavior recognition mainly fall into two categories: (1) Traditional survival analysis methods, such methods often highly rely on assumptions and are difficult to capture complex relationships, especially the complex non-linear relationships in multi-views; (2) Dynamic deep survival methods, such methods have advantages in adaptively fitting complex relationships without assuming the survival time distribution or other factors, but it is difficult to ensure monotonicity at the time and sample levels.

[0003] In recent years, with the increasing adoption of digital, networked, and intelligent technologies in various industries, various data information has been recorded in real time in various information systems, forming multi-view data. Multi-view data can complement, cooperate with, and verify each other from multiple perspectives to describe the same object more accurately.

[0004] In this context, the demand for dynamic abnormal behavior recognition has become increasingly prominent. Multi-view data provides a more comprehensive data analysis perspective for dynamic abnormal behavior recognition, capturing the changes in behavior over time, thereby improving the accuracy and efficiency of abnormal behavior recognition. However, the existing solutions for multi-view modeling cannot effectively remove redundant information. For multi-view modeling, there are mainly two methods: (1) Direct splicing, this method directly splices the view features and embeddings obtained from different views, but ignores information redundancy; (2) Attention mechanism, this method can adaptively identify and capture the key information of multi-view data for downstream tasks, but while reducing redundant information, it will also lead to a reduction in the importance weights of complementary information, and cannot accurately retain complementary information and remove redundant information. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a method and system for dynamically identifying abnormal behaviors based on multi-view decoupled representation learning, which solves the technical problem that existing methods cannot accurately retain complementary information and remove redundant information.

[0007] (2) Technical solutions

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0009] A method for dynamically identifying abnormal behaviors based on multi-view decoupled representation learning, based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer; the method includes:

[0010] Collect and preprocess various view data related to the task, and obtain the feature vectors corresponding to each view;

[0011] Take each of the feature vectors as the input of the corresponding specific view encoder, and obtain the corresponding feature embedding vector;

[0012] Take each of the feature embedding vectors as the input of the Transformer encoder, obtain the corresponding enhanced feature embedding vector, and splice each of the enhanced feature embedding vectors to obtain a multi-view embedding vector;

[0013] Take the multi-view embedding vector as the input of the discrete multi-view encoder, and obtain the matching representation with a preset codebook;

[0014] Take the matching representation as the input of the view fusion decoder, find redundant features in the discrete latent variables through pairwise redundant probes, dynamically obtain a mask vector, so as to extract unique elements from the matching representation and set repeated elements to zero, and obtain a masked representation;

[0015] Take the masked representation as the input of the prediction layer, obtain the original outputs at multiple future moments, and generate the risk probability for predicting the occurrence of abnormal behaviors through a risk accumulation function.

[0016] Preferably, the step of taking the multi-view embedding vector as the input of the discrete multi-view encoder and obtaining the matching representation with a preset codebook includes:

[0017] Adopt the nearest neighbor search method to obtain a set of codebook indices corresponding to the latent vectors of the multi-view embedding vector and the codebook;

[0018] Take the codebook indices as discrete latent variables, so as to map each embedding of the multi-view embedding vector to the latent vector closest to it in the codebook, and obtain the matching representation.

[0019] Preferably, the masked representation is used as the input of the prediction layer to obtain the raw outputs at multiple future time instants, and a risk accumulation function is used to generate a risk probability for predicting the occurrence of abnormal behavior, which includes:

[0020] Using the masked representation as the input of the prediction layer to generate a set of raw outputs predicting the future m number of time instants , where o j represents the time instant j of the raw output;

[0021] Passing through the risk accumulation function to generate a risk probability for predicting the occurrence of abnormal behavior , and calculating the risk probability formula at any time instant u as follows:

[0022]

[0023] where exp is the exponential function.

[0024] Preferably, in the pre-training stage of the model, a codebook loss, a consistency loss, and an orthogonality loss are constructed for the discrete multi-view encoder; where:

[0025] The codebook loss is expressed as:

[0026]

[0027] where sg[·] represents the stop gradient operation, represents the square of the L2 norm, and here the enhanced feature embedding vector after the stop gradient operation is calculated and the matching representation the loss of the L2 distance between them;

[0028] The consistency loss is expressed as:

[0029]

[0030] Here, the loss of the L2 distance between the enhanced feature embedding vector and the matching representation after the stop gradient operation is calculated;

[0031] The orthogonality loss is expressed as:

[0032]

[0033] where, Denotes the L2 norm; triu(∙) denotes returning the upper triangular part of the matrix and setting other elements to zero; Denotes that the vectors in the codebook are normalized; Denotes that the vectors in the codebook are normalized and transposed.

[0034] Preferably, in the pre-training stage of the model, multi-period BCE losses are also constructed, expressed as:

[0035]

[0036] where y t denotes the true label at time t and takes 0 or 1; p t denotes the risk probability of abnormal behavior occurring at time t ; log denotes the logarithmic function.

[0037] Preferably, in the pre-training stage of the model, multi-period BCE losses are used to judge whether an individual as a sample has a risk during the observation period based on the true label;

[0038] It is defined that if the sample i has not experienced a risk event, then δ i = 0, otherwise δ i = 1; between two comparable samples i and sample j , if the risk time T i of sample i is earlier than the risk time T j of sample j , and the risk probability i of sample is higher than the risk probability j of sample , then the sorting is correct, otherwise the sorting is wrong; the function clamp(∙) denotes restricting the input to [0, ∞) and is used to exclude comparable sample pairs with correct sorting;

[0039] Construct the focal survival loss for comparable sample pairs with wrong sorting, expressed as:

[0040]

[0041] where, denotes a comparable sample pair, denotes the indicator function; denotes sample i at time A = min( T , CThe predicted probability of A = min( T , C ) represents taking the actual event occurrence time T and the observation end time C and taking the smaller value;

[0042] σ represents the Sigmoid function, and since the operation in the formula is non-differentiable, σ (x) is approximated as:

[0043]

[0044] where γ is a tuning parameter used to control the sensitivity to comparable changes, σ (x) ranges from [0, 1).

[0045] Preferably, in the pre-training stage of the model, taking the multi-period BCE loss as the main objective and the focal survival loss as the auxiliary objective, an adaptive gradient balancing method is designed; specifically including:

[0046] Let and respectively represent the gradient vectors of the multi-period BCE loss and the focal survival loss at the τ-th iteration;

[0047] Calculate and the cosine similarity between them, determine whether there is a conflict between the two, and update to , and its formula is:

[0048]

[0049] where proj u (v) represents the projection of the vector v onto the vector u here, the gradient vector is projected onto the gradient vector ; cosine represents the cosine similarity.

[0050] An abnormal behavior dynamic recognition system based on multi-view decoupled representation learning, based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer; the system includes:

[0051] A preprocessing module for collecting and preprocessing various view data related to a task and obtaining a feature vector corresponding to each view;

[0052] An embedding module for using each of the feature vectors as an input to the corresponding specific view encoder to obtain a corresponding feature embedding vector;

[0053] An enhancement module for using each of the feature embedding vectors as an input to the Transformer encoder to obtain a corresponding enhanced feature embedding vector, and concatenating each of the enhanced feature embedding vectors to obtain a multi-view embedding vector;

[0054] A matching module for using the multi-view embedding vector as an input to the discrete multi-view encoder to obtain a matching representation with a preset codebook;

[0055] A masking module for using the matching representation as an input to the view fusion decoder, finding redundant features in the discrete latent variables through pairwise redundant probes, dynamically obtaining a mask vector to extract unique elements from the matching representation and setting duplicate elements to zero to obtain a masked representation;

[0056] A prediction module for using the masked representation as an input to the prediction layer to obtain raw outputs at multiple future moments and generating a risk probability for predicting the occurrence of abnormal behavior through a risk accumulation function.

[0057] A storage medium storing a computer program for dynamically identifying abnormal behavior based on multi-view decoupled representation learning, wherein the computer program causes a computer to execute the abnormal behavior dynamic identification method as described above.

[0058] An electronic device, comprising:

[0059] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the abnormal behavior dynamic identification method as described above.

[0060] (III) Advantageous Effects

[0061] The present invention provides an abnormal behavior dynamic identification method and system based on multi-view decoupled representation learning. Compared with the prior art, the following advantageous effects are achieved:

[0062] The present invention takes the multi-view embedding vector as the input of the discrete multi-view encoder, obtains the matching representation with a preset codebook, enhances the decoupling and generalization capabilities of the vector features. On this basis, an adaptive masking mechanism for discrete representation is constructed, which can effectively remove redundant information in multi-view data and retain complementary information, improving the robustness and generalization capabilities of the model, and can obtain more effective and non-redundant representations, so as to dynamically and accurately identify the abnormal behaviors of individuals. Description of the Drawings

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a block diagram of a method for dynamically identifying abnormal behaviors based on multi-view decoupled representation learning provided by an embodiment of the present invention;

[0065] Figure 2 It is a flowchart of a method for dynamically identifying abnormal behaviors based on multi-view decoupled representation learning provided by an embodiment of the present invention. Detailed Embodiments

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0067] By providing a method and system for dynamically identifying abnormal behaviors based on multi-view decoupled representation learning in the embodiments of the present application, the technical problems that the existing methods cannot accurately retain complementary information and remove redundant information are solved, and further the technical problem of how to adapt to both time monotonicity and sample monotonicity is solved, achieving the purpose of assisting individuals and organizations to obtain non-redundant information and thus make accurate decisions.

[0068] The overall idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:

[0069] The mask vector quantization method based on orthogonal constraints designed in the embodiments of the present invention can promote the pairwise orthogonality of codebook vectors, improve the decoupling and generalization capabilities of vector features. On this basis, an adaptive mask mechanism for discrete representations is constructed, which can effectively remove redundant information in multi-view representations and retain complementary information, improving the robustness and generalization capabilities of the model, and obtaining more effective and non-redundant representations, so as to dynamically and accurately identify the abnormal behaviors of individuals, providing a more reliable feature basis for downstream tasks.

[0070] Furthermore, the multi-period survival probability prediction method with joint perception of time-sample monotonicity designed in the embodiments of the present invention realizes cross-time monotonicity by constructing a risk accumulation function, that is, the predicted survival probability decreases monotonically with time; at the same time, a focal survival loss is constructed in combination with the multi-period BCE loss to achieve ranking monotonicity across samples. The combination of the two can accurately capture the risk rankings of samples between different time points, so as to more accurately predict the occurrence time point of abnormal behaviors. However, there is a potential conflict between the two, so an adaptive gradient balancing method is designed to solve this conflict, ensuring that the model can effectively balance these two objectives during the parameter optimization process.

[0071] In addition, it can be understood that the samples in the embodiments of the present invention can be but are not limited to individuals. Abnormal behaviors reflect abnormal changes in an individual's recent behaviors. Detecting these abnormal changes is of great significance for reducing the occurrence of high-risk events. At the same time, studying an individual's abnormal behaviors also has important application values for various tasks such as commodity recommendation, community prediction, and anomaly warning. Since the occurrence of behaviors has an obvious time sequence relationship, therefore, the present invention effectively identifies whether an individual's abnormal behavior occurs and its specific occurrence time by constructing a risk accumulation function and a focal survival loss function.

[0072] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0073] Example 1:

[0074] As Figure 1 shown, the embodiments of the present invention provide an abnormal behavior dynamic recognition method based on multi-view decoupled representation learning. Based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer; the method includes:

[0075] S1. Collect and preprocess various view data related to the task, and obtain the feature vectors corresponding to each view;

[0076] S2. Use each of the feature vectors as the input to the corresponding specific view encoder to obtain the corresponding feature embedding vector;

[0077] S3. Use each of the feature embedding vectors as the input to the Transformer encoder to obtain the corresponding enhanced feature embedding vector, and splice each of the enhanced feature embedding vectors to obtain a multi-view embedding vector;

[0078] S4. Use the multi-view embedding vector as the input to the discrete multi-view encoder to obtain a matching representation with a preset codebook;

[0079] S5. Use the matching representation as the input to the view fusion decoder, find redundant features in the discrete latent variables through pairwise redundant probes, dynamically obtain a mask vector, extract unique elements from the matching representation and set duplicate elements to zero to obtain a masked representation;

[0080] S6. Use the masked representation as the input to the prediction layer to obtain the original outputs at multiple future moments, and generate a risk probability for predicting the occurrence of abnormal behavior through a risk accumulation function.

[0081] In the embodiment of the present invention, the multi-view embedding vector is used as the input to the discrete multi-view encoder to obtain a matching representation with a preset codebook, which improves the decoupling and generalization ability of the vector features. On this basis, an adaptive mask mechanism for discrete representation is constructed, which can effectively remove redundant information in multi-view data and retain complementary information, improving the robustness and generalization ability of the model, and can obtain more effective and non-redundant representations, so as to dynamically and accurately identify the abnormal behavior of individuals.

[0082] As Figure 2 shown, Figure 2 the complete process of the provided dynamic abnormal behavior recognition method based on multi-view decoupled representation learning is given. Next, each step of the above solution will be introduced in detail in combination with Figure 2 :

[0083] In step S1, collect and preprocess various view data related to the task to obtain the feature vector corresponding to each view.

[0084] In this step, collect k types of view data related to the task, and use data processing methods such as missing value filling and normalization to process the collected multi-view data set, which can improve the quality of the multi-view data set and ensure the integrity of the data. After data processing, each view obtains the corresponding feature vector. Define the feature vector , where num represents the index of the view, R represents the set of real numbers, qRepresents the number of features of the corresponding view.

[0085] Exemplarily, it is set that the task of the embodiment of the present invention is dynamic abnormal behavior recognition, and the multi-view data here can be information such as numerical values, texts, and relational networks related to individual behaviors.

[0086] In step S2, each of the feature vectors is used as the input of the corresponding specific view encoder to obtain the corresponding feature embedding vector.

[0087] In this step, based on the specific view encoder pre-constructed for each view, each view feature is mapped to the corresponding feature embedding vector, thereby enhancing the features of the vector; specifically including:

[0088] First, for the feature vector perform numerical embedding, and the process of numerical embedding is as follows:

[0089]

[0090] where v i ∈R d represents a trainable weight vector, b i represents the bias term of the feature , d represents the dimension.

[0091] Secondly, after performing numerical embedding on q feature values, obtain the feature embedding vector .

[0092] In step S3, each of the feature embedding vectors is used as the input of the Transformer encoder to obtain the corresponding enhanced feature embedding vector, and each of the enhanced feature embedding vectors is concatenated to obtain a multi-view embedding vector.

[0093] In this step, the feature embedding vector is input into the Transformer encoder to form a new enhanced feature embedding matrix , fully considering the global dependence relationship between features, concatenating each enhanced feature embedding vector, and making the concatenated multi-view feature embedding vector , that is is concatenated by obtained from each view, where f represents the total number of feature quantities of all views, .

[0094] In step S4, the multi-view embedding vector is used as the input of the discrete multi-view encoder to obtain a matching representation with a preset codebook.

[0095] In an alternative embodiment, this step includes:

[0096] S41. Using the nearest neighbor search method, obtain a set of codebook indices corresponding to the latent vectors of the multi-view embedding vectors and the codebook.

[0097] Design a codebook based on a discrete latent space for mapping continuous feature vectors into a discrete representation form. The discrete multi-view encoder maps the multi-view embedding vectors into a set of codebook indices, and each codebook index corresponds to a latent vector in the codebook. Let the codebook of the latent embedding , where l represents the number of features of the codebook, d represents the dimension, and the embedding vector is compared through and .

[0098] Use the nearest neighbor search method to find the closest discrete embedding in the codebook to replace the continuous representation, generating a set of discrete latent variables (i.e., codebook indices) .

[0099] Calculate the discrete latent variable Z The formula is as follows:

[0100]

[0101] where, represents minimizing the value in the expression, here referring to minimizing the Euclidean distance between and .

[0102] S42. Use the codebook indices as discrete latent variables to map each embedding of the multi-view embedding vectors to the closest latent vector in the codebook, obtaining a matching representation.

[0103] Based on the discrete latent variable Z , by mapping each embedding in to the closest latent vector in to generate a matching representation .

[0104] It can be understood that due to the existence of duplicate elements in the matching representation , there is redundant information within the data view and across views inside it.

[0105] In step S5, use the matching representation as the input of the view fusion decoder, find the redundant features in the discrete latent variables through pairwise redundant probes, dynamically obtain the mask vector, extract the unique elements from the matching representation and set the duplicate elements to zero, obtaining a masked representation.

[0106] In this step, a view fusion decoder based on an adaptive masking mechanism is designed to remove redundant information within and across data views, specifically including:

[0107] First, is used as the input to the decoder, and through paired redundancy probes (i.e., through paired comparison and analysis), redundant features in the discrete latent variable Z are found, and the masking vector is dynamically obtained to ensure that unique elements are extracted from and duplicate elements are set to zero; the formula for the masking vector is as follows:

[0108]

[0109] where m i is the m -th element in the masking vector i .

[0110] Secondly, based on the given masking vector, the masked representation can be generated by the Hadamard product of the matching representation and the masking vector, and the formula is as follows:

[0111]

[0112] where represents element-wise multiplication.

[0113] It should be noted that so far, through the above method in the embodiment of the present invention, the problem of how to decouple to extract complementary information while discarding redundant information has been successfully solved, and then a more efficient and non-redundant representation has been obtained. This method not only enables the model to learn a more comprehensive feature representation, but also improves the overall quality of the data. It effectively fuses complementary information from different views, thus providing a solid foundation for in-depth analysis and accurate modeling.

[0114] In step S6, the masked representation is used as the input to the prediction layer to obtain the original outputs at multiple future moments, and the risk probability for predicting the occurrence of abnormal behavior is generated through a risk accumulation function.

[0115] In this step, after obtaining the extracted representation , a risk accumulation function is constructed to map the extracted representation to the risk probabilities at different future moments, realizing the accurate identification and dynamic tracking of abnormal behaviors, and at the same time realizing cross-time monotonicity.

[0116] In an optional implementation manner, this step includes:

[0117] S61. Represent the mask as the input of the prediction layer to generate a set of original outputs m for predicting the future several moments (the output without passing through the sigmoid function), where o j represents the original output at moment j .

[0118] S62. Generate the risk probability for predicting the occurrence of abnormal behavior through the risk accumulation function , and calculate the risk probability formula at any moment u as follows:

[0119]

[0120] where exp is the exponential function.

[0121] It should be noted that in the above risk probability formula, the exponential function applied to o j is non - negative, preserving the risk ranking, the summation function gives time monotonicity, the external exponential function returns values aligned with the survival probability range, while preserving time monotonicity, and ensuring that the final output increases monotonically with time (i.e., 0 < p 1 ≤ p 2 ≤ ⋯ ≤ p m < 1).

[0122] And, the subscript j in o j corresponds to the time period (t j-1 , t j ), but since the cumulative risk probability is calculated, the subscript u in p u corresponds to the time period (0, t u ), and there is a difference between the two.

[0123] The risk accumulation function introduced in the embodiments of the present invention can accurately predict the possibility of the occurrence of abnormal behavior by accumulating the risk probabilities of each time period, thereby dynamically capturing the evolution process of abnormal behavior in time - series data. In addition, the function can adapt to various data structures and application scenarios. By adjusting the prediction layer and the risk accumulation function, it can be flexibly applied to different prediction tasks, enabling the model to more accurately capture the risk trend changing with time, thereby improving the accuracy and reliability of prediction.

[0124] In particular, in an optional embodiment, in the pre - training stage of the model:

[0125] The embodiments of the present invention construct a codebook loss, a consistency loss, and an orthogonality loss for the discrete multi - view encoder; where:

[0126] (1) Codebook Loss

[0127] Due to the non-differentiability of the argmin operation, backpropagation cannot be performed, the gradient of cannot be propagated to To solve this problem, the straight-through estimator method is adopted. By copying the gradient to and This means that the codebook itself cannot be directly trained by gradient update. Therefore, the codebook is optimized by calculating the codebook loss between

[0128]

[0129] where sg[·] represents the stop-gradient operation, represents the square of the L2 norm. Here, the enhanced feature embedding vector after the stop-gradient operation is calculated and the matching representation The loss of the L2 distance between them.

[0130] It should be noted that as Figure 2 shown, the codebook loss is mainly the loss term used to update the codebook, which minimizes the Euclidean distance between the codebook vector and the encoder output vector, thereby ensuring that the codebook vector can be as close as possible to the encoder output.

[0131] (2) Consistency Loss

[0132] Since the codebook has no dimension, to prevent the codebook from growing arbitrarily, by influencing the parameter update of the encoder to help the encoder learn to generate a latent representation closer to the codebook vector, the formula for calculating the consistency loss is as follows:

[0133]

[0134] Here, the loss of the L2 distance between the enhanced feature embedding vector and the matching representation after the stop-gradient operation is calculated.

[0135] (3) Orthogonal Loss

[0136] The orthogonal loss is used to train the codebook. By adding a vector orthogonality constraint to the codebook, it promotes the pairwise orthogonality of the codebook vectors, making each element in the codebook as independent as possible, thereby enhancing the decoupling and generalization ability of the vector features. The formula for calculating the orthogonal loss is as follows:

[0137]

[0138] Among them, represents the L2 norm; triu(∙) represents returning the upper triangular part of the matrix and setting other elements to zero; represents the vector in the codebook after being normalized; represents the vector in the codebook after being normalized and transposed.

[0139] In addition, the embodiments of the present invention also construct a multi-period BCE loss, which focuses on risk prediction at a single moment. By independently evaluating the predictions at each moment, it can accurately reflect the risk state of the sample at a specific moment and achieve efficient risk ranking. However, since the predictions at each time point are independent, the multi-period BCE loss fails to capture the dependencies between time points and thus cannot directly reflect the dynamic trend of risk changes in the time series; specifically as follows:

[0140] (4) Multi-period BCE loss

[0141] The core of the multi-period BCE loss lies in calculating the error between the predicted risk probability and the true label for each moment respectively, and its calculation formula is as follows:

[0142]

[0143] where, y t represents the true label at moment t , taking 0 or 1; p t represents the risk probability for predicting the occurrence of abnormal behavior at moment t ; log represents the logarithmic function.

[0144] It should be noted that as Figure 2 shown, the multi-period BCE loss measures the matching degree between the predicted probability p t at each moment and the true label y t by accumulating the losses at all moments, so as to achieve independent evaluation of multi-period risks.

[0145] Furthermore, in order to adapt to the monotonicity of the samples, the embodiments of the present invention construct a focal survival loss function to impose penalties on samples with ranking errors, achieve ranking monotonicity across samples, and thus improve the accuracy of the model's risk probability prediction.

[0146] Taking this sample as an example, the multi-period BCE loss is used to judge whether the sample has a risk during the observation period based on the true label; it is defined that if the sample i has not yet had a risk event, then mark δ i =0, otherwise δ i =1; for two comparable samples i and sample jBetween samples, if i the risk time T i of sample j is earlier than the risk time T j of sample i and the risk probability of sample j is higher than the risk probability of sample

[0147]

[0148] The focal survival loss is mainly used to make up for the deficiencies of multi-period BCE. This loss can accurately capture the dependencies between time points and ensure that the model output satisfies the monotonicity between samples by punishing misranked sample pairs. Its calculation formula is as follows:

[0149]

[0150] Among them, represents comparable sample pairs, represents the indicator function; represents the prediction probability of sample i at time A =min( T , C ), A =min( T , C ) represents taking the smaller value of the actual event occurrence time T and the observation end time C ;

[0151] σ σ

[0152] (x) is approximated as:

[0152]

[0153] Among them, γ is the tuning parameter used to control the sensitivity to changes in comparable pairs, σ (x) ranges from [0, 1).

[0154] It should be noted that the focal survival loss can avoid the interference of non-occurring risks and correctly ranked sample pairs on the model, making the model more focused on misranked sample pairs. By punishing misranked sample pairs, it significantly enhances the model's ability to capture changes in abnormal behaviors over time, thereby more accurately predicting the occurrence time of abnormal behaviors.

[0155] It should be noted that after constructing the multi-period BCE loss and the focal survival loss, the two are jointly trained, and there will actually be conflicts in the parameter optimization process.

[0156] To solve the above problems, as Figure 2 shown, in the pre-training stage of the model in the embodiments of the present invention, the multi-period BCE loss is used as the main target, and the focal survival loss is used as the auxiliary target, and an adaptive gradient balancing method is designed to remove the conflicting components through orthogonalization while retaining the vertical component of the auxiliary loss gradient, thereby solving the potential conflict between these two losses; specifically including:

[0157] Let and respectively represent the gradient vectors of the multi-period BCE loss and the focal survival loss at the τ-th iteration;

[0158] Calculate and the cosine similarity between them, determine whether there is a conflict between the two, and update to , and its formula is:

[0159]

[0160] where, proj u (v) represents the projection of the vector v onto the vector u here, project the gradient vector onto the gradient vector ; cosine represents the cosine similarity. When the cosine similarity between the two is less than 0, it means that the directions of the two vectors are opposite or irrelevant, and the projection operation is performed; in other cases, the directions of the two are the same or relevant, and the projection operation is not performed.

[0161] So far, the total loss of the pre-trained model in the embodiments of the present invention can be obtained:

[0162]

[0163] The model is trained by minimizing the above total loss until the model converges. Based on the converged model, the entire process of the above abnormal behavior dynamic recognition method can be executed.

[0164] Embodiment 2:

[0165] An embodiment of the present invention provides an abnormal behavior dynamic recognition system based on multi-view decoupled representation learning. Based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer. The system includes:

[0166] A preprocessing module for collecting and preprocessing multiple view data related to the task to obtain a feature vector corresponding to each view;

[0167] An embedding module for using each of the feature vectors as an input to the corresponding specific view encoder to obtain a corresponding feature embedding vector;

[0168] An enhancement module for using each of the feature embedding vectors as an input to the Transformer encoder to obtain a corresponding enhanced feature embedding vector, and concatenating each of the enhanced feature embedding vectors to obtain a multi-view embedding vector;

[0169] A matching module for using the multi-view embedding vector as an input to the discrete multi-view encoder to obtain a matching representation with a preset codebook;

[0170] A mask module for using the matching representation as an input to the view fusion decoder, finding redundant features in the discrete latent variables through pairwise redundant probes, dynamically obtaining a mask vector to extract unique elements from the matching representation and setting repeated elements to zero to obtain a masked representation;

[0171] A prediction module for using the masked representation as an input to the prediction layer to obtain raw outputs at multiple future moments, and generating a risk probability for predicting the occurrence of abnormal behavior through a risk accumulation function.

[0172] Embodiment 3:

[0173] An embodiment of the present invention provides a storage medium storing a computer program for dynamic recognition of abnormal behavior based on multi-view decoupled representation learning, wherein the computer program causes a computer to execute the abnormal behavior dynamic recognition method as described in Embodiment 1.

[0174] Embodiment 4:

[0175] An embodiment of the present invention provides an electronic device, including:

[0176] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the abnormal behavior dynamic recognition method as described in Embodiment 1.

[0177] It is understandable that the abnormal behavior dynamic recognition system, storage medium and electronic device provided by the embodiments of the present invention corresponding to the abnormal behavior dynamic recognition method based on multi-view decoupled representation learning. The explanations, examples, beneficial effects, etc. of the relevant content can refer to the corresponding parts in the abnormal behavior dynamic recognition method, and will not be elaborated here.

[0178] In summary, compared with the prior art, the following beneficial effects are achieved:

[0179] 1. The orthogonal constraint-based mask vector quantization method designed in the embodiments of the present invention can promote the pairwise orthogonality of the codebook vectors, improve the decoupling and generalization capabilities of the vector features. On this basis, the adaptive mask mechanism for discrete representations constructed can effectively remove the redundant information in the multi-view representations and retain the complementary information, improving the robustness and generalization capabilities of the model, and obtaining more effective and non-redundant representations, so as to dynamically and accurately identify the abnormal behaviors of individuals, providing a more reliable feature basis for downstream tasks.

[0180] 2. The multi-period survival probability prediction method with joint perception of time-sample monotonicity designed in the embodiments of the present invention realizes cross-time monotonicity by constructing a risk accumulation function, that is, the predicted survival probability decreases monotonically with time; at the same time, a focal survival loss is constructed in combination with the multi-period BCE loss to achieve cross-sample ranking monotonicity. The combination of the two can accurately capture the risk rankings of samples between different time points, so as to more accurately predict the occurrence time point of abnormal behaviors. However, there is a potential conflict between the two, so an adaptive gradient balancing method is designed to solve this conflict, ensuring that the model can effectively balance these two goals during the parameter optimization process.

[0181] 3. The samples in the embodiments of the present invention can be, but are not limited to, individuals. Abnormal behaviors reflect the abnormal changes in an individual's recent behaviors. Detecting these abnormal changes is of great significance for reducing the occurrence of high-risk events. At the same time, studying an individual's abnormal behaviors also has important application values for various tasks such as commodity recommendation, community prediction, and abnormal warning. Since the occurrence of behaviors has an obvious temporal relationship, the present invention effectively identifies whether an individual's abnormal behavior occurs and its specific occurrence time by constructing a risk accumulation function and a focal survival loss function.

[0182] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0183] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dynamic recognition of abnormal behavior based on multi-view decoupled representation learning, characterized in that: Based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer; the method includes: Collect and preprocess various view data related to the task and obtain the feature vector corresponding to each view; Taking each of the feature vectors as input of the corresponding specific view encoder to obtain a corresponding feature embedding vector; Using each of the feature embedding vectors as an input of the Transformer encoder to obtain a corresponding enhanced feature embedding vector, and concatenating each of the enhanced feature embedding vectors to obtain a multi-view embedding vector; Using the multi-view embedding vector as input of the discrete multi-view encoder to obtain a matching representation with a preset codebook; Taking the matching representation as the input of the view fusion decoder, finding redundant features in discrete latent variables through pairwise redundant probes, dynamically obtaining a mask vector to extract unique elements from the matching representation and set repeated elements to zero, and obtaining a mask representation; The mask representation is used as the input of the prediction layer to obtain the original output of multiple future moments, and the risk probability for predicting the occurrence of abnormal behavior is generated through a risk accumulation function.

2. The abnormal behavior dynamic identification method according to claim 1, characterized in that: The step of using the multi-view embedding vector as an input of the discrete multi-view encoder to obtain a matching representation with a preset codebook comprises: Using a nearest neighbor search method, obtaining a set of codebook indexes corresponding to the multi-view embedding vector and the potential vector of the codebook; The codebook index is used as a discrete latent variable to map each embedding of the multi-view embedding vector to the closest latent vector in the codebook to obtain a matching representation.

3. The abnormal behavior dynamic identification method according to claim 2, characterized in that: The method uses the mask representation as the input of the prediction layer, obtains the original output at multiple moments in the future, and generates a risk probability for predicting the occurrence of abnormal behavior through a risk accumulation function; including: The mask represents As input to the prediction layer, a set of predicted future m The original output at that moment , where o j Indicates time j The original output of The risk accumulation function is used to generate the risk probability for predicting the occurrence of abnormal behavior. , calculated at any time u The risk probability formula is as follows: Where exp is an exponential function.

4. The abnormal behavior dynamic identification method according to claim 3, characterized in that: In the pre-training stage of the model, a codebook loss, a consistency loss and an orthogonal loss are constructed for the discrete multi-view encoder; wherein: The codebook loss is expressed as: Among them, sg[·] means stopping the gradient operation, Represents the square of the L2 norm, where the enhanced feature embedding vector is calculated after the gradient stop operation Matching expression The loss of L2 distance between them; The consistency loss is expressed as: The enhanced feature embedding vector is calculated here The matching representation after the gradient stop operation The loss of L2 distance between them; The orthogonal loss is expressed as: in, represents the L2 norm; triu(∙) means returning the upper triangular part of the matrix and setting other elements to zero; Indicates that the codebook The vectors in are normalized; Indicates that the codebook The vectors in are normalized and transposed.

5. The abnormal behavior dynamic identification method according to claim 4, characterized in that: In the pre-training stage of the model, a multi-period BCE loss is also constructed, expressed as: Among them, y t Indicates time t The true label is 0 or 1; p t Indicates the time used to predict t The risk probability of abnormal behavior occurring; log represents the logarithmic function.

6. The abnormal behavior dynamic identification method according to claim 5, characterized in that: In the pre-training stage of the model, a multi-period BCE loss is used to determine whether the individual as a sample has a risk during the observation period based on the true label; If the sample is defined i If no risk event has occurred, mark δ i =0, otherwise δ i =1; in two comparable samples i and samples j If the sample i Risk time T i Earlier than the sample j Risk time T j , and the sample i The risk probability Higher than sample j The risk probability , then the sorting is correct, otherwise the sorting is wrong; the function clamp(∙) means limiting the input to [0,∞), which is used to exclude comparable sample pairs with correct sorting; Construct the focal survival loss for pairs of comparable samples with ordering errors, expressed as: Among them, the use of represents comparable sample pairs, represents the indicator function; Representation sample i In time A =min( T , C ), A =min( T , C ) indicates the actual time of event occurrence T and observation end time C The smaller value of ; σ Represents the Sigmoid function, and since in the formula The operation is not differentiable. σ (x) is approximately: in, γ is a tuning parameter used to control the sensitivity of comparable changes. σ The range of (x) is [0, 1).

7. The abnormal behavior dynamic identification method according to claim 6, characterized in that: In the pre-training stage of the model, the multi-period BCE loss is used as the main target, the focal survival loss is used as the auxiliary target, and an adaptive gradient balancing method is designed; specifically, it includes: make and They represent the gradient vectors of multi-period BCE loss and focal survival loss at the τth iteration respectively; calculate and The cosine similarity between them is used to determine whether there is a conflict between the two, and the Gram-Schmidt orthogonalization is used to transform Updated to , the formula is: Among them, proj u (v) represents a vector v In vector u The projection on Projection to the gradient vector Above; cosine represents cosine similarity.

8. An abnormal behavior dynamic recognition system based on multi-view decoupled representation learning, characterized in that: Based on a pre-trained model, the model includes multiple specific view encoders, a Transformer encoder, a discrete multi-view encoder, a view fusion decoder based on an adaptive mask mechanism, and a prediction layer; the system includes: The preprocessing module is used to collect and preprocess various view data related to the task and obtain the feature vector corresponding to each view; An embedding module, configured to use each of the feature vectors as an input of the corresponding specific view encoder to obtain a corresponding feature embedding vector; An enhancement module, used to take each of the feature embedding vectors as an input of the Transformer encoder, obtain a corresponding enhanced feature embedding vector, and concatenate each of the enhanced feature embedding vectors to obtain a multi-view embedding vector; A matching module, configured to use the multi-view embedding vector as an input of the discrete multi-view encoder to obtain a matching representation with a preset codebook; A mask module, configured to use the matching representation as an input of the view fusion decoder, find redundant features in discrete latent variables through pairwise redundant probes, dynamically obtain a mask vector to extract unique elements from the matching representation and set repeated elements to zero, and obtain a mask representation; The prediction module is used to use the mask representation as the input of the prediction layer, obtain the original output of multiple future moments, and generate the risk probability for predicting the occurrence of abnormal behavior through the risk accumulation function.

9. A storage medium, characterized in that: It stores a computer program for dynamic identification of abnormal behavior based on multi-view decoupled representation learning, wherein the computer program enables a computer to execute the method for dynamic identification of abnormal behavior as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a program for executing the abnormal behavior dynamic identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal behavior recognition method and device for multi-view instance-semantic consensus mining

    CN117690192A

  • Cross-lingual unsupervised classification with multi-view transfer learning

    US20210390270A1