A method and system for abnormal behavior recognition and early warning based on deep learning

By using the corpus and word vector analysis of rules and regulations in text data processing, screening and correcting abnormal behavior factors, and using the RNN neural network model, the problem of low accuracy in abnormal behavior recognition of the same word in different contexts is solved, and more efficient abnormal behavior recognition is achieved.

CN119227681BActive Publication Date: 2025-05-09SHANXI PROVINCIAL INVESTMENT GRP INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411351993.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-05-09
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

The existing method of abnormal behavior recognition based on deep learning in text data processing has reduced the accuracy of abnormal behavior recognition because the same word means opposite in different contexts.

Method used

By obtaining recorded text data in daily work, using the corpus of rules and regulations and word vectors, analyzing the abnormal behavior factors of words, filtering out suspected subject words and non-topic words, removing interfering words, using context deviation coefficients to correct the abnormal behavior factors of non-topic words, and using the RNN neural network model to identify abnormal behaviors.

Benefits of technology

It improves the accuracy of abnormal behavior recognition, reduces the impact of different contexts on the abnormality analysis of words, and improves supervision efficiency and supervision accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119227681B_ABST
    Figure CN119227681B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of text data processing, and in particular to a method and system for abnormal behavior identification and early warning based on deep learning, including: obtaining the abnormal behavior factor of each word in the recorded text data through the difference between the word vectors of each recorded text data and the same word in the corpus; determining the subject words and non-subject words according to the distribution of each type of words in the recorded text data and the similarity between the word vectors of each word in the recorded text data and the interference words in the corpus; and obtaining the context deviation coefficient of each non-subject word according to the difference between the corresponding word vectors of the non-subject word and the subject word and the difference between the abnormal behavior factors, correcting the abnormal behavior factors of all non-subject words through the context deviation coefficient, obtaining the abnormal degree of each recorded text data; and performing abnormal behavior identification. The present invention reduces the influence of context on words and improves the accuracy of abnormal behavior identification and early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text data processing, and in particular to a method and system for identifying and warning abnormal behavior based on deep learning. Background Art

[0002] In work supervision, the use of deep learning technology to identify and warn of abnormal behavior can greatly improve work efficiency and supervision accuracy. Traditional supervision methods may rely on manual analysis or rule-based systems, while deep learning models can automatically learn and identify complex behavior patterns, helping supervision agencies to detect potential violations more quickly and accurately.

[0003] In the process of identifying and detecting abnormal behavior on text data in the work monitoring process, deep learning is usually used to identify abnormal behavior on the text data based on the similarity between the text data in the work monitoring process and the text data in the rules and regulations; however, since the same word may have opposite meanings in different subject contexts, by identifying and detecting abnormal behavior on the text data in the work monitoring process and the text data in the rules and regulations, the abnormal behaviors detected are exactly opposite, resulting in a decrease in the accuracy of abnormal behavior identification on the text data in the work monitoring process. Summary of the invention

[0004] The present invention provides an abnormal behavior recognition and early warning method and system based on deep learning to solve the existing problems.

[0005] The abnormal behavior recognition and early warning method and system based on deep learning of the present invention adopt the following technical solutions:

[0006] An embodiment of the present invention provides a method for abnormal behavior recognition and early warning based on deep learning, the method comprising the following steps:

[0007] Get all recorded text data in daily work;

[0008] Obtain a corpus through rules and regulations, obtain all words in the record text data and the corpus and the corresponding word vectors, obtain the abnormal behavior factor of each word in each record text data through the difference between the word vectors of each word in each record text data and the same word in the corresponding corpus; determine the suspected subject words and suspected non-subject words in each record text data according to the distribution of each type of word in all record text data;

[0009] Obtain interference words and corresponding word vectors in the corpus, determine the interference words in each recorded text data by comparing the similarity between each word in each recorded text data and the word vectors of the interference words in the corpus; remove interference words from suspected subject words and suspected non-subject words to obtain subject words and non-subject words in each recorded text data; obtain the context deviation coefficient of each non-subject word in each recorded text data according to the difference between the corresponding word vectors of each non-subject word and all subject words and the difference between abnormal behavior factors, correct the abnormal behavior factors of all non-subject words by the context deviation coefficient of the non-subject word, and obtain the abnormal degree of each recorded text data;

[0010] Abnormal behavior is identified based on the degree of abnormality of each recorded text data.

[0011] Furthermore, the abnormal behavior factor of each word in each recorded text data is obtained by comparing the similarity and difference between the word vectors of each word in each recorded text data and the same word in the corresponding corpus, and the specific steps include the following:

[0012] The word vectors of the same words in the corpus are grouped into a set, which is recorded as the word vector set of each word;

[0013] The cosine similarity between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of proximity between each word and any word vector; the modulus length after vector subtraction between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of difference between each word and any word vector;

[0014] The degree of proximity and difference between each word and any word vector are fused to obtain the abnormal factor between each word and any word vector. The average of the abnormal factors between each word and all word vectors in the word vector set of the same word in the corpus is used as the abnormal behavior factor of each word in each record text data.

[0015] Furthermore, the method of determining suspected subject words and suspected non-subject words in each recorded text data according to the distribution and proportion of each type of words in all recorded text data includes the following specific steps:

[0016] Each category of words in each recorded text data is organized into a set of sequences according to the order of the text content, recorded as the word sequence of each category of words; the distances between adjacent words in the word sequence of each category of words are obtained, and the distances between all adjacent words in the word sequence of each category of words are organized into a set, recorded as the distance set of each category of words in each recorded text data;

[0017] Obtain the topic feature factor of each type of word in each record text data through the word frequency of each type of word in each record text data in the corresponding record text data, the inverse document frequency of each type of word in each record text data, and the standard deviation of all distances in the distance set corresponding to each type of word in each record text data;

[0018] Among them, the word frequency is positively correlated with the topic feature factor, the inverse document frequency is positively correlated with the topic feature factor, and the standard deviation is negatively correlated with the topic feature factor;

[0019] The first A-category words with the largest topic feature factor in each recorded text data are regarded as suspected topic words in the corresponding recorded text data; all words other than the topic words in the recorded text data are recorded as suspected non-topic words;

[0020] Among them, A is the preset parameter.

[0021] Furthermore, the obtaining of interference words in the corpus includes the following specific steps:

[0022] Use time, place and people in the corpus as interference words.

[0023] Furthermore, the method of determining the interference words in each recorded text data by comparing the similarity between each word in each recorded text data and the word vector of the interference words in the corpus includes the following specific steps:

[0024] Cluster all words in each record text data using the K-means clustering algorithm based on the word vectors between different words in each record text data to obtain several clusters of all words;

[0025] The mean of the cosine similarities between the word vectors of each word in each cluster in each recorded text data and the word vectors of all interference words in the corpus is recorded as the first similarity of each word in each cluster, and the mean of the first similarities of all words in each cluster is recorded as the second similarity of each cluster;

[0026] The word corresponding to the cluster with the second largest similarity is recorded as the interference word in each record text data.

[0027] Furthermore, the context deviation coefficient of each non-topic word in each recorded text data is obtained according to the difference between the word vectors corresponding to each non-topic word and all the topic words in each recorded text data, and the difference between the abnormal behavior factors, including the specific steps as follows:

[0028] Perform vector sum operation on the word vectors of all the subject words in each record text data to obtain the total word vector of the subject words in each record text data;

[0029] The modulus length after vector subtraction of the word vector of each non-topic word in each record text data from the total word vector of the topic words is recorded as the first difference between each non-topic word and all the topic words; the mean value of the difference between the abnormal behavior factors of each non-topic word and all the topic words in each record text data is recorded as the second difference between each non-topic word and all the topic words; the first difference and the second difference between each non-topic word and all the topic words are merged to obtain the context deviation of each non-topic word in each record text data; the context deviation of each non-topic word in each record text data is linearly normalized to obtain the context deviation coefficient of each non-topic word in each record text data.

[0030] Furthermore, the abnormal behavior factors of all non-topic words are corrected by the contextual deviation coefficient of the non-topic words to obtain the abnormal degree of each recorded text data, including the following specific steps:

[0031] Correct the abnormal behavior factors of all non-topic words by using the contextual deviation coefficient of the non-topic words, and obtain the corrected abnormal behavior factor of each non-topic word in each recorded text data;

[0032] The average of the normalized abnormal behavior factors corrected by all non-topic words in each record text data is taken as the abnormality degree of each record text data.

[0033] Furthermore, the abnormal behavior factors of all non-topic words are corrected by the contextual deviation coefficient of the non-topic words to obtain the corrected abnormal behavior factor of each non-topic word in each recorded text data, including the following specific steps:

[0034] The difference between the mean of the abnormal behavior factors of all the subject words in each recorded text data and the abnormal behavior factor of each non-subject word is recorded as the first difference of each non-subject word, and the product of the first difference of each non-subject word and the context deviation coefficient is recorded as the first abnormal change value of each non-subject word;

[0035] The sum of the abnormal behavior factor of each non-topic word in each recorded text data and the first abnormal change value is recorded as the corrected abnormal behavior factor of each non-topic word in each recorded text data.

[0036] Furthermore, the abnormal behavior identification is performed based on the abnormality level of each recorded text data, and the specific steps include the following:

[0037] The RNN neural network model is trained by the abnormality degree of all recorded text data to obtain the RNN neural network model after training, wherein the loss function in the RNN neural network model is a cross entropy loss function;

[0038] The recorded text data to be detected is input into the trained RNN neural network model, and 0 and 1 are output; when the output result is 1, it is determined that the detected recorded text data has abnormal behavior, and when the output result is 0, it is determined that the detected recorded text data does not have abnormal behavior.

[0039] The present invention also provides an abnormal behavior identification and early warning system based on deep learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned abnormal behavior identification and early warning methods based on deep learning.

[0040] The beneficial effects of the technical solution of the present invention are as follows: the present invention obtains the abnormal behavior factor of each word in each recorded text data through the difference between the word vectors of each word in each recorded text data and the same word in the corresponding corpus, and preliminarily analyzes the initial abnormal degree of the word in the recorded text data; determines the suspected subject words and suspected non-subject words in each recorded text data according to the distribution of each type of words in all recorded text data, and preliminarily obtains the context subject words in the recorded text data; removes interference words from suspected subject words and suspected non-subject words according to the similarity between the word vectors of each word in each recorded text data and interference words in the corpus, obtains the subject words and non-subject words in each recorded text data, and determines the context subject words in the recorded text data. The final contextual keywords in the recorded text data are reduced to reduce the influence of interference words in the recorded text data on the screening of contextual keywords; according to the difference between the corresponding word vectors of each non-topic word and all topic words in each recorded text data, and the difference between abnormal behavior factors, the contextual deviation coefficient of each non-topic word in each recorded text data is obtained, and the abnormal behavior factors of all non-topic words are corrected by the contextual deviation coefficient of the non-topic word to obtain the abnormal degree of each recorded text data, and the influence of different contexts on the abnormal degree analysis of words is reduced by the difference between topic words and non-topic words; abnormal behavior is identified by the abnormal degree of each recorded text data, which improves the accuracy of abnormal behavior identification and warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 This is a flowchart of the steps of an abnormal behavior recognition and early warning method based on deep learning of the present invention;

[0043] Figure 2 Flowchart for early warning of abnormal behavior identification. DETAILED DESCRIPTION

[0044] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a method and system for identifying and warning abnormal behaviors based on deep learning proposed by the present invention, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0045] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0046] The following is a detailed description of a specific scheme of an abnormal behavior recognition and early warning method and system based on deep learning provided by the present invention in conjunction with the accompanying drawings.

[0047] See also Figure 1 , which shows a flowchart of a method for abnormal behavior recognition and early warning based on deep learning provided by an embodiment of the present invention, the method comprising the following steps:

[0048] Step S001: Collect all recorded text data during the work monitoring process.

[0049] It should be noted that since everyone will have more or less abnormal behaviors in daily work, such as not attending meetings on time, not complying with rules and regulations, etc., in the process of detecting various behavioral anomalies, it is necessary to analyze whether there are abnormal behaviors based on various meeting activity records and work logs in daily work. Therefore, it is necessary to first collect various text data for supervision and inspection.

[0050] Specifically, work logs, meeting records, and activity records during the work process are collected and used as record text data; wherein, a daily work log is used as a record text data, a meeting record is used as a record text data, and an activity record is used as a record text data.

[0051] At this point, all recorded text data during the work monitoring process are obtained.

[0052] Step S002: Obtain a corpus through rules and regulations, acquire the recorded text data, all the words in the corpus, and their corresponding word vectors. Based on the difference in word vectors between each word in each recorded text data and the same word in the corresponding corpus, obtain the abnormal behavior factor for each word in each recorded text data. Determine the suspected topic words and suspected non-topic words in each recorded text data according to the distribution of each type of word in all recorded text data.

[0053] It should be noted that when detecting abnormal behaviors in recorded text data, usually the similarity between the recorded text data and the text data of the rules and regulations is used to identify and detect abnormal behaviors. However, since the semantics corresponding to the same word are different in different situations, it is necessary to analyze the semantics of the same word in different contexts. For example, when "reporting" someone's illegal behavior here, it has a positive meaning, while when "reporting" may be used as a means of maliciously framing or falsely accusing others without good motives and purposes, it has a negative meaning here. Therefore, it is necessary to adjust the abnormal behavior of each word according to different contexts.

[0054] Furthermore, it should be noted that in order to measure the similarity between words, it is necessary to perform preprocessing on the collected recorded text data to obtain word vectors and analyze the similarity features between words through the word vectors.

[0055] Specifically, clean all recorded text data through the text cleaning toolkit NLTK, then segment all text record data through the Jieba word segmentation tool, and then remove the stop words (such as "de", "le", "shi", etc.) in the text record data through the stop word list. Among them, the processes of using the cleaning toolkit NLTK, the Jieba word segmentation tool, and the stop word list to preprocess the text are all well-known technologies and will not be elaborated here in detail.

[0056] Obtain the word vector of each word in the recorded text data through the Word2vec algorithm; among them, the Word2vec algorithm is a well-known technology and will not be elaborated here in detail.

[0057] It should be noted that in the process of identifying and detecting abnormal behaviors, the degree of abnormality of each word in the recorded text data is analyzed by the difference in word vectors between the same words in the recorded text data and the text data of the rules and regulations. Since the words in the rules and regulations are all normal words, the abnormal behavior factor of each word in each recorded text data is obtained based on the difference in word vectors between each word in each recorded text data and the same word in the corresponding corpus.

[0058] Preferably, the text data of the rules and regulations is used as a corpus, and the text in the corpus is segmented and the word vector of each word is obtained. The word vectors of the same words in the corpus are grouped into a set, which is recorded as the word vector set of each word. Among them, the process of segmenting the text in the corpus and obtaining the word vector is the same as the above-mentioned process of recording text data, and will not be described in detail here.

[0059] Further, as an embodiment, the specific calculation method of the abnormal behavior factor of each word in each recorded text data is:

[0060] The word vectors of the same words in the corpus are grouped into a set, which is recorded as the word vector set of each word;

[0061] The cosine similarity between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of proximity between each word and any word vector; the modulus length after vector subtraction between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of difference between each word and any word vector;

[0062] The degree of proximity and difference between each word and any word vector are fused to obtain the abnormal factor between each word and any word vector. The average of the abnormal factors between each word and all word vectors in the word vector set of the same word in the corpus is used as the abnormal behavior factor of each word in each record text data.

[0063] In one embodiment of the present invention, it is specifically expressed by the formula:

[0064]

[0065] In the formula, Z c,v Represents the word vector of the vth word in the cth record text data, Z c,v,v′ Indicates that the vth word in the cth text data corresponds to the v′th word vector in the word vector set of the same word in the corpus, Y(Z c,v ,Z c,v,v′ ) represents Z c,v and Z c,v,v′ The cosine similarity between c,v -Z c,v,v′ ‖ represents Z c,v and Z c,v,v′ The modulus length after vector subtraction, n c,v represents the number of all word vectors in the word vector set of the same word in the corpus corresponding to the vth word in the cth text data, exp() represents an exponential function with a natural constant as the base, Q c,vRepresents the abnormal behavior factor of the vth word in the cth record text data.

[0066] Among them, since the words in the corpus are obtained from the rules and regulations, that is, the word vectors of all words in the corpus represent normal behavior. Therefore, when the cosine similarity between two word vectors is larger, it means that the two word vectors are more similar, which means that the possibility of abnormal occurrence of each word in the recorded text data is smaller; conversely, the more dissimilar the two word vectors are, the greater the possibility of abnormal occurrence of each word in the recorded text data. When the modulus length after vector subtraction between two word vectors is smaller, it means that the two word vectors are more similar, that is, the possibility of abnormal occurrence of each word in the recorded text data is smaller; conversely, the possibility of abnormal occurrence of each word in the recorded text data is greater.

[0067] At this point, the abnormal behavior factor of each word in each recorded text data is obtained.

[0068] It should be noted that since the same word represents different semantics in different contexts, it is necessary to obtain the subject words in each recorded text data, and to correct the abnormal behavior factor of each word in the recorded text data through the semantic differences between the subject words and non-subject words in each recorded text data, and to perform abnormal behavior identification on subsequent recorded text data through the corrected abnormal behavior factors.

[0069] It should be further explained that when the number of texts containing a word in all recorded text data is less, it means that the word is rarer, and the rarer a word is in all recorded text data, the greater the possibility that the word is a subject context word; and the rarity of each word can be reflected by the inverse document frequency of each word in the recorded text data. The more times each word appears in each recorded text data, the greater the possibility that the word is a subject context word, and the number of times each word appears can be reflected by the word frequency of each word. In text data, subject words may be distributed more widely and evenly, so as to cover more details and related content. Therefore, according to the distribution of each type of word in all recorded text data, the suspected subject words and suspected non-subject words in each recorded text data are determined.

[0070] Preferably, all the same words in each recorded text data are recorded as a class of words, so that several classes of words in each recorded text data can be obtained; each class of words in each recorded text data is formed into a sequence according to the order of the text content, which is recorded as the word sequence of each class of words; the distance between adjacent words in the word sequence of each class of words is obtained, and the distances between all adjacent words in the word sequence of each class of words are formed into a set, which is recorded as the distance set of each class of words in each recorded text data. Among them, the number of words between two words is used as the distance between adjacent words in each class of words.

[0071] Further, as an embodiment, the specific calculation method of the topic feature factor of each type of word in each record text data is:

[0072] Obtain the topic feature factor of each type of word in each record text data through the word frequency of each type of word in each record text data in the corresponding record text data, the inverse document frequency of each type of word in each record text data, and the standard deviation of all distances in the distance set corresponding to each type of word in each record text data;

[0073] Among them, the word frequency is positively correlated with the topic feature factor, the inverse document frequency is positively correlated with the topic feature factor, and the standard deviation is negatively correlated with the topic feature factor.

[0074] In one embodiment of the present invention, it is specifically expressed by the formula:

[0075]

[0076] Where TF c,k Indicates the frequency of the k-th word in the c-th record text data, IDF c,k represents the inverse document frequency of the k-th word in the c-th record text data among all the record text data, σd c,k represents the standard deviation of all distances in the distance set corresponding to the k-th word in the c-th record text data, F c,k Represents the topic feature factor of the k-th word in the c-th record text data.

[0077] Among them, when the frequency of each type of words in each record text data is greater in the corresponding record text data, it means that the possibility of this type of words as subject words is greater, that is, the subject feature factor of this type of words is greater; otherwise, it means that the possibility of this type of words as subject words is smaller, that is, the subject feature factor of this type of words is smaller. When the inverse document frequency of each type of words in each record text data is greater, it means that the possibility of this type of words as subject words is greater, that is, the subject feature factor of this type of words is greater; otherwise, it means that the possibility of this type of words as subject words is smaller, that is, the subject feature factor of this type of words is smaller. When the standard deviation of all distances in the distance set corresponding to each type of words in each record text data is smaller, it means that the distribution of this type of words in the corresponding record text data is more uniform, and the possibility of this type of words as subject words is greater, that is, the subject feature factor of this type of words is greater; otherwise, it means that the distribution of this type of words in the corresponding record text data is more uneven, and the possibility of this type of words as subject words is smaller, that is, the subject feature factor of this type of words is smaller.

[0078] At this point, the topic feature factors of each type of words in each record text data are obtained.

[0079] A parameter A is preset, wherein this embodiment is described by taking A=4 as an example, and this embodiment is not specifically limited, wherein A can be determined according to specific implementation conditions.

[0080] The first A-category words with the largest topic feature factor in each record text data are taken as suspected topic words in the corresponding record text data; all words other than the topic words in the record text data are recorded as suspected non-topic words.

[0081] At this point, the suspected subject words and suspected non-subject words in each record text data are obtained.

[0082] Step S003: Obtain interference words and corresponding word vectors in the corpus, and determine the interference words in each record text data through the similarity between each word in each record text data and the word vector of the interference words in the corpus; remove interference words from suspected subject words and suspected non-subject words to obtain subject words and non-subject words in each record text data; obtain the context deviation coefficient of each non-subject word in each record text data based on the difference between each non-subject word in each record text data and the corresponding word vectors of all subject words and the difference between abnormal behavior factors, and correct the abnormal behavior factors of all non-subject words through the context deviation coefficient of the non-subject word to obtain the abnormal degree of each record text data.

[0083] It should be noted that since the text data are composed of work logs, meeting records, activity records, etc., and there will be a large number of words such as time, place and characters in these texts, these words generally appear in the text, and these words are insignificant to the abnormal behavior of the text. However, since the recorded text data analyzes the abnormal behavior of the text through the abnormal behavior factor of each word, the abnormal behavior factors of words such as time, place and characters will cause errors in the judgment of abnormal behavior of the text. Therefore, it is necessary to eliminate the recognition of abnormal behavior of the text by the said words.

[0084] It should be further explained that in order to eliminate abnormal behavior factors of words such as time, place and person in the abnormal behavior recognition and detection of recorded text data, the words should be removed from suspected subject words and suspected non-subject words to obtain the true subject words and non-subject words.

[0085] Preferably, all words in each recorded text data are clustered by a K-means clustering algorithm according to the word vectors between different words in each recorded text data to obtain a number of clusters of all words; wherein, in this embodiment, the number of clusters is 5, but the number of clusters is not specifically limited, and the implementer can determine it according to the specific situation. wherein, the K-means clustering algorithm is a well-known technology and is not specifically limited here.

[0086] Furthermore, words such as time, place, and person are selected from the corpus and recorded as interference words in the corpus. The cosine similarity between the word vectors of the interference words in the corpus and the word vectors of each word in each cluster in each recorded text data is calculated. The average of the cosine similarities between the word vectors of each word in each cluster in each recorded text data and the word vectors of all interference words in the corpus is recorded as the first similarity of each word in each cluster, and the average of the first similarities of all words in each cluster is recorded as the second similarity of each cluster.

[0087] The words corresponding to the cluster with the second largest similarity are recorded as interference words in each recorded text data, and the suspected subject words remaining after deleting the interference words in each recorded text data are recorded as subject words, and the remaining suspected non-subject words are recorded as non-subject words.

[0088] At this point, the subject words and non-subject words in each record text data are obtained.

[0089] It should be noted that since the same word has different meanings in different contexts, and the subject word can represent the context of a text to a certain extent, the context deviation coefficient of each non-subject word in each record text data is obtained based on the difference between the corresponding word vectors of each non-subject word and all subject words in each record text data, and the difference between the abnormal behavior factors.

[0090] Preferably, a vector sum operation is performed on the word vectors of all the subject words in each recorded text data to obtain the total word vector of the subject words in each recorded text data.

[0091] Further, as an embodiment, the context deviation coefficient of each non-topic word in each record text data is specifically calculated as follows:

[0092] The modulus length after vector subtraction of the word vector of each non-topic word in each record text data from the total word vector of the topic words is recorded as the first difference between each non-topic word and all the topic words; the mean value of the difference between the abnormal behavior factors of each non-topic word and all the topic words in each record text data is recorded as the second difference between each non-topic word and all the topic words; the first difference and the second difference between each non-topic word and all the topic words are merged to obtain the context deviation of each non-topic word in each record text data; the context deviation of each non-topic word in each record text data is linearly normalized to obtain the context deviation coefficient of each non-topic word in each record text data.

[0093] In one embodiment of the present invention, it is specifically expressed by the formula:

[0094]

[0095] In the formula, Q c,i represents the abnormal behavior factor of the i-th non-topic word in the c-th record text data, Q c,j represents the abnormal behavior factor of the jth subject word in the cth record text data, || is the absolute value symbol, m1 c represents the number of all keywords in the cth record text data, Z c,i Represents the word vector of the i-th non-topic word in the c-th record text data, Z c represents the total word vector of the topic words in the c-th record text data, ‖Z c,i -Z c ‖ represents Z c,i and Z c The modulus length after vector subtraction, P c,i represents the context deviation coefficient of the i-th non-topic word in the c-th record text data, and norm() represents the linear normalization function.

[0096] Among them, the greater the difference between the abnormal behavior factors of the subject words and non-subject words in each recorded text data, the greater the contextual deviation of the non-subject words in the recorded text data, that is, the greater the contextual deviation coefficient of each non-subject word in the recorded text data; conversely, the smaller the contextual deviation coefficient of each non-subject word in the recorded text data. ‖Z1 c -Z2 c ‖ can be used to represent the difference between the word vector of each non-topic word and the total word vector of the topic word. The larger the difference, the larger the context deviation coefficient of each non-topic word; conversely, the smaller the context deviation coefficient of each non-topic word.

[0097] At this point, the context deviation coefficient of each non-topic word in each record text data is obtained.

[0098] It should be noted that when the non-topic words in the recorded text data have a large contextual deviation, there is a deviation when analyzing the abnormal behavior of the recorded text data through the abnormal behavior factor of each word in each recorded text data. Therefore, the abnormal behavior factor of the word is adjusted according to the contextual deviation to obtain whether the recorded text data has abnormal behavior. Therefore, the abnormal behavior factors of all non-topic words are corrected by the contextual deviation coefficient of the non-topic words to obtain the abnormal degree of each recorded text data.

[0099] Preferably, as an embodiment, the specific calculation method of the abnormal behavior factor after correction of each non-topic word in each recorded text data is:

[0100] The difference between the mean of the abnormal behavior factors of all the subject words in each recorded text data and the abnormal behavior factor of each non-subject word is recorded as the first difference of each non-subject word, and the product of the first difference of each non-subject word and the context deviation coefficient is recorded as the first abnormal change value of each non-subject word;

[0101] The sum of the abnormal behavior factor of each non-topic word in each recorded text data and the first abnormal change value is recorded as the corrected abnormal behavior factor of each non-topic word in each recorded text data.

[0102] In one embodiment of the present invention, it is specifically expressed by the formula:

[0103]

[0104] In the formula, Q c,i represents the abnormal behavior factor of the i-th non-topic word in the c-th record text data, represents the mean of the abnormal behavior factors of all the subject words in the cth record text data, P c,i represents the contextual deviation coefficient of the i-th non-topic word in the c-th record text data, Q′ c,i Represents the abnormal behavior factor after correction of the i-th non-topic word in the c-th record text data.

[0105] Further, as an embodiment, the specific calculation method of the abnormality degree of each recorded text data is:

[0106] The average of the normalized abnormal behavior factors corrected by all non-topic words in each record text data is taken as the abnormality degree of each record text data.

[0107] In one embodiment of the present invention, it is specifically expressed by the formula:

[0108]

[0109] In the formula, Q′ c,i represents the abnormal behavior factor after correction of the i-th non-topic word in the c-th record text data, m2 c represents the number of all non-topic words in the cth record text data, norm() represents the linear normalization function, T c Indicates the degree of abnormality of the cth record text data.

[0110] At this point, the abnormality level of each recorded text data is obtained.

[0111] Step S004: Identify abnormal behavior based on the abnormality level of each recorded text data.

[0112] It should be noted that the abnormal degree of the recorded text data after analysis is used as the data set in the deep learning process, so as to complete the abnormal behavior identification of the recorded text data through the neural network model.

[0113] Preferably, the RNN neural network model is trained by the abnormality degree of all recorded text data to obtain the trained RNN neural network model, wherein the loss function in the RNN neural network model is a cross entropy loss function.

[0114] The trained RNN neural network model is used to input the recorded text data to be detected, and outputs 0 and 1; when the output result is 1, it is determined that the detected recorded text data has abnormal behavior, and when the output result is 0, it is determined that the detected recorded text data does not have abnormal behavior. The abnormal behavior identification and warning flow chart is as follows Figure 2 shown.

[0115] At this point, abnormal behavior identification is completed.

[0116] It should be noted that the exp(-x) model used in this embodiment is only used to indicate that the negative correlation and the result of the constraint model output are in the interval (0,1). In specific implementation, it can be replaced with other models with the same purpose. This embodiment only uses the exp(-x) model as an example for description without making specific limitations on it, where x refers to the input of the model.

[0117] At this point, this embodiment is completed.

[0118] This embodiment provides an abnormal behavior recognition and early warning system based on deep learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, an abnormal behavior recognition and early warning method based on deep learning in steps S001 to S004 is implemented.

[0119] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for abnormal behavior recognition and early warning based on deep learning, characterized in that: The method comprises the following steps: Get all recorded text data in daily work; Obtain a corpus through rules and regulations, obtain all words in the record text data and the corpus and the corresponding word vectors, obtain the abnormal behavior factor of each word in each record text data through the difference between the word vectors of each word in each record text data and the same word in the corresponding corpus; determine the suspected subject words and suspected non-subject words in each record text data according to the distribution of each type of word in all record text data; Obtain interference words and corresponding word vectors in the corpus, determine the interference words in each recorded text data by comparing the similarity between each word in each recorded text data and the word vectors of the interference words in the corpus; remove interference words from suspected subject words and suspected non-subject words to obtain subject words and non-subject words in each recorded text data; obtain the context deviation coefficient of each non-subject word in each recorded text data according to the difference between the corresponding word vectors of each non-subject word and all subject words and the difference between abnormal behavior factors, correct the abnormal behavior factors of all non-subject words by the context deviation coefficient of the non-subject word, and obtain the abnormal degree of each recorded text data; Identify abnormal behavior based on the abnormality level of each recorded text data; The abnormal behavior factor of each word in each recorded text data is obtained by comparing the word vectors of each word in each recorded text data with the same word in the corresponding corpus, and the specific steps include the following: The word vectors of the same words in the corpus are grouped into a set, which is recorded as the word vector set of each word; The cosine similarity between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of proximity between each word and any word vector; the modulus length after vector subtraction between the word vector of each word in each recorded text data and any word vector in the word vector set corresponding to the same word in the corpus is recorded as the degree of difference between each word and any word vector; The degree of proximity and difference between each word and any word vector is integrated to obtain the abnormal factor between each word and any word vector. The average of the abnormal factors between each word and all word vectors in the word vector set of the same word in the corpus is used as the abnormal behavior factor of each word in each record text data. The context deviation coefficient of each non-topic word in each recorded text data is obtained according to the difference between the word vectors corresponding to each non-topic word and all the topic words in each recorded text data, and the difference between the abnormal behavior factors, including the following specific steps: Perform vector sum operation on the word vectors of all the subject words in each record text data to obtain the total word vector of the subject words in each record text data; The modulus length after vector subtraction of the word vector of each non-topic word in each record text data from the total word vector of the topic words is recorded as the first difference between each non-topic word and all the topic words; the mean value of the difference between the abnormal behavior factors of each non-topic word and all the topic words in each record text data is recorded as the second difference between each non-topic word and all the topic words; the first difference and the second difference between each non-topic word and all the topic words are merged to obtain the context deviation of each non-topic word in each record text data; the context deviation of each non-topic word in each record text data is linearly normalized to obtain the context deviation coefficient of each non-topic word in each record text data.

2. According to claim 1, a method for abnormal behavior recognition and early warning based on deep learning is characterized in that: The method of determining suspected subject words and suspected non-subject words in each recorded text data according to the distribution and proportion of each type of words in all recorded text data includes the following specific steps: Each category of words in each recorded text data is organized into a set of sequences according to the order of the text content, recorded as the word sequence of each category of words; the distances between adjacent words in the word sequence of each category of words are obtained, and the distances between all adjacent words in the word sequence of each category of words are organized into a set, recorded as the distance set of each category of words in each recorded text data; Obtain the topic feature factor of each type of word in each record text data through the word frequency of each type of word in each record text data in the corresponding record text data, the inverse document frequency of each type of word in each record text data, and the standard deviation of all distances in the distance set corresponding to each type of word in each record text data; Among them, the word frequency is positively correlated with the topic feature factor, the inverse document frequency is positively correlated with the topic feature factor, and the standard deviation is negatively correlated with the topic feature factor; The topmost topic feature factor in each record text data is Class words are regarded as suspected subject words in the corresponding recorded text data; all words other than the subject words in the recorded text data are recorded as suspected non-subject words; in, are preset parameters.

3. According to the method for abnormal behavior recognition and early warning based on deep learning as claimed in claim 1, it is characterized in that: The specific steps of obtaining interference words in the corpus are as follows: Use time, place and people in the corpus as interference words.

4. According to the method for abnormal behavior recognition and early warning based on deep learning as claimed in claim 1, it is characterized in that: The method of determining the interference words in each recorded text data by comparing the similarity between each word in each recorded text data and the word vector of the interference words in the corpus includes the following specific steps: Cluster all words in each record text data using the K-means clustering algorithm based on the word vectors between different words in each record text data to obtain several clusters of all words; The mean of the cosine similarities between the word vectors of each word in each cluster in each recorded text data and the word vectors of all interference words in the corpus is recorded as the first similarity of each word in each cluster, and the mean of the first similarities of all words in each cluster is recorded as the second similarity of each cluster; The word corresponding to the cluster with the second largest similarity is recorded as the interference word in each record text data.

5. According to the method for abnormal behavior recognition and early warning based on deep learning according to claim 1, it is characterized in that: The abnormal behavior factors of all non-topic words are corrected by the context deviation coefficient of the non-topic words to obtain the abnormal degree of each recorded text data, including the following specific steps: Correct the abnormal behavior factors of all non-topic words by using the contextual deviation coefficient of the non-topic words, and obtain the corrected abnormal behavior factor of each non-topic word in each recorded text data; The average of the normalized abnormal behavior factors corrected by all non-topic words in each record text data is taken as the abnormality degree of each record text data.

6. According to claim 5, a method for abnormal behavior recognition and early warning based on deep learning is characterized in that: The abnormal behavior factors of all non-topic words are corrected by using the contextual deviation coefficient of the non-topic words to obtain the corrected abnormal behavior factor of each non-topic word in each recorded text data, including the following specific steps: The difference between the mean of the abnormal behavior factors of all the subject words in each recorded text data and the abnormal behavior factor of each non-subject word is recorded as the first difference of each non-subject word, and the product of the first difference of each non-subject word and the context deviation coefficient is recorded as the first abnormal change value of each non-subject word; The sum of the abnormal behavior factor of each non-topic word in each recorded text data and the first abnormal change value is recorded as the corrected abnormal behavior factor of each non-topic word in each recorded text data.

7. The abnormal behavior recognition and early warning method based on deep learning according to claim 1 is characterized in that: The specific steps of identifying abnormal behavior by the abnormal degree of each recorded text data are as follows: The RNN neural network model is trained by the abnormality degree of all recorded text data to obtain the RNN neural network model after training, wherein the loss function in the RNN neural network model is a cross entropy loss function; The recorded text data to be detected is input into the trained RNN neural network model, and 0 and 1 are output; when the output result is 1, it is determined that the detected recorded text data has abnormal behavior, and when the output result is 0, it is determined that the detected recorded text data does not have abnormal behavior.

8. An abnormal behavior recognition and early warning system based on deep learning, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the abnormal behavior identification and early warning method based on deep learning as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Mass tourism web text semantic analysis method based on model fusion

    CN115099241A

  • Financial risk early warning method based on natural language processing

    CN115391498A