Chinese latent language recognition method and system based on dynamic word embedding

Through a method based on dynamic word embedding, the RoBERTa-wwm-ext model is used to generate context-sensitive word vectors, combined with new word discovery technology and vector spatial mapping, the problem of limited scope and poor effect of cryptographic recognition in the existing technology is solved, and more accurate and adaptive cryptographic recognition is achieved.

CN119990118APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510055566.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has little research on data diversity, emerging cryptographic language discovery and context dependence, resulting in limited scope of cryptographic language recognition, difficulty in recognizing emerging cryptographic languages, and poor recognition effects.

Method used

Using Chinese cryptographic recognition method based on dynamic word embedding, using information platform data and new word discovery technology, context-sensitive word vectors are generated through the large-scale pre-trained deep learning model RoBERTa-wwm-ext in Chinese, combined with vector spatial mapping and similarity analysis, accurate recognition of Chinese cryptographic words is achieved.

Benefits of technology

It significantly improves the accuracy and adaptability of cryptographic language recognition, can effectively identify emerging cryptographic languages, improves the understanding of complex semantics, and enhances the practicality of cryptographic language recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990118A_ABST
    Figure CN119990118A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese latent language recognition method based on dynamic word embedding, and belongs to the technical fields of natural language processing, deep learning, data mining, cross-context text analysis, network security and the like. The method comprises the steps of obtaining to-be-detected text data; through a new word discovery technology and a boundary enhancement technology, obtaining new words from to-be-detected text data so as to construct a new word dictionary; according to the new word dictionary and the basic dictionary, performing word segmentation processing on the to-be-detected text data to obtain a word segmentation result; the word segmentation result comprises new words and known words; comparing the known word with the official corpus, and identifying whether the known word is a latent language or not; and comparing the new word with a predefined category, and identifying whether the new word is a latent language or not. By integrating platform data, the data diversity is remarkably improved, and the adaptability of the model in different environments is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to technical fields such as natural language processing, deep learning, data mining, cross-context text analysis and network security, and in particular to a Chinese argot recognition method and system based on dynamic word embedding. Background Art

[0002] The reason why code words are abused is that they have the following characteristics:

[0003] 1. Concealment: The main feature of code language is that its meaning is not easy to be understood by outsiders. It distorts or replaces the surface meaning of common words so that only specific groups or insiders can understand it accurately, thus avoiding monitoring and identification.

[0004] 2. Dynamicity: Code words are constantly changing and evolving. With the development of society and technology, new code words continue to emerge, and old code words may be modified or abandoned. Therefore, the code word recognition system must be highly adaptable to keep up with the pace of its development.

[0005] 3. Context dependence: The meaning of lingo is usually closely related to a specific context. The same word or expression may represent different meanings in different situations. Therefore, the recognition of lingo requires not only considering the word itself, but also combining it with contextual information to accurately understand its true meaning.

[0006] The main research methods for cryptanalysis detection and recognition can be summarized as follows:

[0007] 1. Dictionary / rule-based recognition method: This method relies on a pre-defined lingo vocabulary or rules (such as regular expressions) to identify lingo, and is used to identify common lingo and lingo with specific grammatical structures or lexical patterns. This method is the most direct and simple lingo recognition method.

[0008] 2. Identification method based on statistical model: Statistical model is based on the statistical characteristics of a large amount of data, and identifies codewords by calculating the frequency of words in the text, collocation patterns, etc. This type of method includes traditional NLP techniques, such as n-gram model, maximum entropy model, hidden Markov model (HMM), etc.

[0009] 3. Word vector-based recognition method: Use static word embedding technology (such as Word2Vec, GloVe, FastText, etc.) to map words to a vector space, and identify cryptic language by calculating the similarity between word vectors. However, when encoding data in this way, only fixed expressions can be generated for each word, and different expressions cannot be flexibly generated according to the context.

[0010] Currently, there is little research on the existing technologies in terms of data diversity, emerging codeword discovery and context dependency. There are problems such as limited codeword recognition range, difficulty in identifying emerging codewords and poor recognition effect. Summary of the invention

[0011] The purpose of the present invention is to provide a Chinese lingo recognition method and system based on dynamic word embedding, aiming to use information platform data to identify potential emerging lingo through new word discovery and boundary enhancement technology, and use RoBERTa-wwm-ext, a deep learning model based on large-scale pre-training in Chinese, to generate word vector expressions based on Chinese vocabulary, and complete the Chinese lingo recognition task through vector space mapping, similarity analysis and other steps. On the basis of existing lingo recognition technology, the present invention integrates multi-domain data and enhances the recognition ability of emerging lingo; using dynamic word vector generation technology, different word vectors are generated according to different contexts, and the meaning of words in the context is understood more accurately, which significantly improves the accuracy and practicality of lingo recognition.

[0012] In order to solve the above technical problems, the present invention provides a Chinese argot recognition method based on dynamic word embedding, comprising the following steps:

[0013] Get the text data to be detected;

[0014] Through new word discovery technology and boundary enhancement technology, new word vocabulary is obtained from the text data to be detected, so as to build a new word dictionary;

[0015] According to the new word dictionary and the basic dictionary, the text data to be detected is segmented to obtain a segmentation result; the segmentation result includes new words and known words;

[0016] Compare the known words with the formal corpus to identify whether the known words are code words;

[0017] Compare the new words with predefined categories to identify whether the new words are cryptic.

[0018] Preferably, the new word discovery technology adopts an information entropy-based method to measure the context diversity of candidate words by calculating the left and right information entropies of the candidate words;

[0019] The boundary enhancement technology accurately identifies word boundaries by capturing local information of candidate words.

[0020] Preferably, the calculation formula of the information entropy is:

[0021]

[0022] Where: H(X) represents information entropy, which indicates the uncertainty of random variable X; P(x i) represents the probability of event xi occurring; n represents the total number of possible events.

[0023] Preferably, based on the formal corpus, the known word is compared with the formal corpus to identify whether the known word is a codeword, which specifically includes the following steps:

[0024] Vectorize the known words in the formal corpus to obtain the formal corpus word vectors;

[0025] Perform vectorization processing on known words in the text data to be detected to obtain known word vectors;

[0026] After the formal corpus word vectors and known word vectors are mapped into vector space, cosine similarity comparison is performed;

[0027] If the cosine similarity is greater than or equal to a preset threshold value t1, the corresponding known word is determined to be a non-cryptographic word;

[0028] If the cosine similarity is lower than the first preset threshold t1, the corresponding known word is determined to be a codeword.

[0029] Preferably, based on the formal corpus, the following steps are also included:

[0030] The known word vectors of the known words determined to be secret words are respectively calculated with the word vectors of several predefined categories for cosine similarity;

[0031] Add the known words to the predefined category with the highest cosine similarity and update the word vector of the predefined category.

[0032] Preferably, based on the formal corpus, the new word is compared with the predefined category to identify whether the new word is a codeword, which specifically includes the following steps:

[0033] Vectorize the new words in the text data to be detected to obtain new word vectors;

[0034] Calculate the cosine similarity between the new word vector and the word vectors of several predefined categories;

[0035] If the cosine similarity between the new word vector and the word vectors of all predefined categories is lower than the second preset threshold t2, the corresponding new word is determined to be a non-cryptographic word;

[0036] If the cosine similarity between the new word vector and the word vectors of all predefined categories is greater than or equal to the second preset threshold t2, the corresponding new word is determined to be a codeword, and the corresponding new word is added to the predefined category with the highest cosine similarity, and the word vector of the predefined category is updated.

[0037] Preferably, the vector space mapping process specifically comprises the following steps:

[0038] Based on the orthogonal Procrustes method, the mean square error between the source vector set and the target vector set is minimized by finding the optimal orthogonal transformation matrix Q.

[0039] The mean square error is:

[0040] ||XQ-Y|| 2 =tr((XQ-Y) T (XQ-Y))

[0041] To ensure the orthogonality of the transformation matrix Q, it is necessary to satisfy:

[0042] Q T Q=I

[0043] By X T The singular value decomposition of Y is obtained, that is:

[0044] X T Y=U∑V T

[0045] Thus we get:

[0046] Q=UV T

[0047] Where: X is the original vector set; Q is the orthogonal transformation matrix; Y is the target vector set; I is the identity matrix; U and V are orthogonal matrices; Σ is a diagonal matrix; T is the transpose of the matrix.

[0048] Preferably, the calculation formula of the cosine similarity is:

[0049]

[0050] Where: is the summation symbol; K is the summation variable; n is the summation upper limit; E i 、E j are two word vectors for cosine similarity calculation; e i,k 、e j,k are the dimension values ​​of the two word vectors.

[0051] The present invention also provides a Chinese argot recognition system based on dynamic word embedding, comprising:

[0052] An acquisition module is used to acquire text data to be detected;

[0053] A new word dictionary building module is used to obtain new word vocabulary from the text data to be detected through new word discovery technology and boundary enhancement technology, so as to build a new word dictionary;

[0054] A word segmentation processing module is used to perform word segmentation processing on the text data to be detected according to the new word dictionary and the basic dictionary to obtain a word segmentation result; the word segmentation result includes new words and known words;

[0055] A known word comparison module is used to compare the known word with the formal corpus to identify whether the known word is a codeword;

[0056] The new word comparison module is used to compare the new word with the predefined categories and identify whether the new word is a codeword.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. Data diversity and model scalability: Existing lingo recognition technologies are mainly focused on a single field, lacking data diversity, which results in limited model scalability. This invention significantly improves data diversity and enhances the adaptability of the model in different environments by integrating information platform data.

[0059] 2. New word discovery capability: The new word discovery and boundary enhancement technology introduced in the present invention can identify emerging lingo in a specific field and prepare to identify the boundaries of emerging words from the inside, thereby effectively improving the accuracy of lingo recognition, especially when facing unknown or newly emerging lingo.

[0060] 3. Contextual understanding and dynamic word embedding: In order to accurately understand the meaning of lingo in different contexts, the present invention adopts the dynamic word embedding model RoBERTa-wwm-ext based on Chinese pre-training. This model is designed for Chinese tasks and has been trained with large-scale Chinese data. It can accurately distinguish the meaning of words according to context changes, significantly improving the effect of lingo recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.

[0062] Figure 1 It is a flow chart of a Chinese argot recognition method based on dynamic word embedding of the present invention. DETAILED DESCRIPTION

[0063] Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific implementation disclosed below.

[0064] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0065] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0066] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0067] like Figure 1 As shown, the present invention provides a Chinese argot recognition method based on dynamic word embedding, comprising the following steps:

[0068] Get the text data to be detected;

[0069] Through new word discovery technology and boundary enhancement technology, new word vocabulary is obtained from the text data to be detected, so as to build a new word dictionary;

[0070] According to the new word dictionary and the basic dictionary, the text data to be detected is segmented to obtain a segmentation result; the segmentation result includes new words and known words;

[0071] Compare the known words with the formal corpus to identify whether the known words are code words;

[0072] Compare the new words with predefined categories to identify whether the new words are cryptic.

[0073] Preferably, the new word discovery technology adopts an information entropy-based method to measure the context diversity of candidate words by calculating the left and right information entropies of the candidate words;

[0074] The boundary enhancement technology accurately identifies word boundaries by capturing local information of candidate words.

[0075] Preferably, the calculation formula of the information entropy is:

[0076]

[0077] Where: H(X) represents information entropy, which indicates the uncertainty of random variable X; P(x i ) represents the probability of event xi occurring; n represents the total number of possible events.

[0078] Preferably, based on the formal corpus, the known word is compared with the formal corpus to identify whether the known word is a codeword, which specifically includes the following steps:

[0079] Vectorize the known words in the formal corpus to obtain the formal corpus word vectors;

[0080] Perform vectorization processing on known words in the text data to be detected to obtain known word vectors;

[0081] After the formal corpus word vectors and known word vectors are mapped into vector space, cosine similarity comparison is performed;

[0082] If the cosine similarity is greater than or equal to a preset threshold value t1, the corresponding known word is determined to be a non-cryptographic word;

[0083] If the cosine similarity is lower than the first preset threshold t1, the corresponding known word is determined to be a codeword.

[0084] Preferably, based on the formal corpus, the following steps are also included:

[0085] The known word vectors of the known words determined to be secret words are respectively calculated with the word vectors of several predefined categories for cosine similarity;

[0086] Add the known words to the predefined category with the highest cosine similarity and update the word vector of the predefined category.

[0087] Preferably, based on the formal corpus, the new word is compared with the predefined category to identify whether the new word is a codeword, which specifically includes the following steps:

[0088] Vectorize the new words in the text data to be detected to obtain new word vectors;

[0089] Calculate the cosine similarity between the new word vector and the word vectors of several predefined categories;

[0090] If the cosine similarity between the new word vector and the word vectors of all predefined categories is lower than the second preset threshold t2, the corresponding new word is determined to be a non-cryptographic word;

[0091] If the cosine similarity between the new word vector and the word vectors of all predefined categories is greater than or equal to the second preset threshold t2, the corresponding new word is determined to be a codeword, and the corresponding new word is added to the predefined category with the highest cosine similarity, and the word vector of the predefined category is updated.

[0092] Preferably, the vector space mapping process specifically comprises the following steps:

[0093] Based on the orthogonal Procrustes method, the mean square error between the source vector set and the target vector set is minimized by finding the optimal orthogonal transformation matrix Q.

[0094] The mean square error is:

[0095] ||XQ-Y|| 2 =tr((XQ-Y) T (XQ-Y))

[0096] To ensure the orthogonality of the transformation matrix Q, it is necessary to satisfy:

[0097] Q T Q=I

[0098] By X T The singular value decomposition of Y is obtained, that is:

[0099] X T Y=U∑V T

[0100] Thus we get:

[0101] Q=UV T

[0102] Where: X is the original vector set; Q is the orthogonal transformation matrix; Y is the target vector set; I is the identity matrix; U and V are orthogonal matrices; Σ is a diagonal matrix; T is the transpose of the matrix.

[0103] Preferably, the calculation formula of the cosine similarity is:

[0104]

[0105] Where: is the summation symbol; K is the summation variable; n is the summation upper limit; E i 、E j are two word vectors for cosine similarity calculation; e i,k 、e j,k are the dimension values ​​of the two word vectors.

[0106] The present invention also provides a Chinese argot recognition system based on dynamic word embedding, comprising:

[0107] An acquisition module, used to acquire text data to be detected;

[0108] A new word dictionary building module is used to obtain new word vocabulary from the text data to be detected through new word discovery technology and boundary enhancement technology, so as to build a new word dictionary;

[0109] A word segmentation processing module is used to perform word segmentation processing on the text data to be detected according to the new word dictionary and the basic dictionary to obtain a word segmentation result; the word segmentation result includes new words and known words;

[0110] A known word comparison module is used to compare the known word with the formal corpus to identify whether the known word is a codeword;

[0111] The new word comparison module is used to compare the new word with the predefined categories and identify whether the new word is a codeword.

[0112] The present invention proposes a Chinese lingo recognition technology based on dynamic word embedding, aiming to improve the accuracy and adaptability of lingo recognition. Traditional lingo recognition methods usually rely on static word embedding or dictionaries, which cannot effectively deal with the variability and context dependence of lingo. In order to overcome this limitation, the present invention introduces a dynamic word embedding technology suitable for Chinese lingo recognition for the first time, and generates context-sensitive word vectors based on the Chinese large-scale pre-trained deep learning model RoBERTa-wwm-ext. This method can generate different word vectors according to different contexts and effectively handle the changes in lexical meanings in lingo. In addition, this study introduces new word discovery and boundary enhancement technology, and identifies potential emerging lingo through in-depth mining of the information platform. This technology improves the adaptability of the lingo recognition system to dynamic changes in lingo. Through vector space mapping and similarity analysis, the recognition effect is further optimized, enabling the system to accurately identify new lingo expressions and enhance the understanding of complex semantics.

[0113] In order to better illustrate the technical effect of the present invention, the present invention provides the following specific embodiments to illustrate the above technical process:

[0114] Embodiment 1: A Chinese argot recognition method based on dynamic word embedding has significant advantages in accuracy, adaptability and real-time performance of argot recognition, and can provide valuable information for intelligence analysis, specifically including:

[0115] Module 1: Comparative Corpus Construction Module

[0116] In order to construct word vectors for normal corpora, this module obtains a large amount of text data from the information platform through web crawlers. After strict data cleaning and preprocessing of the collected data, the RoBERTa-wwm-ext model is used to generate high-quality word vectors. These word vectors are then stored in a KV-type database for efficient retrieval and management. To further improve the efficiency of data storage and retrieval, the system chooses to use the Redis database, making full use of its high-performance memory storage characteristics, supporting persistent storage of large-scale data, and enhancing the scalability and overall processing capabilities of the system through a cluster structure.

[0117] Module 2: Word vector generation and new word discovery module

[0118] To effectively identify lingo, we use a unified encoding model to process the input text. Specifically, we use the dynamic word embedding model RoBERTa-wwm-ext to generate word vectors, and use the vector averaging method to generate vector representations for the segmented words. However, since lingo may contain a large number of new words, directly generating word vectors may lead to segmentation errors or semantic deviations. To this end, before generating word vectors, we introduced new word discovery and boundary enhancement technology to improve the recognition of new words.

[0119] New word discovery technology uses an entropy-based method to measure the contextual diversity of candidate words by calculating the left and right entropies of candidate words. Entropy is an indicator that reflects the uncertainty of random variables and can effectively evaluate the degree of freedom of combination of words and their contexts. Boundary enhancement technology captures local information of words and accurately identifies word boundaries, thereby effectively improving the accuracy of new word discovery. By combining entropy analysis and boundary enhancement technology, we can more accurately identify new words in lingo and further improve the semantic expression quality of subsequent encoding processes. This strategy significantly enhances the model's adaptability to lingo and recognition accuracy.

[0120]

[0121] Where:

[0122] H(X): Information entropy, which represents the uncertainty of random variable X.

[0123] P(x i ): The probability of event xi occurring.

[0124] n: The total number of possible events.

[0125] Module 3: Vector Space Mapping Module

[0126] In the similarity comparison, since the word vectors generated in different corpora may lack comparability, the present invention performs vector space mapping, specifically by introducing the Orthogonal Procrustes Problem. This method aims to minimize the mean square error between the source vector set and the target vector set by finding the optimal orthogonal transformation matrix Q.

[0127] Specifically, the mean square error can be expressed as:

[0128] ||XQ-Y|| 2 =tr((XQ-Y) T (XQ-Y))

[0129] To ensure the orthogonality of the transformation matrix Q, it is necessary to satisfy:

[0130] Q T Q=I

[0131] This problem is solved by T The singular value decomposition (SVD) of Y is obtained, that is,

[0132] X T Y=U∑V T

[0133] Thus we get:

[0134] Q=UV T

[0135] This enables the vectors to be mapped to the same plane and solves the problem of large errors when the vectors are directly compared.

[0136] Where: X is the original vector set; Q is the orthogonal transformation matrix; Y is the target vector set; I is the identity matrix; U and V are orthogonal matrices; Σ is a diagonal matrix; T is the transpose of the matrix.

[0137] Module 4 Similarity comparison module

[0138] For known words, after being mapped into vector space, cosine similarity is used for comparison. This method is concise, efficient and can accurately measure semantic similarity.

[0139] For example: two word vectors E i [e i,1 ,e i,2 ,...,e i,n ] and E j [e j,1 ,e j,2 ,...,e j,n ] can be expressed as:

[0140]

[0141] Where: is the summation symbol; K is the summation variable; n is the summation upper limit; E i 、E j are two word vectors for cosine similarity calculation; e i,k 、e j,k is the dimension value of the two word vectors;

[0142] Crawl the information platform to obtain the formal corpus; for known words, obtain its word vector expression in this formal corpus. Then obtain the word vector expression of the known word in the text data to be detected. Compare the cosine similarity of the two word vector expressions, that is, compare the word vectors of the same vocabulary in different contexts. If the similarity of the two word vectors is relatively large, it means that the meaning of the word in the two contexts is similar; otherwise, the meaning is quite different. That is: if the similarity between the two is greater than or equal to the first preset threshold t1, it is determined that the meaning of the word in the two contexts is consistent, and it is considered not to be a codeword; if the similarity is lower than the first preset threshold t1, it is considered to be a codeword.

[0143] Next, we will classify the known words that are determined to be codewords. The classification process is as follows:

[0144] There are 10 predefined categories, examples of the category format: Fruit: [banana, ..., pear];

[0145] During initialization, each predefined category contains several words, and the word vector of the category has been calculated. The word vector of the category will change dynamically according to the words it contains. During classification, you only need to compare the word vector of the known word in the text to be detected with the word vector of the category. After comparing the known word with the word vectors of the 10 categories, add the known word to the predefined category with the highest similarity, and update the word vector of the predefined category.

[0146] For example, the known word to be detected is red. Red is compared with the word vectors of 10 predefined categories. It is found that the word vector has the highest similarity with the fruit category. Red is added to the fruit category and the word vector of the fruit category is updated. For example, fruit: [banana, red..., pear]. The original word vector of the fruit category is (0.3, 0.5, 0.7), the word vector of red is (0.1, 0.1, 0.1), and the word vector of the fruit category after the update is (0.2, 0.3, 0.4).

[0147] For the classification of lingo, similarity is calculated between it and the word vectors of 10 predefined categories, and the category with the highest similarity is classified, and the word vector of the category is updated.

[0148] For emerging words, the similarity is directly compared with the word vectors of each predefined category. If the maximum similarity of all predefined categories is lower than the second preset threshold t2, the word is determined to be non-cryptic; if the similarity of a predefined category exceeds the second preset threshold t2, it is classified into the category with the highest similarity and the word vector of this category is updated.

[0149] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules, modules or units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units, modules or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0150] The units may or may not be physically separated, and the components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0151] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0152] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit (CPU), the above-mentioned functions defined in the method of the present invention are executed. It should be noted that the above-mentioned computer-readable medium of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of an electrical, magnetic, optical, electromagnetic, infrared segment, or semiconductor, or any combination of the above.

[0153] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0154] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A Chinese argot recognition method based on dynamic word embedding, characterized in that: The following steps are involved: Get the text data to be detected; Through new word discovery technology and boundary enhancement technology, new word vocabulary is obtained from the text data to be detected, so as to build a new word dictionary; According to the new word dictionary and the basic dictionary, the text data to be detected is segmented to obtain a segmentation result; the segmentation result includes new words and known words; Compare the known words with the formal corpus to identify whether the known words are code words; Compare the new words with predefined categories to identify whether the new words are cryptic.

2. The Chinese argot recognition method based on dynamic word embedding according to claim 1 is characterized in that: The new word discovery technology adopts an information entropy-based method to measure the context diversity of candidate words by calculating the left and right information entropies of the candidate words; The boundary enhancement technology accurately identifies word boundaries by capturing local information of candidate words.

3. The Chinese argot recognition method based on dynamic word embedding according to claim 2 is characterized in that: The calculation formula of the information entropy is: Where: H(X) represents information entropy, which indicates the uncertainty of random variable X; P(x i ) represents the probability of event xi occurring; n represents the total number of possible events.

4. The Chinese argot recognition method based on dynamic word embedding according to claim 3 is characterized in that: Based on the formal corpus, the known words are compared with the formal corpus to identify whether the known words are code words, which specifically includes the following steps: Vectorize the known words in the formal corpus to obtain the formal corpus word vectors; Perform vectorization processing on known words in the text data to be detected to obtain known word vectors; After the formal corpus word vectors and known word vectors are mapped into vector space, cosine similarity comparison is performed; If the cosine similarity is greater than or equal to a preset threshold value t1, the corresponding known word is determined to be a non-cryptographic word; If the cosine similarity is lower than the first preset threshold t1, the corresponding known word is determined to be a codeword.

5. The Chinese argot recognition method based on dynamic word embedding according to claim 4 is characterized in that: Based on the formal corpus, the following steps are also included: The known word vectors of the known words determined to be secret words are respectively calculated with the word vectors of several predefined categories for cosine similarity; Add the known words to the predefined category with the highest cosine similarity and update the word vector of the predefined category.

6. The Chinese argot recognition method based on dynamic word embedding according to claim 5 is characterized in that: Based on the formal corpus, the new words are compared with the predefined categories to identify whether the new words are code words, which specifically includes the following steps: Vectorize the new words in the text data to be detected to obtain the new word vectors; Calculate the cosine similarity between the new word vector and the word vectors of several predefined categories; If the cosine similarity between the new word vector and the word vectors of all predefined categories is lower than the second preset threshold t2, the corresponding new word is determined to be a non-cryptographic word; If the cosine similarity between the new word vector and the word vectors of all predefined categories is greater than or equal to the second preset threshold t2, the corresponding new word is determined to be a codeword, and the corresponding new word is added to the predefined category with the highest cosine similarity, and the word vector of the predefined category is updated.

7. The Chinese argot recognition method based on dynamic word embedding according to claim 6 is characterized in that: The vector space mapping process specifically includes the following steps: Based on the orthogonal Procrustes method, the mean square error between the source vector set and the target vector set is minimized by finding the optimal orthogonal transformation matrix Q. The mean square error is: ||XQ-Y|| 2 =tr((XQ-Y) T (XQ-Y)) To ensure the orthogonality of the transformation matrix Q, it is necessary to satisfy: Q T Q=I By X T The singular value decomposition of Y is obtained, that is: X T Y=U∑V T Thus we get: Q=UV T Where: X is the original vector set; Q is the orthogonal transformation matrix; Y is the target vector set; I is the identity matrix; U and V are orthogonal matrices; Σ is a diagonal matrix; T is the transpose of the matrix.

8. The Chinese argot recognition method based on dynamic word embedding according to claim 7 is characterized in that: The calculation formula of the cosine similarity is: Where: is the summation symbol; K is the summation variable; n is the summation upper limit; Ei and Ej are two word vectors for cosine similarity calculation; i,k 、e j,k are the dimension values ​​of the two word vectors.

9. A Chinese argot recognition system based on dynamic word embedding, used to implement the Chinese argot recognition method based on dynamic word embedding as claimed in any one of claims 1 to 8, characterized in that: include: An acquisition module, used to acquire text data to be detected; A new word dictionary building module is used to obtain new word vocabulary from the text data to be detected through new word discovery technology and boundary enhancement technology, so as to build a new word dictionary; A word segmentation processing module is used to perform word segmentation processing on the text data to be detected according to the new word dictionary and the basic dictionary to obtain a word segmentation result; the word segmentation result includes new words and known words; A known word comparison module is used to compare the known word with the formal corpus to identify whether the known word is a codeword; The new word comparison module is used to compare the new word with the predefined categories and identify whether the new word is a codeword.