Word set acquisition method, intelligent measurement terminal application detection method and device based on word set, equipment and storage medium

By inserting interval encoding into the interface call sequence of the intelligent measurement terminal and building a word set, using word embedding vectors to obtain the correlation degree of the model identification interface encoding, the problem of difficulty in detecting abnormal applications in the existing technology is solved, and more efficient abnormal application detection and malicious code analysis are achieved.

CN120278152AActive Publication Date: 2025-07-08GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510749796.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify abnormal applications of intelligent measurement terminals, especially in interface call combinations, which leads to an increase in the security risks of power grids.

Method used

By obtaining the interface call sequence of the sample application, the insertion interval coding process the problem that the time interval is too long, the word embedding vector acquisition model is used for self-supervised training, the word set is constructed to identify the correlation degree of interface coding, and abnormal application detection is performed in combination with word segmentation technology.

Benefits of technology

It improves the accuracy of intelligent measurement terminal abnormal application detection, can effectively identify interface call order combination and irrelevant API insertion, and enhances the ability to locate and analyze malicious code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278152A_ABST
    Figure CN120278152A_ABST
Patent Text Reader

Abstract

The invention relates to a word set acquisition method, an intelligent measurement terminal application detection method and device based on a word set, equipment and a readable storage medium, relates to the technical field of power grids, and can improve the abnormal application detection accuracy of an intelligent measurement terminal. The method comprises the steps of obtaining a sample interface calling sequence; according to an interval code, a first interception interval and a first interface code serving as an interval center in the sample interface calling sequence, a first interface code sequence is obtained from the sample interface calling sequence, self-supervision training is conducted on the word embedding vector obtaining model according to the first interface code sequence, and when training is finished, the word embedding vector obtaining model is obtained; determining a word embedding vector corresponding to each interface code according to a word embedding vector acquisition model; according to the sample interface calling sequences, obtaining candidate words, according to the word embedding vectors of the interface codes, determining the coding association degree of the interface codes contained in the candidate words, and according to the candidate words with the coding association degree meeting a preset condition, constructing a word set of multiple application types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of power grids, and particularly to a method for obtaining a word set for intelligent measurement terminal application detection, an intelligent measurement terminal application detection method based on the word set, a device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] As a key component of the power grid, the intelligent measurement terminal connects the master station system and various metering devices, and plays an indispensable role in the accurate capture of the power grid's panoramic data. With the digitization and intelligence of the power industry, the intelligent measurement terminal can flexibly possess various required functions by installing various application software. However, the flexible installation of application software also brings potential security risks.

[0003] In the related art, the vocabulary-level semantic analysis technology in natural language processing can be used to analyze the interface relationships of the interfaces (Application Programming Interface, API) called by the applications on the intelligent measurement terminal during operation. For example, the interface is embedded and encoded as a word, and then the word embedding vectors of the interfaces are clustered, and the application detection result is determined according to the clustering result.

[0004] However, the above method is difficult to effectively identify the interface call combinations of abnormal applications and accurately detect abnormal applications on the intelligent measurement terminal. Summary of the Invention

[0005] Based on this, it is necessary to provide a method for obtaining a word set for intelligent measurement terminal application detection, an intelligent measurement terminal application detection method based on the word set, a device, a computer device, a computer-readable storage medium, and a computer program product for the above technical problems.

[0006] In a first aspect, this application provides a method for obtaining a word set for intelligent measurement terminal application detection, and the method includes:

[0007] Obtain the sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include the interface codes and interval codes corresponding to each platform interface called in sequence when the sample applications are running; the interval codes are inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold;

[0008] According to the interval coding, the first truncation interval with a preset length, and the first interface coding as the interval center in the sample interface call sequence, obtain the first interface coding sequence from the sample interface call sequence. Perform self-supervised training on the word embedding vector acquisition model according to multiple first interface coding sequences. At the end of the training, determine the word embedding vectors corresponding to each interface coding according to the word embedding vector acquisition model; the word embedding vector acquisition model is used to determine the word embedding vectors according to each interface coding in the first interface coding sequence and predict the first interface coding according to the word embedding vectors.

[0009] According to each sample interface call sequence, obtain candidate words containing at least two interface codings. According to the word embedding vectors corresponding to each interface coding, determine the coding correlation degrees of the interface codings included in each candidate word. Construct word sets of multiple application types according to the candidate words whose coding correlation degrees meet the preset conditions; the word sets are used to detect the application types of intelligent measurement terminals.

[0010] In one embodiment, the determining the coding correlation degrees of the interface codings included in each candidate word according to the word embedding vectors corresponding to each interface coding, and constructing word sets of multiple application types according to the candidate words whose coding correlation degrees meet the preset conditions includes:

[0011] For each candidate word, obtain the adjacency entropy corresponding to the candidate word. According to the word embedding vectors corresponding to the interface codings included in the candidate word and the occurrence probability of the interface coding sequence corresponding to the candidate word, determine the overall tightness of the candidate word, and determine the coding correlation degrees of the interface codings included in the candidate word according to the adjacency entropy and the overall tightness of the candidate word.

[0012] Add each candidate word with a coding correlation degree greater than a preset threshold to the word library;

[0013] Perform word segmentation on the sample interface call sequences under each application type according to the word library, and obtain the word sets of each application type according to the word segmentation results.

[0014] In one embodiment, the determining the overall tightness of the candidate word according to the word embedding vectors corresponding to the interface codings included in the candidate word and the occurrence probability of the interface coding sequence corresponding to the candidate word includes:

[0015] Divide the candidate word to obtain multiple groups of front-segment interface coding sequences and back-segment interface coding sequences;

[0016] For each set of the front - end interface coding sequences and the back - end interface coding sequences, determine the splitting probability according to the occurrence probabilities of the candidate word, the front - end interface coding sequence, and the back - end interface coding sequence respectively, and determine the semantic similarity between the front - end interface coding sequence and the back - end interface coding sequence according to the word embedding vectors corresponding to each interface coding included in the candidate word;

[0017] Determine the overall tightness of the candidate word according to the splitting probabilities and the semantic similarities corresponding to multiple sets of the front - end interface coding sequences and the back - end interface coding sequences.

[0018] In one embodiment, the tokenizing the sample interface call sequences for each application type according to the thesaurus, and obtaining the word sets for each application type according to the tokenization results includes:

[0019] According to the interval coding, the second intercepting interval with a preset length, and the second interface coding as the center of the interval in the sample interface call sequence, obtain the second interface coding sequence from the current sample interface call sequence in reverse;

[0020] Judge whether there is a matching word for the second interface coding sequence in the thesaurus; the matching word is an interface coding combination including the first interface coding and the last interface coding in the current second interface coding sequence;

[0021] If so, take the current second interface coding sequence as the tokenization result, split it from the current sample interface call sequence, add it to the word set corresponding to the corresponding application type, and if the current sample interface call sequence is not completely split, return to execute the step of obtaining the second interface coding sequence from the current sample interface call sequence;

[0022] If not, remove the first interface coding from the current second interface coding sequence, and return to execute the step of judging whether there is a matching word in the thesaurus.

[0023] In one embodiment, the obtaining the first interface coding sequence from the sample interface call sequence according to the interval coding, the first intercepting interval with a preset length, and the first interface coding as the center of the interval in the sample interface call sequence includes:

[0024] In the sample interface call sequence, perform sequence interception with the first interface coding as the center of the interval and the first intercepting interval with a preset length of 2N + 1 to obtain an intercepted sequence; N is a positive integer;

[0025] When there is the interval code among the N interface codes before the first interface code, set each interface code in the intercepted sequence before the interval code to the interval code, and when there is the interval code among the N interface codes after the first interface code, set each interface code in the intercepted sequence after the interval code to the interval code;

[0026] Obtain the first interface code sequence according to the intercepted sequence obtained by setting the interval code.

[0027] In a second aspect, the present application also provides an intelligent measurement terminal application detection method based on a word set, and the method includes:

[0028] Obtain the actual interface call sequence corresponding to the target application to be detected by the intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to multiple platform interfaces sequentially called during the operation of the target application;

[0029] Segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code;

[0030] If each word segment of the actual interface call sequence is included in the word set of the normal type application, determine that the target application is the normal type application;

[0031] If some word segments of the actual interface call sequence are included in the word set of the abnormal type application, determine that the target application is the abnormal type application, and determine the abnormal code in the target application according to the interface call codes corresponding to each interface code in some word segments;

[0032] Wherein, each word set is obtained according to the word set acquisition method for intelligent measurement terminal application detection described in any one of the above.

[0033] In a third aspect, the present application also provides a word set acquisition device for intelligent measurement terminal application detection, and the device includes:

[0034] A sample acquisition module, configured to acquire the sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include interface codes and interval codes corresponding to each platform interface sequentially called during the operation of the sample applications; the interval code is inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold;

[0035] A word embedding acquisition module, configured to obtain a first interface code sequence from the sample interface call sequence according to the interval coding, the first truncation interval with a preset length, and the first interface code as the interval center in the sample interface call sequence, and perform self-supervised training on the word embedding vector acquisition model according to a plurality of the first interface code sequences. At the end of the training, determine the word embedding vectors corresponding to each of the interface codes according to the word embedding vector acquisition model; the word embedding vector acquisition model is configured to determine word embedding vectors according to each interface code in the first interface code sequence, and predict the first interface code according to the word embedding vectors;

[0036] A word set acquisition module, configured to obtain candidate words including at least two interface codes according to each of the sample interface call sequences, determine the coding correlation degree of each interface code included in each of the candidate words according to the word embedding vectors corresponding to each of the interface codes, and construct word sets of a plurality of the application types according to the candidate words that meet the preset conditions; the word sets are used to detect the application types of intelligent measurement terminal applications.

[0037] Fourthly, the present application further provides an intelligent measurement terminal application detection device based on word sets, and the device includes:

[0038] An actual sequence acquisition module, configured to obtain an actual interface call sequence corresponding to a target application to be detected by an intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to a plurality of platform interfaces sequentially called during the operation of the target application;

[0039] A word segmentation module, configured to segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code;

[0040] A first recognition module, configured to determine that the target application is the normal type application if each of the word segments of the actual interface call sequence is included in the word set of the normal type application;

[0041] A second recognition module, configured to determine that the target application is the abnormal type application if some of the word segments of the actual interface call sequence are included in the word set of the abnormal type application, and determine the abnormal code in the target application according to the interface call codes corresponding to each of the interface codes in some of the word segments;

[0042] Wherein, each of the word sets is obtained according to the word set acquisition method for intelligent measurement terminal application detection described in any one of the above.

[0043] Fifthly, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the method described in any one of the above when executing the computer program.

[0044] In a sixth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0045] In a seventh aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0046] The above method for obtaining a word set for intelligent measurement terminal application detection, the method for intelligent measurement terminal application detection based on the word set, the device, the computer device, the computer-readable storage medium, and the computer program product. On the one hand, by inserting interval coding in the sample call sequence, it is possible to avoid using other interfaces with very low correlation as the context of the current interface, avoid adding noise to the word embedding vector, improve the accuracy and expression effect according to the word embedding vector, and effectively handle the problem of too long interface call time intervals. On the other hand, by obtaining the word embedding vector corresponding to the interface coding, determining the coding correlation degree of each interface coding included in the candidate word according to the word embedding vector, and constructing word sets for multiple application types accordingly, each called interface can be regarded as a word, and the interface call combination can be regarded as a word. The word sets constructed through unsupervised training combine the interface call sequences with strong relevance into words, which can effectively enhance the ability to represent the combination of interface call orders and handle the problem of inserting irrelevant API calls. Subsequently, by combining with the word segmentation technology, accurate detection of interface call exceptions can be achieved, effectively improving the detection accuracy of abnormal applications of intelligent measurement terminals. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of a method for obtaining a word set for intelligent measurement terminal application detection in an embodiment;

[0049] Figure 2 It is a schematic flowchart of a method for intelligent measurement terminal application detection based on a word set in an embodiment;

[0050] Figure 3 It is a schematic flowchart of another method for intelligent measurement terminal application detection based on a word set in an embodiment;

[0051] Figure 4 The structural block diagram of a word set acquisition device for intelligent measurement terminal application detection in an embodiment;

[0052] Figure 5 The structural block diagram of an intelligent measurement terminal application detection device based on a word set in an embodiment;

[0053] Figure 6 The internal structure diagram of a computer device in an embodiment;

[0054] Figure 7 The internal structure diagram of another computer device in an embodiment. Detailed implementation manners

[0055] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] In order to enable those skilled in the art to better understand the present application, the related technologies will be introduced first below.

[0057] An intelligent measurement terminal is an integrated and intelligent power device that can collect electrical parameters such as voltage, current, power factor, frequency, and harmonics of the power grid in real time, and can upload these data to the power grid dispatching center or cloud platform through a communication interface. At the same time, it can also process and analyze the collected data, providing functions for evaluating, predicting, and alarming the operation status of the power grid. In addition, the intelligent measurement terminal also has a remote control function, which can receive instructions from the superior system to realize remote regulation and optimization of power grid equipment. This function makes the power grid management more convenient and efficient, and improves the safety and reliability of the power grid.

[0058] With the digitization and intelligence of the power industry, higher requirements are put forward for the flexibility of intelligent measurement terminals to flexibly respond to and carry increasingly rich functional requirements. In this context, a combination mode of a configurable software platform and diverse application programs emerges as the times require, enabling intelligent measurement terminals to flexibly adapt to different application scenarios and functional requirements by simply installing or uninstalling various application modules. However, the flexibility and diversity of application programs also bring potential security risks. To ensure the safe and stable operation of the power grid, the application programs installed on the intelligent measurement terminal (such as before being put on the shelf and / or after being deployed to the intelligent measurement terminal) need to be fully tested and verified to identify abnormal codes in the applications.

[0059] Abnormal codes in applications (such as malicious codes) are software programs designed to damage, interfere with, or endanger computer systems. Malicious codes seriously threaten the operation safety of the power grid through one or more ways including but not limited to attacking computer systems and destroying system functions. Therefore, it is particularly important to conduct strict code detection on the application programs of intelligent metering terminals. To effectively identify and prevent such threats, the detection methods can include signature detection (comparison based on a known malicious code signature library), heuristic detection (inferring potential maliciousness through behavior patterns), behavior analysis (monitoring the specific behaviors during program operation), and cloud detection (rapidly identifying unknown threats using cloud big data and machine learning technologies), etc. As a key device in the power grid system, the intelligent metering terminal has a unique application program operating environment. The communication between the application program and the master station and metering devices, as well as the data exchange between application programs, are realized through specific software platform interfaces. Therefore, an effective method for detecting abnormal codes is to observe and record the interface call behaviors during the operation of the application program, and identify potential malicious codes by comparing the differences between normal operations and potential abnormal behaviors.

[0060] Using interface call behaviors to identify malicious codes mainly utilizes the pattern differences in the types, quantities, and API call sequences of APIs in abnormal codes and normal codes, including methods such as sequence matching, association rules, hybrid feature encoders, clustering, word embedding encoding, and semantic analysis.

[0061] In related technologies, when analyzing the interfaces called during the operation of an application using the lexical-level semantic analysis technology in natural language processing, first, the API call sequence is statically extracted, and each API is embedded and encoded as a word to obtain the word embedding vector of the API. Among them, word embedding is a technology in natural language processing (Natural Language Processing, NLP), which maps words or phrases in the vocabulary to real-valued vectors in a high-dimensional space. These vectors can generally capture the semantic relationships between words, that is, similar words will be close to each other in the vector space. The goal of word embedding is to convert language symbols (such as words) into a form that can be understood and processed by a computer, so as to enable effective calculation and reasoning. Then, the word embedding vectors of the APIs are clustered, the application program is converted into the encoding of the cluster, and finally, the abnormal application is identified through the cluster encoding and integration model included in the application program.

[0062] However, the above method is difficult to effectively identify the interface call combinations of abnormal applications, has weak ability to extract the API call combination features of abnormal codes, is not conducive to the positioning of abnormal codes and the analysis of their API call features, and is difficult to accurately detect abnormal applications of intelligent metering terminals.

[0063] Based on this, it is necessary to provide a method for obtaining a word set for intelligent measurement terminal application detection, an intelligent measurement terminal application detection method based on the word set, a device, a computer device, a computer-readable storage medium, and a computer program product for the above technical problems.

[0064] In one embodiment, as Figure 1 shown, a method for obtaining a word set for intelligent measurement terminal application detection is provided. In this embodiment, taking the application of this method to a terminal as an example, it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0065] S101, obtain the sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include the interface codes and interval codes corresponding to each platform interface called in sequence during the operation of the sample application; the interval code is inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold.

[0066] In specific implementation, a list of intelligent measurement terminal software platform interfaces (which can also be simply referred to as APIs) can be established in advance, and each API can be encoded to obtain the interface code corresponding to each platform interface. In addition, a code representing interruption (i.e., the interval code) can be added. In some examples, the APIs can be encoded using one-hot encoding.

[0067] Then, multiple sample applications with pre-labeled application types can be obtained. Each application type can include one or more applications. For the convenience of distinction, the applications used to construct the word set are called sample applications. In some embodiments, the multiple application types can include normal type applications and abnormal type applications. Among them, normal type applications refer to applications that can run normally and do not attack the system, and abnormal type applications can refer to applications that contain abnormal codes and / or attack the system. For abnormal type applications, different abnormal type applications can also be divided according to the specific situation and specific manner of the occurrence of the abnormality.

[0068] After obtaining sample applications under multiple application types, in some embodiments, the sample applications with labeled application types can be run separately in a sandbox. For each sample application, the software platform interfaces called during the runtime of the sample application can be recorded in sequence. Based on the interface codes corresponding to the platform interfaces and the call time intervals between the platform interfaces, a call timing sequence of APIs is generated, which is also referred to as an interface call sequence. Herein, the call time interval is the interval between the call time of any called platform interface and the call time of the previous called platform interface. For the sake of distinction, the interface call sequence of the sample application is referred to as the sample interface call sequence. Furthermore, a training dataset can be constructed based on multiple sample interface call sequences.

[0069] It can be understood that in step S101, multiple sample applications with labeled application types can be prepared first as a training set. Among them, the multiple samples include normal application programs and abnormal application programs of different types. Then, the sample applications of each application type are run separately in a sandbox, and the codes of the software platform APIs called during the runtime of each sample application are recorded in sequence, as well as the time interval between the current API and the previous API call. Based on this, an interface call sequence of APIs is generated, and the call order of each called platform interface and the call time interval between two adjacent platform interfaces are recorded in this sequence. Furthermore, the sample interface call sequences obtained after all application programs of each application type are run can constitute the training dataset of this application type.

[0070] After obtaining multiple original sample interface call sequences, it can be determined whether an interval code needs to be inserted into the original sample interface call sequence according to the call time interval, so as to obtain the sample interface call sequence used in the subsequent steps. Specifically, for every two adjacent platform interfaces, when the call time interval between the two adjacent platform interfaces is greater than a preset threshold, an interval code can be inserted between the interface codes corresponding to the two adjacent platform interfaces in the original sample interface call sequence, and its call time is set to half of the call time interval between the two adjacent platform interfaces. For example, if the call time interval corresponding to two adjacent platform interfaces is t1 - t2 and t1 - t2 exceeds the preset threshold, an interval code can be inserted at the position of (t1 - t2) / 2.

[0071] S102, according to the interval codes in the sample interface call sequence, the first truncation interval with a preset length, and the first interface code as the center of the interval, obtain a first interface code sequence from the sample interface call sequence, and perform self-supervised training on the word embedding vector acquisition model according to multiple first interface code sequences. At the end of the training, determine the word embedding vectors corresponding to each interface code according to the word embedding vector acquisition model; the word embedding vector acquisition model is used to determine the word embedding vectors according to each interface code in the first interface code sequence and predict the first interface code according to the word embedding vectors.

[0072] In the related art, the call time interval of the API is not considered during the process of obtaining word embeddings. That is, when the call time interval between two adjacent platform interfaces is too long, the possibility of relevance between the two platform interfaces is very small. If they are used as context for each other, noise will be added to the word embeddings, affecting the expression effect of the word embeddings. In this regard, in this embodiment, an interval encoding can be inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold. That is, for an overly long call interval, this time interval can be regarded as a separator in natural language, and the multiple interface encodings can be more reasonably segmented by inserting the interval encoding, which can make the subsequent feature extraction more accurate.

[0073] Specifically, in this step, according to the interval encoding in the sample interface call sequence, the first truncation interval of a preset length, and the first interface encoding as the center of the interval, an interface encoding sequence is obtained from the sample interface call sequence. That is, first, according to the first truncation interval of a preset length and the first interface encoding as the center of the interval, the corresponding truncation sequence can be determined. Then, according to whether the interval encoding appears in the truncation sequence, the consecutive interface encodings including the first interface encoding in the truncation sequence can be determined, and the interface encoding sequence can be determined according to the consecutive interface encodings. For the convenience of distinction, the interface encoding sequence used in step S102 is called the first interface encoding sequence. In some examples, the first interface encoding sequence can be understood as a sequence containing multiple interface encodings centered on the first interface encoding, which includes the first interface encoding and the context highly relevant to the first interface encoding.

[0074] Then, the self-supervised training of the word embedding vector acquisition model can be performed according to multiple first interface encoding sequences. Specifically, for example, for each first interface encoding sequence, all the interface encodings within the first interface encoding sequence can be merged into a vector as the input of the word embedding vector acquisition model. For example, all the interface encodings within the first interface encoding sequence are concatenated, and the concatenated vector is input into the model. After obtaining the input, the word embedding vector acquisition model can perform feature extraction and output the corresponding feature extraction result in the hidden layer. This feature extraction result can be used as the word embedding vector of the first interface encoding. Then, the word embedding acquisition model can predict the first interface encoding according to the word embedding vector and output it. According to the difference between the predicted result of the first interface encoding and the actual first interface encoding, the model parameters are adjusted. When the training end condition is met (such as the number of model iterations reaches the number threshold or the difference is less than the difference threshold), each first interface encoding sequence can be input into the word embedding vector acquisition model to obtain the word embedding vectors corresponding to each first interface encoding.

[0075] In some examples, a continuous bag-of-words model can be used to train word embedding vectors. Specifically, all the interface encodings within the first interface encoding sequence can be combined into a vector as the input of the continuous bag-of-words model, and the first interface encoding is used as the output. The continuous bag-of-words model is trained in a self-supervised manner. When the training end condition is met, the vector output by the hidden layer in the continuous bag-of-words model is used as the word embedding vector.

[0076] S103. According to each sample interface call sequence, obtain candidate words containing at least two interface encodings. According to the word embedding vectors corresponding to each interface encoding, determine the encoding correlation degree of each interface encoding included in each candidate word. Construct word sets of multiple application types based on the candidate words whose encoding correlation degree meets the preset conditions; the word sets are used to detect the application types of the intelligent metering terminal applications.

[0077] After obtaining the word embedding vectors of each interface encoding, the interface encoding can be regarded as a word, and the combination of interface encodings can be regarded as a word to construct a vocabulary for interface calls. In this embodiment, multiple candidate words can be obtained according to each sample interface call sequence, and each candidate word can contain at least two interface encodings.

[0078] Then, according to the word embedding vectors corresponding to each interface encoding, determine the encoding correlation degree of each interface encoding included in each candidate word, and filter out the candidate words whose encoding correlation degree meets the preset conditions. Based on the candidate words whose encoding correlation degree meets the preset conditions, construct word sets corresponding to multiple application types respectively, so that the application types corresponding to the applications on the intelligent metering terminal can be detected according to the word sets subsequently.

[0079] When the abnormal code adds irrelevant calls between API calls, the detection ability of the related technologies mentioned above is not strong. However, in this embodiment, by regarding the combination of multiple platform interfaces called during the running of the sample application as candidate words and identifying the relevance of each interface in the candidate words according to the word embedding vectors, the abnormal code that inserts irrelevant API calls between key API calls can be effectively detected.

[0080] In the above method for obtaining a word set for intelligent measurement terminal application detection, sample interface call sequences of sample applications under multiple application types can be obtained. Among them, the sample interface call sequence includes the interface codes and interval codes corresponding to each platform interface sequentially called during the operation of the sample application. The interval code is inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold. Then, according to the interval codes in the sample interface call sequence, the first truncation interval with a preset length, and the first interface code as the center of the interval, a first interface code sequence can be obtained from the sample interface call sequence. The word embedding vector acquisition model is self-supervised trained based on multiple first interface code sequences. At the end of the training, the word embedding vectors corresponding to each interface code are determined according to the word embedding vector acquisition model. Furthermore, according to each sample interface call sequence, candidate words containing at least two interface codes are obtained. According to the word embedding vectors corresponding to each interface code, the coding correlation degree of each interface code included in each candidate word is determined. Multiple word sets for detecting the application types of intelligent measurement terminals are constructed according to the candidate words whose coding correlation degree meets the preset conditions. In this embodiment, on the one hand, by inserting interval codes in the sample call sequence, it is possible to avoid using other interfaces with very low correlation as the context of the current interface, avoid adding noise to the word embedding vector, improve the accuracy and expression effect of the word embedding vector, and effectively handle the problem of too long interface call time intervals. On the other hand, by obtaining the word embedding vectors corresponding to the interface codes, determining the coding correlation degree of each interface code included in the candidate word according to the word embedding vector, and constructing multiple word sets of the application types accordingly, each called interface can be regarded as a word, and the interface call combination can be regarded as a word. The word set constructed through unsupervised training combines the interface call sequences with strong relevance into words, which can effectively enhance the ability to represent the interface call order combination and handle the problem of inserting irrelevant API calls. Subsequently, by combining with the word segmentation technology, accurate detection of interface call anomalies can be achieved, effectively improving the detection accuracy of abnormal applications of intelligent measurement terminals.

[0081] In one embodiment, in step S103, according to the word embedding vectors corresponding to each interface code, determining the coding correlation degree of each interface code included in each candidate word, and constructing multiple word sets of application types according to the candidate words whose coding correlation degree meets the preset conditions may include the following steps:

[0082] S1031, for each candidate word, obtain the adjacency entropy corresponding to the candidate word. According to the word embedding vectors corresponding to each interface code included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word, determine the overall tightness of the candidate word, and according to the adjacency entropy and the overall tightness of the candidate word, determine the coding correlation degree of each interface code included in the candidate word.

[0083] In a specific implementation, multiple candidate words can be obtained. In an exemplary embodiment, the maximum string length can be set to 2M + 1 (M is a positive integer), and then a segment of length 2M + 1 is intercepted from the sample interface call sequence centered on the current interface code (also referred to as the third interface code). If there are interval codes among the M interface codes before the current interface code, the interval codes and all the interface codes before them can be removed. If there are interval codes among the M interface codes after the current interface code, the interval codes and all the interface codes after them can be removed. In this way, it is possible to avoid combining interface codes with too long call intervals into the same candidate word, reducing the data noise and interference contained in the candidate words. After removing the interface codes (including interval codes or the interface codes before and / or after the interval codes) and obtaining the corresponding code sequence (also referred to as the third interface code sequence), all interface code combinations containing the current interface code can be selected as candidate words in the third interface code sequence. Specifically, for example, if the third interface code sequence obtained after processing is ABC (where A, B, and C are different interface codes respectively, and the interface code B is the current interface code), the candidate words obtained can include AB, BC, and ABC.

[0084] Then, on the one hand, the adjacency entropy of the candidate word can be obtained. In some exemplary embodiments, the adjacency entropy can be determined in the following manner :

[0085]

[0086] where represent the left adjacency entropy and the right adjacency entropy of the candidate word w respectively. Exemplarily, they can be calculated using the following formula:

[0087]

[0088]

[0089] where is the left adjacent character set of the candidate word w, k is the number of characters in the left adjacent character set, is the right adjacent character set of the candidate word w, is the number of characters in the right adjacent character set, represents the probability that the left adjacent word appears, represents the probability that the right adjacent word appears, and is calculated using the following formula:

[0090]

[0091]

[0092] where The number of occurrences of the represented character on the left side of the word segment w, The number of occurrences of the represented character on the right side of the word segment w.

[0093] On the other hand, the overall tightness of the candidate word can be determined according to the word embedding vectors corresponding to the respective interface codes included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word. Specifically, according to the word embedding vectors corresponding to the respective interface codes, the tightness of each interface code included in the candidate word in terms of semantics and / or grammar can be measured. At the same time, according to the occurrence probability of the interface code sequence included in the candidate word, it can be measured whether it is appropriate to divide each of the interface codes into the same combination. Thus, the overall tightness of multiple interface codes in the candidate word can be determined based on the information from both aspects.

[0094] Then, according to the adjacency entropy and the overall tightness of the candidate word, the coding correlation degree of each interface code included in the candidate word can be determined. In some examples, the coding correlation degree is also referred to as the candidate word score, and it can be calculated in the following manner:

[0095]

[0096] where is the overall tightness of the candidate word w, represents the maximum-minimum normalization operation, which can be shown as follows:

[0097]

[0098] and is the maximum and minimum values of and

[0099] S1032. Add each candidate word with a coding correlation degree greater than the preset threshold to the word library.

[0100] After determining the coding correlation degree of each candidate word, for each candidate word, if its corresponding coding correlation degree is greater than the preset threshold, the candidate word can be added to the word library. If its corresponding coding correlation degree is less than or equal to the preset threshold, the candidate word is discarded.

[0101] S1033. Segment the sample interface call sequences under each application type according to the word library, and obtain the word sets for each application type according to the segmentation results.

[0102] After the thesaurus is constructed, the thesaurus can be used to segment the sample interface call sequences under each application type, and word sets for different application types can be constructed according to the segmented words that appear in the sample interface call sequences of different application types.

[0103] In this embodiment, on the one hand, by considering the adjacency entropy, word embedding vectors, and occurrence probabilities of candidate words to determine the coding correlation degree, comprehensively considering multiple factors can more comprehensively and accurately measure the tight relationship between each interface code in the candidate words, avoiding the one-sidedness of single-factor analysis, and thus more accurately mining the potential connections between interface codes; on the other hand, by adding candidate words with a coding correlation degree greater than a preset threshold to the thesaurus, words with higher correlation degrees and representativeness can be screened out, making the thesaurus more refined and accurate, improving the quality and practicality of the thesaurus, and helping to better adapt to the characteristics and requirements of different application types.

[0104] In an exemplary embodiment, in step S1031, according to the word embedding vectors corresponding to each interface code included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word, determining the overall tightness of the candidate word may include the following steps:

[0105] Divide the candidate word to obtain multiple groups of front-segment interface code sequences and back-segment interface code sequences; for each group of front-segment interface code sequences and back-segment interface code sequences, determine the splitting probability according to the occurrence probabilities of the candidate word, the front-segment interface code sequence, and the back-segment interface code sequence respectively, and determine the semantic similarity between the front-segment interface code sequence and the back-segment interface code sequence according to the word embedding vectors corresponding to each interface code included in the candidate word; determine the overall tightness of the candidate word according to the splitting probabilities and the semantic similarities corresponding to multiple groups of front-segment interface code sequences and back-segment interface code sequences.

[0106] In practical applications, the candidate word can be divided based on each interface code included in the candidate word to obtain multiple groups of front-segment interface code sequences and back-segment interface code sequences. For example, for the i-th character (that is, the i-th interface code that makes up the candidate word), the front-segment interface code sequence and the back-segment interface code sequence can be obtained. The front-segment interface code sequence and the back-segment interface code sequence together form a candidate word with a length of n, that is, the candidate word .

[0107] For each group of front-segment interface code sequences and back-segment interface code sequences, on the one hand, the splitting probability can be determined according to the occurrence probabilities of the candidate word, the front-segment interface code sequence, and the back-segment interface code sequence respectively. In some examples, the splitting probability can be calculated in the following way:

[0108]

[0109] Among them, is the occurrence probability of the above candidate word in the training set, and are the occurrence probabilities of the strings composed of and in all intercepted API strings. The splitting probability can measure the probability of the candidate word appearing as a whole. If and have a smaller ratio, it indicates that after splitting the candidate word into two parts, the probabilities of these two parts appearing separately are relatively high, while the probability of appearing as a whole is relatively low. Then the candidate word may be split into more independent lexical units. On the contrary, if and have a larger ratio, it indicates that after splitting the candidate word into two parts, the probabilities of these two parts appearing separately are relatively low, while the probability of appearing as a whole is relatively high. Then the candidate word may be more suitable to appear as a whole.

[0110] On the other hand, the semantic similarity between the front-end interface code sequence and the back-end interface code sequence can be determined according to the word embedding vectors corresponding to the respective interface codes included in the candidate word. In some examples, the cosine similarity of the mean vectors of the front-end interface code sequence and the back-end interface code sequence can be calculated, which can measure the semantic and syntactic similarity degree of the characters in the front and back parts. The higher the similarity, the closer the semantic and syntactic connection between the front and back parts. The cosine similarity of the mean vectors of the front-end interface code sequence and the back-end interface code sequence can be calculated in the following way:

[0111]

[0112] Among them, and are the means of the word embedding vectors of all characters in the strings and .

[0113] Furthermore, the overall tightness of the candidate word can be determined according to the splitting probabilities and semantic similarities corresponding to multiple groups of front-end interface code sequences and back-end interface code sequences. In one example, the overall tightness can be determined in the following way , the overall tightness is also called the vector enhanced mutual information of the candidate word:

[0114]

[0115] Among them, is the weight and can be set according to empirical values.

[0116] In this embodiment, according to the splitting probabilities corresponding to multiple groups of front - segment interface coding sequences and back - segment interface coding sequences and the semantic similarity, the overall tightness of candidate words is determined. It can comprehensively consider the probability distribution of candidate words and the semantic and syntactic similarities of their character compositions to evaluate the rationality of a candidate word as a complete unit. Through the combination of the two, candidate words can be analyzed and processed more comprehensively.

[0117] In one embodiment, in step S1033, the sample interface call sequences under each application type are segmented according to the word library, and word sets for each application type are obtained according to the segmentation results. The following steps may be included:

[0118] According to the interval coding in the sample interface call sequence, the second interception interval with a preset length, and the second interface coding as the center of the interval, a second interface coding sequence is obtained from the current sample interface call sequence in reverse; it is judged whether the word library includes a matching word for the second interface coding sequence; the matching word is all interface coding combinations including the first interface coding and the last interface coding in the current second interface coding sequence; if so, the current second interface coding sequence is segmented from the current sample interface call sequence as the segmentation result and added to the word set corresponding to the said application type, and when the current sample interface call sequence is not completely segmented, the step of obtaining the second interface coding sequence from the current sample interface call sequence is returned for execution; if not, the first interface coding in the current second interface coding sequence is removed, and the step of judging whether the word library includes a matching word is returned for execution.

[0119] In natural language processing, Chinese word segmentation is the process of splitting a continuous sequence of Chinese characters into meaningful words or phrases. Since there is no space as a natural separator between words in the Chinese writing system, Chinese word segmentation is a basic step when performing natural language processing tasks such as text analysis, information retrieval, machine translation, etc. Except for interfaces with overly long call times, there is also a lack of corresponding segmentation for multiple consecutive interface codes in the sample interface call sequence.

[0120] In this embodiment, the reverse maximum matching algorithm can be used to segment the sample interface call sequence. That is, the second interface code sequence can be obtained from the current sample interface call sequence reversely according to the interval code, the second intercept interval with a preset length, and the second interface code as the interval center in the sample interface call sequence. Specifically, for the sample interface call sequence, the last 2M + 1 (M is a positive integer) characters (i.e., interface codes) can be taken as the segment to be matched. If there is an interval code among the M interface codes before the current interface code, the interval code and all the interface codes before it can be removed. If there is an interval code among the M interface codes after the current interface code, the interval code and all the interface codes after it can be removed, thus obtaining the second interface code sequence.

[0121] Then, it can be determined whether any matching word of the second interface code sequence is included in the thesaurus. The matching word refers to all combinations of interface codes that include the first interface code and the last interface code in the current second interface code sequence. For example, if the first interface code and the last interface code of the second interface code sequence are "A" and "G", respectively, it can be queried in the thesaurus whether all combinations of interface codes in the second interface code sequence that include the first interface code "A" and the last interface code "G" exist in the thesaurus. If so, the matching is successful, and the second interface code sequence can be segmented from the sample interface call sequence as a word and added to the word set of the application type corresponding to the sample interface call sequence. If the sample interface call sequence has been completely segmented, it ends. If the current sample interface call sequence has not been segmented completely, the step of obtaining the second interface code sequence from the current sample interface call sequence is returned for execution.

[0122] If no matching word of the second interface code sequence is included in the thesaurus, the first interface code in the current second interface code sequence can be removed, and the judgment step is returned again to determine whether any matching word of the second interface code sequence is included in the thesaurus. Specifically, for example, if all combinations of interface codes in the second interface code sequence that include the first interface code "A" and the last interface code "G" do not exist in the thesaurus, the first interface code of the second interface code sequence can be removed (i.e., remove the first interface code "A"). If only one interface code remains in the current second interface code sequence after removing the first interface code, it can be determined that the remaining interface code forms a word independently, and then it is judged whether the current sample interface call sequence has been completely segmented. If so, it can end; otherwise, the step of obtaining the second interface code sequence from the current sample interface call sequence is returned for execution. If two or more interface codes remain, the judgment step is returned again to determine whether any matching word of the current second interface code sequence is included in the thesaurus.

[0123] By matching the candidate words in the word library with the interface codes in the sample interface call sequence, it is possible to segment all the sample interface call sequences used for training, form a normal word set with the words used by normal applications, and remove the words in the normal word set from the words used by abnormal applications to form an abnormal word set corresponding to each type of abnormal application. In one example, the abnormal word set can be understood as the feature of malicious code.

[0124] In this embodiment, on the one hand, during the word segmentation process, by determining whether the word library contains the matching words of the second interface code sequence, the existing word library information can be used to guide the word segmentation operation. If the match is successful, the sequence can be directly segmented as the word segmentation result, ensuring that the segmented result conforms to the knowledge in the word library, and improving the rationality and consistency of the word segmentation result. On the other hand, when there is no matching word in the word library, by removing the first interface code in the current second interface code sequence and re-judging, this dynamic adjustment mechanism enables the algorithm to continuously try different combinations to find the word segmentation result that conforms to the word library. By gradually narrowing the range of the intercepted sequence, it is possible to find a suitable word segmentation method as much as possible, avoiding incorrect word segmentation caused by a single non-match judgment, and improving the flexibility and adaptability of the word segmentation. On the other hand, adding the segmented second interface code sequence to the word set corresponding to the application type helps to continuously enrich and improve the word set. As different sample interface call sequences are processed, the word set can cover more interface code combinations, thus better reflecting the characteristics and rules of different application types and providing more comprehensive and accurate data support for subsequent application type detection based on the word set.

[0125] In an exemplary embodiment, in step S102, to obtain the first interface code sequence from the sample interface call sequence according to the interval code, the first intercept interval with a preset length, and the first interface code as the interval center in the sample interface call sequence, the following steps may be included:

[0126] In the sample interface call sequence, perform sequence interception with the first interface code as the interval center and the first intercept interval with a preset length of 2N + 1 to obtain an intercepted sequence; when there is an interval code among the N interface codes before the first interface code, set each interface code before the interval code in the intercepted sequence to the interval code, and when there is an interval code among the N interface codes after the first interface code, set each interface code after the interval code in the intercepted sequence to the interval code; obtain the first interface code sequence according to the intercepted sequence after setting the interval code.

[0127] Where N is a positive integer.

[0128] In practical applications, in the sample interface call sequence, a first truncation interval with the first interface code as the interval center and a preset length of 2N + 1 can be used to truncate the sequence to obtain a truncated sequence. Specifically, the current interface code for which context information is to be extracted, that is, the first interface code, can be determined. Then, with the first interface code as the center, all interface codes within a window of length 2N + 1 are truncated in the sample interface call sequence.

[0129] If there are interval codes among the N interface codes before the first interface code, then all interface codes within the current window and before the interval codes can be set to interval codes. If there are interval codes among the N codes after the first interface code, then all interface codes within the current window and after the interval codes can be set to interval codes. Thus, the first interface code can be separated from irrelevant interface codes with too long call time intervals, avoiding introducing interference information by taking the coding information with too long call time intervals as the context of the first interface code.

[0130] In one embodiment, as Figure 2 shown, a method for detecting intelligent measurement terminal applications based on a word set is also provided. In this embodiment, taking this method applied to a server as an example, it can be understood that this method can also be applied to a terminal and can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0131] S201, obtain the actual interface call sequence corresponding to the target application to be detected by the intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to multiple platform interfaces called sequentially during the operation of the target application.

[0132] In specific implementation, the application to be detected by the intelligent measurement terminal can be determined. For the sake of distinction, this application is called the target application. In some examples, the target application can include other applications that are completely different from the sample application, or can also include an updated version of the sample application.

[0133] After determining the target application, the target application can be run and the interface call sequence of the target application can be obtained. For the sake of distinction, the interface call sequence of the target application is called the actual interface call sequence. The actual interface call sequence includes interface codes corresponding to multiple platform interfaces called sequentially during the operation of the target application. In some examples, interval codes may not be inserted into the actual interface call sequence, that is, the interface codes corresponding to the platform interfaces called sequentially during the operation of the target application can be simply recorded.

[0134] S202, segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code.

[0135] Then, the actual interface call sequence can be segmented to obtain corresponding word segments. In some exemplary embodiments, the word segmenting process of the actual interface call sequence can be exactly the same as that of the sample interface call sequence segmentation. For details, reference can be made to the introduction of the foregoing embodiments, which will not be elaborated here. In other embodiments, other word segmentation algorithms can also be used to segment the actual interface call sequence. For example, it can be segmented according to other maximum matching algorithms (such as forward maximum matching algorithm, bidirectional matching algorithm, etc.).

[0136] S203, if each word segment of the actual interface call sequence is included in the word set of normal type applications, then determine that the target application is a normal type application.

[0137] S204, if some word segments of the actual interface call sequence are included in the word set of abnormal type applications, then determine that the target application is an abnormal type application, and determine the abnormal code in the target application according to the interface call codes corresponding to the interface encodings in the some word segments.

[0138] Among them, each word set can be obtained according to any one or more of the foregoing embodiments of the word set acquisition method for intelligent measurement terminal application detection.

[0139] After obtaining the word segments, it can be determined whether the word segments in the actual interface call sequence are all in the word set of normal type applications. If so, it can be determined that the target application is a normal application program.

[0140] If not, it can be checked whether the word segments outside the word set of normal type applications are included in the word set of abnormal type applications of known abnormal types. If so, it can be determined that the target application is the above abnormal type application, and then the word segments included in the word set of abnormal type applications can be determined, and in the target application code, the call code corresponding to the interface encoding of the word segment can be determined, whereby the location of the abnormal code in the target application can be obtained. If some word segments do not appear in any word set, it can be determined that the target application is an unknown type application.

[0141] In the above method for detecting applications of intelligent measurement terminals based on word sets, the actual interface call sequence corresponding to the target application to be detected by the intelligent measurement terminal can be obtained. The actual interface call sequence includes interface codes corresponding to multiple platform interfaces sequentially called during the operation of the target application. Then, the actual interface call sequence can be segmented to obtain corresponding word segments. Each word segment consists of at least one interface code. If each word segment of the actual interface call sequence is included in the word set of normal type applications, it is determined that the target application is a normal type application. If some word segments of the actual interface call sequence are included in the word set of abnormal type applications, it is determined that the target application is an abnormal type application, and the abnormal code in the target application is determined according to the interface call codes corresponding to the interface codes in the some word segments. Among them, each word set can be obtained according to the word set acquisition method for detecting applications of intelligent measurement terminals in any of the foregoing embodiments. In this embodiment, on the one hand, by pre - constructing word sets for normal type applications and abnormal type applications, it is possible to better represent the interface call combinations of various types of applications based on the candidate words composed of multiple interface codes in the word sets, so that the characteristics of abnormal code API calls can be extracted more clearly and effectively. On the other hand, by segmenting the actual interface call sequence of the target application into different segments, it is easier to identify the API call positions of abnormal codes, accurately locate and analyze abnormal codes, and improve the detection accuracy and detection efficiency of abnormal applications of intelligent measurement terminals.

[0142] To enable those skilled in the art to better understand the above steps, the following gives an exemplary illustration of the embodiments of the present application through an example, but it should be understood that the embodiments of the present application are not limited thereto.

[0143] In the related art, when analyzing the interfaces called during the operation of an application using the lexical - level semantic analysis technology in natural language processing, first, the call sequence of the API is statically extracted, and each API is embedded and encoded as a word to obtain the word embedding vector of the API. Then, the word embedding vectors of the API are clustered, and the application program is converted into the encoded clusters. Finally, the abnormal application is identified through the cluster codes included in the application program and the integrated model.

[0144] However, the above methods have the following deficiencies: (1) Embedding and encoding the API as a single word results in relatively weak representation ability for the combination of API call sequences; (2) The time interval between API calls is not considered during word embedding. That is, when the time interval between two adjacent API calls is too long, the likelihood of their relevance is very small. Using each other as context is equivalent to adding noise during word embedding, affecting the effect of word embedding; (3) When malicious code inserts irrelevant calls between API calls, the detection ability of the method proposed in the paper is not strong; (4) In the paper, the word embedding vectors of APIs are clustered, and then whether malicious code exists is judged by the clustering codes included in the application. The ability to extract the combined features of API calls of malicious code is not strong, which is not conducive to the positioning of malicious code and the analysis of its API call characteristics.

[0145] In response, the present application provides an application detection method for an intelligent measurement terminal. As Figure 3 shown, the method may include the following steps:

[0146] S301, Establish a list of APIs of the intelligent measurement terminal software platform and encode each API.

[0147] S302, Run the application program with a marked type for training the initial model in a sandbox, and record the time intervals of the encoded calls of the APIs of the software platform called during the running of the application program in sequence to form an API call sequence.

[0148] S303, Treat each API as a character and use a model (such as the model of the word2vc framework) to train the character embedding vector of the API.

[0149] S304, Construct a word library for API calls based on the character embedding vectors of APIs.

[0150] S305, Use the word library to segment the API call sequences of the application programs in the training set.

[0151] S306, Segment all the API call sequences of the application programs for training. The words used by normal application programs form a normal word set. The words used by abnormal application programs that are removed from the normal word set form an abnormal word set corresponding to each type of abnormal application program.

[0152] S307, For a new application program to be tested (i.e., the target application), first segment its API call sequence, and check whether all the words it uses are in the normal word set. If so, it is judged as a normal application program. Otherwise, check whether the words outside the normal word set are included in the abnormal word sets of known types. If so, it is judged as the above abnormal type. If there are words not included in all word sets, it is judged as an unknown type.

[0153] Compared with the related art, the present application has the following beneficial effects:

[0154] (1) In the present application, APIs are regarded as words, and the call combinations of APIs are regarded as a word. By constructing a word library without supervision, API call sequences with strong relevance are combined into words, which has a stronger ability to represent the combination of API call orders and can well handle the problems of too long call time intervals and the insertion of irrelevant APIs.

[0155] (2) Through the constructed word library, the present application can more clearly and effectively extract the characteristics of malicious code API calls.

[0156] (3) By segmenting the API call sequences of the application program through word segmentation, the present application can more easily locate and analyze malicious code.

[0157] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0158] Based on the same inventive concept, the embodiments of the present application also provide a word set acquisition device for intelligent measurement terminal application detection for implementing the word set acquisition method for intelligent measurement terminal application detection described above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the word set acquisition device for intelligent measurement terminal application detection provided below can refer to the limitations on the word set acquisition method for intelligent measurement terminal application detection in the above text, and will not be repeated here.

[0159] In an exemplary embodiment, as Figure 4 shown, a word set acquisition device for intelligent measurement terminal application detection is provided, including:

[0160] A sample acquisition module 401, configured to acquire sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include interface encodings and interval encodings corresponding to each platform interface called in sequence during the operation of the sample application; the interval encoding is inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold.

[0161] A word embedding acquisition module 402, configured to obtain a first interface code sequence from the sample interface call sequence according to the interval encoding, the first truncation interval with a preset length, and the first interface code as the interval center in the sample interface call sequence, and perform self-supervised training on the word embedding vector acquisition model according to a plurality of the first interface code sequences. At the end of the training, determine the word embedding vectors corresponding to each of the interface codes according to the word embedding vector acquisition model; the word embedding vector acquisition model is configured to determine word embedding vectors according to each interface code in the first interface code sequence, and predict the first interface code according to the word embedding vectors;

[0162] A word set acquisition module 403, configured to obtain candidate words including at least two interface codes according to each of the sample interface call sequences, determine the coding correlation degrees of the interface codes included in each of the candidate words according to the word embedding vectors corresponding to each of the interface codes, and construct word sets of the multiple application types according to the candidate words that satisfy a preset condition; the word sets are used to detect the application types of intelligent measurement terminal applications.

[0163] In one embodiment, the word set acquisition module 403 is configured to:

[0164] For each of the candidate words, obtain the adjacency entropy corresponding to the candidate word, determine the overall tightness of the candidate word according to the word embedding vectors corresponding to the interface codes included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word, and determine the coding correlation degrees of the interface codes included in the candidate word according to the adjacency entropy and the overall tightness of the candidate word;

[0165] Add each candidate word with a coding correlation degree greater than a preset threshold to a word library;

[0166] Perform word segmentation on the sample interface call sequences under each of the application types according to the word library, and obtain word sets of the multiple application types according to the word segmentation results.

[0167] In one embodiment, the word set acquisition module 403 is configured to:

[0168] Divide the candidate word to obtain multiple groups of front-segment interface code sequences and back-segment interface code sequences;

[0169] For each group of the front-segment interface code sequences and back-segment interface code sequences, determine a splitting probability according to the occurrence probabilities of the candidate word, the front-segment interface code sequence, and the back-segment interface code sequence, and determine the semantic similarity between the front-segment interface code sequence and the back-segment interface code sequence according to the word embedding vectors corresponding to the interface codes included in the candidate word;

[0170] Determine the overall tightness of the candidate word according to the splitting probability and the semantic similarity corresponding to multiple groups of the front-segment interface coding sequence and the rear-segment interface coding sequence.

[0171] In one embodiment, the word set acquisition module 403 is configured to:

[0172] Reverse-obtain a second interface coding sequence from the current sample interface call sequence according to the interval coding in the sample interface call sequence, the second intercepting interval with a preset length, and the second interface coding as the center of the interval;

[0173] Determine whether the word library includes a matching word for the second interface coding sequence; the matching word is an interface coding combination including the first interface coding and the last interface coding in the current second interface coding sequence;

[0174] If so, use the current second interface coding sequence as the word segmentation result to segment from the current sample interface call sequence, add it to the word set corresponding to the application type, and if the current sample interface call sequence is not completely segmented, return to execute the step of obtaining the second interface coding sequence from the current sample interface call sequence;

[0175] If not, remove the first interface coding in the current second interface coding sequence, and return to execute the step of determining whether the word library includes a matching word.

[0176] In one embodiment, the word embedding acquisition module 402 is configured to:

[0177] In the sample interface call sequence, perform sequence interception with the first interface coding as the center of the interval and a first interception interval with a preset length of 2N+1 to obtain an intercepted sequence; N is a positive integer;

[0178] When the interval coding exists among the N interface codings before the first interface coding, set each interface coding before the interval coding in the intercepted sequence to the interval coding, and when the interval coding exists among the N interface codings after the first interface coding, set each interface coding after the interval coding in the intercepted sequence to the interval coding;

[0179] Obtain a first interface coding sequence according to the intercepted sequence after setting the interval coding.

[0180] Based on the same inventive concept, an embodiment of the present application further provides a word set-based intelligent measurement terminal application detection device for implementing the above-mentioned word set-based intelligent measurement terminal application detection method. The implementation solution provided by this device for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the word set-based intelligent measurement terminal application detection device provided below can refer to the limitations on the word set-based intelligent measurement terminal application detection method in the above text, and will not be elaborated here.

[0181] In an exemplary embodiment, as Figure 5 shown, a word set-based intelligent measurement terminal application detection device is provided, including:

[0182] An actual sequence acquisition module 501, configured to acquire an actual interface call sequence corresponding to a target application to be detected by the intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to multiple platform interfaces sequentially called during the operation of the target application;

[0183] A word segmentation module 502, configured to segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code;

[0184] A first recognition module 503, configured to determine that the target application is the normal type application if each word segment of the actual interface call sequence is included in the word set of the normal type application;

[0185] A second recognition module 504, configured to determine that the target application is the abnormal type application if some word segments of the actual interface call sequence are included in the word set of the abnormal type application, and determine the abnormal code in the target application according to the interface call codes corresponding to each interface code in some word segments;

[0186] Wherein, each word set is obtained according to the word set acquisition method for intelligent measurement terminal application detection described in any one of the above.

[0187] Each module in the above-mentioned word set acquisition device for intelligent measurement terminal application detection and the word set-based intelligent measurement terminal application detection device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0188] In an exemplary embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store interface call sequence data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for obtaining a word set for intelligent metering terminal application detection or a method for intelligent metering terminal application detection based on a word set.

[0189] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 7 shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC), or other technologies. When the computer program is executed by the processor, it implements a method for obtaining a word set for intelligent metering terminal application detection or a method for intelligent metering terminal application detection based on a word set. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0190] Those skilled in the art can understand thatFigure 6 and Figure 7 The structures shown in Figure 7 are merely block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0191] In one embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0192] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0193] In one embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0194] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0195] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0196] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.

[0197] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for obtaining a word set for intelligent measurement terminal application detection, characterized in that The method includes: Obtaining the sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include the interface codes and interval codes corresponding to each platform interface called in sequence during the operation of the sample applications; the interval codes are inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold; According to the interval codes, the first truncation interval with a preset length, and the first interface code as the center of the interval in the sample interface call sequence, obtaining a first interface code sequence from the sample interface call sequence, and performing self-supervised training on the word embedding vector acquisition model according to multiple first interface code sequences. At the end of the training, determining the word embedding vectors corresponding to each interface code according to the word embedding vector acquisition model; the word embedding vector acquisition model is used to determine the word embedding vectors according to each interface code in the first interface code sequence and predict the first interface code according to the word embedding vectors; According to each sample interface call sequence, obtaining candidate words containing at least two interface codes, determining the coding correlation degrees of the interface codes included in each candidate word according to the word embedding vectors corresponding to each interface code, and constructing word sets for multiple application types according to the candidate words whose coding correlation degrees meet the preset conditions; the word sets are used to detect the application types of intelligent metering terminal applications.

2. The method according to claim 1, wherein The determining the coding correlation degrees of the interface codes included in each candidate word according to the word embedding vectors corresponding to each interface code, and constructing word sets for multiple application types according to the candidate words whose coding correlation degrees meet the preset conditions includes: For each candidate word, obtaining the adjacency entropy corresponding to the candidate word, determining the overall tightness degree of the candidate word according to the word embedding vectors corresponding to each interface code included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word, and determining the coding correlation degrees of the interface codes included in the candidate word according to the adjacency entropy and the overall tightness degree of the candidate word; Adding each candidate word with a coding correlation degree greater than a preset threshold to a word library; Performing word segmentation on the sample interface call sequences of each application type according to the word library, and obtaining the word sets of each application type according to the word segmentation results.

3. The method according to claim 2, wherein The determining the overall tightness degree of the candidate word according to the word embedding vectors corresponding to each interface code included in the candidate word and the occurrence probability of the interface code sequence corresponding to the candidate word includes: Dividing the candidate word to obtain multiple groups of front-segment interface code sequences and back-segment interface code sequences; For each group of the front-segment interface code sequence and the back-segment interface code sequence, determining the splitting probability according to the occurrence probabilities of the candidate word, the front-segment interface code sequence, and the back-segment interface code sequence respectively, and determining the semantic similarity between the front-segment interface code sequence and the back-segment interface code sequence according to the word embedding vectors corresponding to each interface code included in the candidate word; Determine the overall tightness of the candidate word according to the splitting probability and the semantic similarity corresponding to multiple groups of the front-end interface coding sequences and the back-end interface coding sequences.

4. The method according to claim 2, wherein Performing word segmentation on the sample interface call sequences of each of the application types according to the word library, and obtaining the word sets of each of the application types according to the word segmentation results, including: Obtain a second interface coding sequence from the current sample interface call sequence in reverse according to the interval coding, the second truncation interval with a preset length, and the second interface coding as the center of the interval in the sample interface call sequence; Determine whether the word library includes a matching word for the second interface coding sequence; the matching word is an interface coding combination including the first interface coding and the last interface coding in the current second interface coding sequence; If so, use the current second interface coding sequence as the word segmentation result to segment from the current sample interface call sequence, add it to the word set of the corresponding application type, and when the current sample interface call sequence is not completely segmented, return to execute the step of obtaining the second interface coding sequence from the current sample interface call sequence; If not, remove the first interface coding in the current second interface coding sequence, and return to execute the step of determining whether the word library includes a matching word.

5. The method according to any one of claims 1 to 4, characterized in that The obtaining of the first interface coding sequence from the sample interface call sequence according to the interval coding, the first truncation interval with a preset length, and the first interface coding as the center of the interval in the sample interface call sequence includes: In the sample interface call sequence, perform sequence truncation with the first interface coding as the center of the interval and a first truncation interval with a preset length of 2N + 1 to obtain a truncated sequence; N is a positive integer; When the interval coding exists in the N interface codes before the first interface coding, set each interface code before the interval coding in the truncated sequence to the interval coding, and when the interval coding exists in the N interface codes after the first interface coding, set each interface code after the interval coding in the truncated sequence to the interval coding; Obtain the first interface coding sequence according to the truncated sequence after setting the interval coding.

6. An intelligent measurement terminal application detection method based on a word set, characterized in that, The method includes: Obtain the actual interface call sequence corresponding to the target application to be detected by the intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to multiple platform interfaces called in sequence during the operation of the target application; Segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code; If each of the word segments of the actual interface call sequence is included in the word set of the normal type application, determine that the target application is the normal type application; If some of the word segments of the actual interface call sequence are included in the word set of the abnormal type application, determine that the target application is the abnormal type application, and determine the abnormal code in the target application according to the interface call codes corresponding to the interface codes in some of the word segments. Among them, each of the above-mentioned word sets is obtained according to the word set acquisition method for intelligent measurement terminal application detection described in any one of claims 1 to 5.

7. A word set acquisition device for intelligent measurement terminal application detection, characterized in that, The device includes: A sample acquisition module, configured to acquire sample interface call sequences of sample applications under multiple application types; the sample interface call sequences include interface codes and interval codes corresponding to each platform interface called in sequence during the operation of the sample applications; the interval codes are inserted when the call time interval between two adjacent platform interfaces is greater than a preset threshold; A word embedding acquisition module, configured to obtain a first interface code sequence from the sample interface call sequences according to the interval codes, a first truncation interval with a preset length, and a first interface code as the center of the interval in the sample interface call sequences, and perform self-supervised training on a word embedding vector acquisition model according to multiple first interface code sequences. When the training ends, determine word embedding vectors corresponding to each of the interface codes according to the word embedding vector acquisition model; the word embedding vector acquisition model is configured to determine word embedding vectors according to each interface code in the first interface code sequence and predict the first interface code according to the word embedding vectors; A word set acquisition module, configured to acquire candidate words including at least two interface codes according to each of the sample interface call sequences, determine the coding correlation degrees of the interface codes included in each of the candidate words according to the word embedding vectors corresponding to each of the interface codes, and construct word sets of multiple application types according to the candidate words that meet the preset conditions; the word sets are used to detect the application types of intelligent measurement terminal applications.

8. An intelligent measurement terminal application detection device based on a word set, characterized in that The device includes: An actual sequence acquisition module, configured to acquire an actual interface call sequence corresponding to a target application to be detected by an intelligent measurement terminal; the actual interface call sequence includes interface codes corresponding to multiple platform interfaces called in sequence during the operation of the target application; A word segmentation module, configured to segment the actual interface call sequence to obtain corresponding word segments; each word segment consists of at least one interface code; A first recognition module, configured to determine that the target application is the normal type application if each of the word segments of the actual interface call sequence is included in the word set of the normal type application; A second recognition module, configured to determine that the target application is the abnormal type application if some of the word segments of the actual interface call sequence are included in the word set of the abnormal type application, and determine the abnormal codes in the target application according to the interface call codes corresponding to the interface codes in some of the word segments; Among them, each of the above-mentioned word sets is obtained according to the word set acquisition method for intelligent measurement terminal application detection described in any one of claims 1 to 5.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5 or the steps of the method described in claim 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 5 or the steps of the method described in claim 6.

Citation Information

Patent Citations

  • Entity word recognition method and device, equipment, storage medium and program product

    CN113656561A

  • Abnormal prediction method and device, storage medium and electronic equipment

    CN116680141A

  • User abnormal behavior detection method and device, equipment and storage medium

    CN117688434A

  • Malicious software analysis and identification method and device

    CN120068073A

  • Text summarization generation method and apparatus, and device and storage medium

    WO2022241950A1