An application detection method, device, equipment and storage medium

By using a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting, and combining preset splicing structures to generate an application overview diagram, the problem of low efficiency and accuracy of existing application detection technologies is solved, and more efficient application detection is achieved.

CN114119365BActive Publication Date: 2025-06-10EVERSEC BEIJING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111328570.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-06-10
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The existing application detection technology lacks holistic considerations when detecting application content, and the deep learning model is single, resulting in low efficiency and accuracy.

Method used

Using a Transformer model based on inter-block feature aggregation and shallow feature short-circuit, we use screenshots of key applications of the application to be tested and spliced ​​according to the preset splicing structure to generate an application overview diagram, and then input it into the trained application detection model for detection.

Benefits of technology

Improve the efficiency and accuracy of application detection, and can more effectively determine whether the application contains specific information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114119365B_ABST
    Figure CN114119365B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an application detection method, device, equipment and storage medium. The method includes: obtaining key application screenshots of the application to be detected, and splicing the key application screenshots according to a preset splicing structure to generate an application overview map; inputting the application overview map into a trained application detection model, and returning the application detection result output by the application detection model; wherein, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short connection. The technical solution of the embodiment of the present invention improves the efficiency and accuracy of application detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of machine learning, and in particular, to an application detection method, device, equipment and storage medium. Background Art

[0002] The purpose of application detection is to effectively detect whether there is specific information in a mobile application. For example, it is to detect whether the application contains gambling content. Currently, application detection technologies can be divided into three categories. One is specific element detection based on traditional computer vision. The second is specific content picture classification or specific element detection based on the convolutional neural network (CNN) model of deep learning. The third is to perform text detection and recognition on the application screenshot based on the optical character recognition (OCR) technology, and judge whether the application is an application including specific content according to the page text information.

[0003] However, when the existing application detection technologies judge the application content, they do not consider the integrity of the application, and the deep learning models used are also relatively single. Therefore, the efficiency and accuracy of application detection need to be improved. Summary of the Invention

[0004] Embodiments of the present invention provide an application detection method, device, equipment and storage medium to improve the efficiency and accuracy of application detection.

[0005] In a first aspect, an embodiment of the present invention provides an application detection method, including:

[0006] Obtain key application screenshots of the application to be detected, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram;

[0007] Input the application overview diagram into the trained application detection model, and return the application detection result output by the application detection model;

[0008] Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting.

[0009] Optionally, obtaining key application screenshots of the application to be detected, and splicing the key application screenshots according to a preset splicing structure to generate an application overview diagram includes:

[0010] Obtain an application screenshot sequence of the application to be detected, and obtain the image gray-scale features of each application screenshot in the application screenshot sequence;

[0011] Cluster the application screenshot sequence into an application loading picture cluster and an application content picture cluster according to the image gray-scale features;

[0012] Select a preset number of key application screenshots from the core area of the picture cluster loaded by the application and the core area of the application content picture cluster respectively;

[0013] Stitch the key application screenshots according to a preset stitching structure to generate an application overview diagram.

[0014] Optionally, input the application overview diagram into a trained application detection model, and return the application detection results output by the application detection model, including:

[0015] Input the application overview diagram into a trained application detection model. Through the application detection model, layer-by-layer segmentation of the application overview diagram is performed in a four-equal-block manner to determine the bottom-layer image blocks;

[0016] Use the word vectors obtained by linearly mapping each bottom-layer image block as the current layer input elements, map the attention matrix for every four current layer input elements, and calculate the self-attention features of each image block in the current layer;

[0017] Perform inter-block feature aggregation processing on the self-attention features of every four image blocks in the current layer to obtain the feature information of each image block in the upper layer;

[0018] Use the feature information of each image block in the upper layer as the current layer input elements, return to perform the operation of mapping the attention matrix for every four current layer input elements and calculating the self-attention features of each image block in the current layer until the self-attention features of the complete application overview diagram are obtained;

[0019] Merge the self-attention features of the downsampled bottom-layer image blocks with the self-attention features of the application overview diagram, make a decision based on the merged self-attention features, obtain the application detection results and return them.

[0020] Optionally, map the attention matrix for every four current layer input elements and calculate the self-attention features of each image block in the current layer, including:

[0021] For every four current layer input elements corresponding to the same image block, calculate the matching query matrix Q, key matrix K, and value matrix V;

[0022] Use a non-linear function to perform position mapping on the four image sub-blocks included in each current layer image block to obtain an intra-block position relationship matrix P;

[0023] According to the formula Z = (Q * K T + P) * V, calculate the self-attention feature Z of each image block in the current layer;

[0024] Among them, K T represents the transpose of the key matrix K.

[0025] Optionally, a non-linear function is used to perform position mapping on the four image sub-blocks included in each current layer image block to obtain an intra-block position relationship matrix, including:

[0026] Generate a pixel position coordinate matrix corresponding to each current layer image block;

[0027] Among the matrix elements of the pixel position coordinate matrix, the x coordinate represents the image sub-block hierarchical coordinate, the y coordinate represents the pixel hierarchical coordinate within the image sub-block, and the matrix elements at the main diagonal position of each level have a coordinate of 0 at the corresponding level;

[0028] Synchronously adjust each matrix element in the pixel position coordinate matrix to a non-negative value, perform hashing processing on the x coordinate, and sum the processed matrix elements for the x coordinate and the y coordinate;

[0029] Use a non-linear function to map the pixel position coordinate matrix after the summation process to an intra-block position relationship matrix.

[0030] Optionally, after inputting the application overview map into the trained application detection model and returning the application detection result output by the application detection model, it further includes:

[0031] In response to the detection backtracking request from the application detection request side, if the application to be tested is a target type application, then return the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request side for display;

[0032] If the application to be tested is a non-target type application, then return the application overview map.

[0033] Optionally, returning the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request side for display includes:

[0034] Obtain the image self-attention feature used when the application detection model makes a judgment, and divide the image self-attention feature according to the multi-layer image blocks of the application overview map;

[0035] Use the image self-attention feature corresponding to the topmost layer image block as the current feature, calculate the score value corresponding to each current feature, and select the image block corresponding to the maximum score value as the current layer active block;

[0036] Use the image self-attention feature of the next layer image block corresponding to the current layer active block as the current feature, return to execute the operation of calculating the score value corresponding to each current feature, and select the image block corresponding to the maximum score value as the current layer active block until the bottom layer active block is determined;

[0037] Generate decision trees corresponding to the active blocks of each layer, and return the decision trees and the application overview diagram marked with the active blocks of each layer to the application detection request end for display.

[0038] In a second aspect, an embodiment of the present invention further provides an application detection device, including:

[0039] A picture splicing module, configured to obtain key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram;

[0040] An application detection module, configured to input the application overview diagram into a trained application detection model, and return the application detection result output by the application detection model;

[0041] Wherein, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting.

[0042] In a third aspect, an embodiment of the present invention further provides a computer device, the device includes:

[0043] One or more processors;

[0044] A storage device, configured to store one or more programs,

[0045] When the one or more programs are executed by the one or more processors, the one or more processors implement the application detection method provided in any embodiment of the present invention.

[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the application detection method provided in any embodiment of the present invention.

[0047] The technical solution of the embodiment of the present invention obtains key application screenshots of the application to be tested, splices the key application screenshots according to a preset splicing structure to generate an application overview diagram; inputs the application overview diagram into a trained application detection model, and returns the application detection result output by the application detection model; wherein, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting, which solves the problems of low efficiency and accuracy in application detection in the prior art, and improves the efficiency and accuracy of application detection. Description of the Drawings

[0048] Figure 1a is a flowchart of an application detection method in Embodiment 1 of the present invention;

[0049] Figure 1b is a splicing example diagram of an application overview diagram in Embodiment 1 of the present invention;

[0050] Figure 1c It is a schematic diagram of feature extraction of a Transformer model in Embodiment 1 of the present invention;

[0051] Figure 2a It is a flowchart of an application detection method in Embodiment 2 of the present invention;

[0052] Figure 2b It is a flowchart of the implementation of an application detection in Embodiment 2 of the present invention;

[0053] Figure 2c It is a schematic diagram of the generation of an intra-block positional relationship matrix in Embodiment 2 of the present invention;

[0054] Figure 2d It is an example diagram of a non-linear mapping function of position information in Embodiment 2 of the present invention;

[0055] Figure 2e It is a schematic diagram of an active block and a decision tree of an application overview diagram in Embodiment 2 of the present invention;

[0056] Figure 3 It is a schematic diagram of the structure of an application detection device in Embodiment 3 of the present invention;

[0057] Figure 4 It is a schematic diagram of the structure of a computer device in Embodiment 4 of the present invention. Detailed implementation manners

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.

[0059] Embodiment 1

[0060] Figure 1a It is a flowchart of an application detection method in Embodiment 1 of the present invention. This embodiment is applicable to the situation of detecting whether a specific information is included in an application. This method can be executed by an application detection device, which can be implemented by hardware and / or software and is generally integrated in a computer device providing application detection services. As Figure 1a shown, this method includes:

[0061] Step 110, obtain key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram.

[0062] In this embodiment, the application to be tested may be an application that requests to detect whether it contains specific information, where the specific information may be gambling information, loan information, etc. The key application screenshots may include multiple application screenshots that can reflect the overall content of the application to be tested. For example, screenshots showing the application loading process and screenshots showing the application usage process after loading are included. The preset splicing structure is specifically designed for the obtained key application screenshots in order to balance the overall integrity of the application, the accuracy of application detection, and the efficiency of application detection.

[0063] In this embodiment, when detecting specific information in the application to be tested, first, according to the preset splicing structure, a specified number of key application screenshots can be selected from the screenshots of the application to be tested. Exemplarily, as Figure 1b shown, if the preset splicing structure requires 2 application loading pictures and 3 application content pictures, then, in combination with the picture similarity, 2 application loading pictures and 3 application content pictures can be selected from all the application screenshots of the application to be tested, and the pictures are spliced according to the preset splicing structure to obtain an application overview diagram as Figure 1b shown.

[0064] Optionally, obtaining the key application screenshots of the application to be tested and splicing the key application screenshots according to the preset splicing structure to generate an application overview diagram may include: obtaining the application screenshot sequence of the application to be tested and obtaining the image gray-scale features of each application screenshot in the application screenshot sequence; clustering the application screenshot sequence into an application loading picture cluster and an application content picture cluster according to the image gray-scale features; respectively selecting a preset number of key application screenshots from the core areas of the application loading picture cluster and the application content picture cluster; and splicing each key application screenshot according to the preset splicing structure to generate an application overview diagram.

[0065] In this embodiment, the Android environment can be used to perform operations such as clicking on the application to be tested using a simulated sandbox, and a preset number of application pictures are intercepted at fixed intervals. For example, 200 application pictures are intercepted to obtain the application screenshot sequence of the application to be tested. The size of each application screenshot is reduced and converted into a 256-dimensional gray-scale histogram, and the image gray-scale features are obtained from each gray-scale histogram. Through a clustering algorithm, the application screenshot sequence is clustered into two clusters using the image gray-scale features. The picture time of the application screenshots located in the core area of each cluster is obtained, and the cluster with the earlier time is used as the application loading picture cluster, and the cluster with the later time is used as the application content picture cluster. Then, a deduplication operation is performed on all the application screenshots, and according to the requirements of the preset splicing structure for the application screenshots, the corresponding number of application screenshots are respectively selected from the core areas of the application loading picture cluster and the application content picture cluster. For example, 2 are extracted from the application loading picture cluster and 3 are extracted from the application content picture cluster as the key application screenshots that reflect the overall integrity of the application. The key application screenshots are spliced according to the preset splicing structure to obtain an application overview diagram.

[0066] Among them, a number of levels are divided logarithmically between white and black, called grayscale, and the grayscale is divided into 256 levels. The grayscale histogram is a function of the grayscale level, representing the number of pixels with each grayscale level in the image, and reflecting the frequency of each grayscale appearance in the image.

[0067] In this embodiment, by setting a preset splicing structure, only one picture is needed to reflect the overall situation of the application to be tested, avoiding the situation where an application detection process is carried out separately for all application screenshots of the application to be tested, and only after comprehensively combining the detection results of all application screenshots can it be determined whether the application to be tested includes specific information, thus improving the application detection efficiency. And, using the Figure 1b splicing structure as shown, including an application content screenshot in the middle of the application overview diagram, can facilitate obtaining mixed information from the application overview diagram when using the application detection model for application detection later, improving the accuracy of the image self-attention features required for the final decision, and thus improving the accuracy of application detection.

[0068] Step 120: Input the application overview diagram into the trained application detection model, and return the application detection result output by the application detection model.

[0069] Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting. The Transformer model consists of an encoding component, a decoding component, and the connection between them, and adopts the Self-Attention mechanism, enabling the model to be trained in parallel and having global information.

[0070] In this embodiment, in order to better adapt to the scenario of using application screenshots for application detection, a Transformer model is used as the application detection model, and when training the Transformer model, the shallow block attention features are short-circuited and mapped to the deep features obtained through inter-block feature aggregation, effectively improving the sensitivity of the model to image details, and further improving the accuracy of application detection.

[0071] Optionally, input the application overview map into the trained application detection model, and return the application detection result output by the application detection model, which may include: input the application overview map into the trained application detection model, and through the application detection model, use the four-equal-block method to layer-by-layer segment the application overview map to determine the bottom-layer image blocks; use the word vectors obtained by linearly mapping each bottom-layer image block as the current-layer input elements, map the attention matrix for every four current-layer input elements, and calculate the self-attention features of each image block in the current layer; perform inter-block feature aggregation processing on the self-attention features of every four image blocks in the current layer to obtain the feature information of each image block in the upper layer; use the feature information of each image block in the upper layer as the current-layer input elements, and return to perform the operation of mapping the attention matrix for every four current-layer input elements and calculating the self-attention features of each image block in the current layer until the self-attention features of the complete application overview map are obtained; merge the self-attention features of the downsampled bottom-layer image blocks with the self-attention features of the application overview map, make a decision based on the merged self-attention features, obtain the application detection result and return it.

[0072] In this embodiment, input the application overview map as shown in Figure 1b into the trained application detection model, that is, the Transformer model. Divide the application overview map with a size of W*H according to the specified number of layers. Each layer evenly divides the traversed image blocks into four pieces, and finally determines the bottom-layer image blocks with the same size. Assume that the size of the bottom-layer image block is Sw*Sh, then the number of finally obtained bottom-layer image blocks is Bn = (W*H) / (Sw*Sh). Then as shown in Figure 1cAs shown, block - level self - attention feature aggregation is performed from the bottom layer to the top layer. First, the bottom layer is regarded as the current layer. Each bottom - layer image block is linearly mapped into a word vector of length d. All word vectors are sliced into blocks and flattened to generate the model input x corresponding to the application overview map. x is a vector of dimension 1 * Bn * d. Each word vector in x is used as an input element of the current layer, and an attention matrix is mapped for every four input elements of the current layer. The self - attention features of each image block in the current layer are calculated according to the attention matrix. Then, feature integration is performed. The self - attention features of all image blocks in the current layer are recorded as Ym, where m represents the level of the image blocks in the current layer, and the level corresponding to the bottom layer is 1. Every four self - attention features are grouped together, and through two - layer 3 * 3 convolution and max - pooling processing, inter - block feature aggregation is achieved, and the feature information of the image blocks in the upper layer corresponding to this feature group is refined. It is judged whether the self - attention features of the complete application overview map of the top layer have been obtained. If not, the feature information of each image block in the upper layer is used as the input element of the current layer, and the operation of mapping the attention matrix for every four input elements of the current layer and calculating the self - attention features of each image block in the current layer is returned until the self - attention features of the complete application overview map are obtained. At this time, the self - attention features Y1 of the bottom - layer image blocks need to be downsampled and short - connected to the self - attention features of the application overview map to obtain the final image self - attention features. This can ensure the effective acquisition of shallow - layer information, which is beneficial for the application detection model to judge whether the application under test contains specific information, such as gambling information, by learning key icon information, such as the Double - Color Ball icon and the Big Lotto icon related to specific gambling information.

[0073] In this embodiment, during the process of self - attention feature aggregation of image blocks, after the self - attention features of the bottom layer are downsampled and merged with the self - attention features of the whole picture, the sensitivity of the application detection model to the details of the application overview map can be effectively improved, making the activation degree of icon information in the application detection model higher.

[0074] The technical solution of the embodiment of the present invention is as follows: by obtaining the key application screenshots of the application under test and splicing the key application screenshots according to a preset splicing structure to generate an application overview map; inputting the application overview map into a trained application detection model, and returning the application detection result output by the application detection model; where the application detection model is a Transformer model based on inter - block feature aggregation and shallow - layer feature short - connection, which solves the problems of low efficiency and accuracy in application detection in the prior art and improves the efficiency and accuracy of application detection.

[0075] Embodiment Two

[0076] Figure 2aIt is a flowchart of an application detection method in the second embodiment of the present invention. This embodiment is further refined on the basis of the above embodiment, providing specific steps for mapping the attention matrix for every four input elements of the current layer, calculating the self-attention features of each image block of the current layer, and, after inputting the application overview map into the trained application detection model and returning the application detection result output by the application detection model, providing specific steps for the judgment basis in response to the detection backtracking request of the application detection request end. The following combines Figure 2a to illustrate an application detection method provided in this embodiment, including the following steps:

[0077] Step 210: Obtain the key application screenshots of the application to be detected, and splice the key application screenshots according to a preset splicing structure to generate an application overview map.

[0078] In this embodiment, in order to enable the Transformer model to better adapt to the scenario of application detection using application screenshots, before obtaining the key application screenshots of the application to be detected and splicing the key application screenshots according to a preset splicing structure to generate an application overview map, the Transformer model can be trained to short-circuit and map the shallow block attention features to the deep features obtained through inter-block feature aggregation, so as to effectively improve the sensitivity of the model to image details and further improve the application detection accuracy.

[0079] As Figure 2b shown, first obtain the training application sample set, obtain the application screenshot sequences of each application sample, select the key application screenshots from the application screenshot sequences, and splice them to generate an application overview map. Among them, the application sample set includes target type application samples containing specific information and non-target type application samples not containing specific information. Input the application overview maps of each application sample into the transformer model for training to adjust the model parameters to the best, so that the efficiency and accuracy of the model for application detection reach the best.

[0080] Step 220: Input the application overview map into the trained application detection model, and use the four-equal-block method to layer-by-layer segment the application overview map through the application detection model to determine the bottom-layer image blocks.

[0081] Step 230: Use the word vectors obtained by linearly mapping each bottom-layer image block as the input elements of the current layer, map the attention matrix for every four input elements of the current layer, and calculate the self-attention features of each image block of the current layer.

[0082] Optionally, for every four current-layer input elements, map the attention matrix and calculate the self-attention features of each image patch in the current layer, which may include: for every four current-layer input elements corresponding to the same image patch, calculate the matching query matrix Q, key matrix K, and value matrix V; use a non-linear function to perform a position mapping on the four image sub-patches included in each current-layer image patch to obtain an intra-block position relationship matrix P; according to the formula Z = (Q * K T + P) * V, calculate the self-attention feature Z of each image patch in the current layer; where K T represents the transpose of the key matrix K.

[0083] In this embodiment, when calculating the self-attention feature of a current-layer image patch, the upper-layer image patch corresponding to the current-layer image patch may first be determined, and then the current-layer input elements corresponding to the four current-layer image patches included in the upper-layer image patch are taken as a group, multiplied by the corresponding weight matrix, and the query matrix Q, key matrix K, and value matrix V matching each current-layer input element are calculated. Then calculate the product Q * K between the matrix Q and the transpose of the matrix K T , determine the relationship between the image patches, and then use a non-linear function to perform a position mapping on the four image sub-patches included in each current-layer image patch to obtain an intra-block position relationship matrix P, add the matrix P to Q * K T , and then multiply by the matrix V to obtain the self-attention feature Z of each image patch in the current layer.

[0084] In this embodiment, during the calculation of the self-attention feature, by using a non-linear function to map the intra-block position information, the position information suitable for the Vision Transformer can be obtained without training.

[0085] Optionally, using a non-linear function to perform a position mapping on the four image sub-patches included in each current-layer image patch to obtain an intra-block position relationship matrix may include: generating a pixel position coordinate matrix corresponding to each current-layer image patch; where in each matrix element of the pixel position coordinate matrix, the x coordinate represents the hierarchical coordinate of the image sub-patch, the y coordinate represents the pixel-level coordinate within the image sub-patch, and the matrix elements at the main diagonal position of each level have a coordinate of 0 at the corresponding level; synchronously adjusting each matrix element in the pixel position coordinate matrix to a non-negative value, performing a hash processing on the x coordinate, and adding the processed matrix elements for the x coordinate and the y coordinate; using a non-linear function to map the pixel position coordinate matrix after the addition processing to an intra-block position relationship matrix.

[0086] In this embodiment, when using a non-linear function to map the in-block position information, a corresponding pixel position coordinate matrix can be generated for the pixels in the current layer image block. Each element in the matrix corresponds to a pixel in the current layer image block, and the position coordinates of each pixel correspond to the two-layer encoding of the image sub-block and the pixel. Among them, the x coordinate represents the image sub-block level coordinate, and the y coordinate represents the pixel level coordinate within the image sub-block. The matrix elements at the main diagonal position of each level have coordinates of 0 at the corresponding level. As Figure 2c shown in Table a in Figure 1b , Table a is equivalent to a 4*2 position matrix, and the two 2*2 blocks in it respectively correspond to Figure 1b the left two image sub-blocks of the first image block at the bottom layer in

[0087] , and there are 4 pixels in each image sub-block. For the image sub-block level coordinate, for the first 2*2 block of the matrix, since the corresponding image sub-block is on the main diagonal of the bottom layer image block, all the x coordinates within the block are 0. For the second 2*2 block of the matrix, since the corresponding image sub-block is at a position one to the left of the main diagonal in the same row of the bottom layer image block, all the x coordinates within the block are -1. If there is a 2*2 block in the matrix corresponding to Figure 2c the second image sub-block above the first image block at the bottom layer in Figure 2c , then for this image sub-block, since it is at a position one to the right of the main diagonal in the same row of the bottom layer image block, all the x coordinates within the block are +1. For the pixel level coordinate within the image sub-block, taking the first 2*2 block of the matrix as an example, since the first element and the fourth element are both on the main diagonal within the block, the y coordinate values of both are 0. For the second element of this block, since it is at a position one to the right of the diagonal in the same row at the element level, the y coordinate is 1. For the third element of this block, since it is at a position one to the left of the diagonal in the same row at the element level, the y coordinate is -1. Figure 2cTable d coordinates. Finally, since it is necessary to limit the range of the coordinate values of the image sub-block levels to avoid overly affecting the inter-block relationship information, the even function (e^x - 1) / (e^x + 1) is selected as the mapping basis and deformed according to business requirements into c’ = η*(e^((c - σ)*ω)) / (e^((c - σ)*ω)+1). According to this formula, the pixel position coordinate matrix of Table d is mapped to the intra-block position relationship matrix of Table e. Among them, σ represents the main diagonal offset value, ω represents the smoothing coefficient, and η represents the position intensity. The function image information is as Figure 2d shown.

[0088] Step 240: Perform inter-block feature aggregation processing on the self-attention features of every four image blocks in the current layer to obtain the feature information of each image block in the upper layer.

[0089] Step 250: Use the feature information of each image block in the upper layer as the input elements of the current layer, and return to perform the operation of mapping the attention matrix for every four current layer input elements and calculating the self-attention features of each image block in the current layer until the self-attention features of the complete application overview map are obtained.

[0090] Step 260: Merge the self-attention features of the downsampled bottom-layer image blocks with the self-attention features of the application overview map, and make a decision based on the merged self-attention features to obtain the application detection result and return it.

[0091] In this embodiment, it is also possible to trace the detection and recognition process of the application to be tested, display to the user the backtracking path of the active block that provides the greatest feature contribution, determine the position of the active block in the application overview map, and graphically display the recognition and detection process of the application.

[0092] Optionally, after inputting the application overview map into the trained application detection model and returning the application detection result output by the application detection model, it may further include: in response to the detection backtracking request of the application detection request end, if the application to be tested is a target type application, then return the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request end for display; if the application to be tested is a non-target type application, then return the application overview map.

[0093] In this embodiment, if the user clicks to view the judgment basis corresponding to the detection result of the application to be tested, it is determined whether the detection result of the application to be tested output by the judgment model is an application of the target type, for example, whether it is a gambling application. If so, the self-attention features of each layer of image patches in the detection process are extracted to draw a self-attention decision tree. At the same time, since the final self-attention features are aggregated in the order of image patches, the position of the active block in the application overview diagram can be traced back. The application overview diagram marked with the judgment active block and the decision tree are used as the judgment basis and returned to the application detection request end for display. If the application to be tested is a non-target type application, only the application overview diagram needs to be returned.

[0094] Optionally, using the application overview diagram marked with the judgment active block and the decision tree as the judgment basis and returning them to the application detection request end for display may include: obtaining the image self-attention features used when the application detection model makes a judgment, and dividing the image self-attention features according to the multi-layer image patches of the application overview diagram; using the image self-attention features corresponding to the top-layer image patches as the current features, calculating the score values corresponding to each current feature, and selecting the image patch corresponding to the maximum score value as the active block of the current layer; using the image self-attention features of the next-layer image patches corresponding to the active block of the current layer as the current features, returning to execute the operation of calculating the score values corresponding to each current feature, and selecting the image patch corresponding to the maximum score value as the active block of the current layer until the bottom-layer active block is determined; generating a decision tree corresponding to each layer of active blocks, and returning the decision tree and the application overview diagram marked with each layer of active blocks to the application detection request end for display.

[0095] In this embodiment, when obtaining the judgment basis, the image self-attention features used when the application detection model finally makes a judgment can be obtained, that is, the features after merging the self-attention features of the bottom-layer image patches after downsampling and the self-attention features of the application overview diagram. First, divide the image self-attention features into four blocks according to the image patches during application prediction, calculate the score values corresponding to the self-attention features of each block, and select the image patch corresponding to the maximum score value as the active block of the current layer. As Figure 2eAs shown, the score values corresponding to each self-attention feature are 0.49, 0.42, 0.44, and 0.46 respectively. Then the active block of the current layer is the image block corresponding to 0.49, that is, the upper left quarter image block in the application overview diagram. Then, the self-attention features of the active block of the current layer are further divided into four blocks according to the image blocks, and the score values corresponding to each self-attention feature are calculated. For example, the score values are 3.28, 0.34, 0.50, and 0.93. Then, continue to select the image block corresponding to 3.28 in the active block of the current layer as the active block, that is, the upper left sixteenth image block in the application overview diagram. And so on, until the bottom active block is determined and the decision path is obtained. Generate decision trees corresponding to the active blocks of each layer, and return the decision trees and the application overview diagrams marked with the active blocks of each layer to the application detection request side for display.

[0096] The technical solution of the embodiment of the present invention is to obtain the key application screenshots of the application to be tested, splice the key application screenshots according to a preset splicing structure to generate an application overview diagram, input the application overview diagram into a trained application detection model, and return the application detection result output by the application detection model. Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting, which solves the problems of low efficiency and accuracy in application detection in the prior art and improves the efficiency and accuracy of application detection.

[0097] Embodiment III

[0098] Figure 3 It is a structural schematic diagram of an application detection device in Embodiment III of the present invention. This embodiment is applicable to the situation of detecting whether a specific information is included in an application. The device can be implemented by hardware and / or software and is generally integrated in a computer device providing application detection services. As Figure 3 shown, the device includes:

[0099] A picture splicing module 310, configured to obtain the key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram;

[0100] An application detection module 320, configured to input the application overview diagram into a trained application detection model, and return the application detection result output by the application detection model;

[0101] Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting.

[0102] The technical solution of the embodiment of the present invention obtains key application screenshots of the application to be tested, splices the key application screenshots according to a preset splicing structure to generate an application overview diagram, inputs the application overview diagram into a trained application detection model, and returns the application detection result output by the application detection model. The application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting, which solves the problems of low efficiency and accuracy in application detection in the prior art and improves the efficiency and accuracy of application detection.

[0103] Optionally, the picture splicing module 310 is used for:

[0104] Obtain a sequence of application screenshots of the application to be tested, and obtain the image gray-scale features of each application screenshot in the sequence of application screenshots;

[0105] Cluster the sequence of application screenshots into an application loading picture cluster and an application content picture cluster according to the image gray-scale features;

[0106] Select a preset number of key application screenshots from the core area of the application loading picture cluster and the core area of the application content picture cluster respectively;

[0107] Splice each key application screenshot according to a preset splicing structure to generate an application overview diagram.

[0108] Optionally, the application detection module 320 includes:

[0109] An image segmentation unit for inputting the application overview diagram into a trained application detection model, and using the application detection model to layer-by-layer segment the application overview diagram in a four-equal-block manner to determine the bottom-layer image blocks;

[0110] A feature calculation unit for using the word vectors obtained by linearly mapping each bottom-layer image block as the current-layer input elements, mapping the attention matrix for every four current-layer input elements, and calculating the self-attention features of each image block in the current layer;

[0111] A feature aggregation unit for performing inter-block feature aggregation processing on the self-attention features of every four image blocks in the current layer to obtain the feature information of each image block in the upper layer;

[0112] A return execution unit for using the feature information of each image block in the upper layer as the current-layer input elements, returning and executing the operation of mapping the attention matrix for every four current-layer input elements and calculating the self-attention features of each image block in the current layer until the self-attention features of the complete application overview diagram are obtained;

[0113] A judgment unit, configured to merge the self-attention features of the downsampled underlying image blocks with the self-attention features of the application overview map, make a judgment based on the merged self-attention features, obtain an application detection result and return it.

[0114] Optionally, a feature calculation unit, configured to:

[0115] A matrix calculation sub-unit, configured to calculate a matching query matrix Q, key matrix K, and value matrix V for every four current layer input elements corresponding to the same image block;

[0116] A position mapping sub-unit, configured to use a non-linear function to perform position mapping on the four image sub-blocks included in each current layer image block to obtain an intra-block position relationship matrix P;

[0117] A feature calculation sub-unit, configured to calculate the self-attention feature Z of each image block in the current layer according to the formula Z = (Q * K T + P) * V;

[0118] where K T represents the transpose of the key matrix K.

[0119] Optionally, a position mapping sub-unit, configured to:

[0120] Generate a pixel position coordinate matrix corresponding to each current layer image block;

[0121] Among the matrix elements of the pixel position coordinate matrix, the x coordinate represents the image sub-block level coordinate, the y coordinate represents the pixel level coordinate within the image sub-block, and the matrix elements at the main diagonal position of each level have a coordinate of 0 at the corresponding level;

[0122] Synchronously adjust each matrix element in the pixel position coordinate matrix to a non-negative value, perform a hash process on the x coordinate, and sum the processed matrix elements for the x coordinate and y coordinate;

[0123] Use a non-linear function to map the pixel position coordinate matrix after the summation process to an intra-block position relationship matrix.

[0124] Optionally, it further includes: a detection backtracking module, configured to, after inputting the application overview map into the trained application detection model and returning the application detection result output by the application detection model,

[0125] In response to a detection backtracking request from the application detection request end, if the application to be detected is a target type application, then return the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request end for display;

[0126] If the application to be detected is a non-target type application, then return the application overview map.

[0127] Optionally, a detection backtracking module is used for:

[0128] Obtain the image self-attention features used when the application detection model makes a decision, and divide the image self-attention features according to the multi-layer image blocks of the application overview diagram;

[0129] Use the image self-attention features corresponding to the top-layer image blocks as the current features, calculate the score values corresponding to each current feature, and select the image block corresponding to the maximum score value as the active block of the current layer;

[0130] Use the image self-attention features of the next-layer image blocks corresponding to the active block of the current layer as the current features, return to execute the operation of calculating the score values corresponding to each current feature, and select the image block corresponding to the maximum score value as the active block of the current layer until the active block of the bottom layer is determined;

[0131] Generate decision trees corresponding to the active blocks of each layer, and return the decision trees and the application overview diagram marked with the active blocks of each layer to the application detection request side for display.

[0132] The application detection device provided by the embodiments of the present invention can execute the application detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0133] Embodiment 4

[0134] Figure 4 It is a schematic structural diagram of a computer device in Embodiment 4 of the present invention. Figure 4 It shows a block diagram of an exemplary device 12 suitable for implementing the embodiments of the present invention. Figure 4 The shown device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0135] As Figure 4 shown, the device 12 is presented in the form of a general-purpose computing device. The components of the device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0136] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0137] Device 12 typically includes a variety of computer system readable media. These media can be any available media accessible to device 12, including volatile and non-volatile media, removable and non-removable media.

[0138] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. Memory 28 can include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0139] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.

[0140] Device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the device 12, and / or communicate with any device that enables the device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, network adapter 20 communicates with other modules of device 12 through bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0141] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, for example, implementing an application detection method provided by an embodiment of the present invention, including:

[0142] Obtain the key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram;

[0143] Input the application overview diagram into the trained application detection model, and return the application detection result output by the application detection model;

[0144] Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting.

[0145] Embodiment Five

[0146] Embodiment Five of the present invention also discloses a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, an application detection method is implemented, including:

[0147] Obtain the key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview diagram;

[0148] Input the application overview diagram into the trained application detection model, and return the application detection result output by the application detection model;

[0149] Among them, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting.

[0150] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0151] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0152] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0153] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0154] Note that the above is only the preferred embodiment of the present invention and the applied technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An application detection method, characterized in that, it includes: Obtain key application screenshots of the application to be detected, and splice the key application screenshots according to a preset splicing structure to generate an application overview map; Input the application overview map into a trained application detection model, and return the application detection result output by the application detection model; wherein, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature shortcut connection; wherein, the obtaining key application screenshots of the application to be detected and splicing the key application screenshots according to a preset splicing structure to generate an application overview map includes: Obtain the application screenshot sequence of the application to be detected, and obtain the image gray-scale features of each application screenshot in the application screenshot sequence; According to the image gray-scale features, cluster the application screenshot sequence into an application loading picture cluster and an application content picture cluster; Select a preset number of key application screenshots from the core areas of the application loading picture cluster and the core areas of the application content picture cluster respectively; Splice each key application screenshot according to a preset splicing structure to generate an application overview map.

2. The method according to claim 1, characterized in that, Inputting the application overview map into a trained application detection model and returning the application detection result output by the application detection model includes: Input the application overview map into a trained application detection model, and through the application detection model, use the method of four equal parts to layer by layer divide the application overview map to determine the bottom layer image blocks; Use the word vectors obtained by linearly mapping each bottom layer image block as the current layer input elements, map the attention matrix for every four current layer input elements, and calculate the self-attention features of each image block in the current layer; Perform inter-block feature aggregation processing on the self-attention features of every four image blocks in the current layer to obtain the feature information of each image block in the upper layer; Use the feature information of each image block in the upper layer as the current layer input elements, and return to execute the operation of mapping the attention matrix for every four current layer input elements and calculating the self-attention features of each image block in the current layer until the self-attention features of the complete application overview map are obtained; Merge the self-attention features of the downsampled bottom layer image blocks with the self-attention features of the application overview map, and make a decision according to the merged self-attention features to obtain the application detection result and return it.

3. The method according to claim 2, characterized in that, Mapping the attention matrix for every four current layer input elements and calculating the self-attention features of each image block in the current layer includes: For every four current layer input elements corresponding to the same image block, calculate the matching query matrix Q, key matrix K, and value matrix V; Use a non-linear function to perform position mapping on the four image sub-blocks included in each current layer image block to obtain an intra-block position relationship matrix P; According to the formula Z = (Q * K T + P) * V, calculate the self-attention feature Z of each image patch in the current layer; where K T represents the transpose of the key matrix K.

4. The method according to claim 3, characterized in that, Using a non-linear function to perform position mapping on the four image sub-blocks included in each current layer image block to obtain an intra-block position relationship matrix includes: Generate a pixel position coordinate matrix corresponding to each current layer image block; Among the matrix elements of the pixel position coordinate matrix, the x coordinate represents the coordinate of the image sub-block level, the y coordinate represents the coordinate of the pixel level within the image sub-block, and the matrix elements at the main diagonal position of each level have a coordinate of 0 at the corresponding level; Synchronously adjust each matrix element in the pixel position coordinate matrix to a non-negative value, perform a hashing process on the x coordinate, and sum the x coordinate and the y coordinate of each processed matrix element; Use a non-linear function to map the pixel position coordinate matrix after the summation process to a block internal position relationship matrix.

5. The method according to claim 1, wherein, after inputting the application overview map into the trained application detection model and returning the application detection result output by the application detection model, it further includes: In response to a detection backtracking request from the application detection request side, if the application to be tested is a target type application, then return the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request side for display; If the application to be tested is a non-target type application, then return the application overview map.

6. The method according to claim 5, wherein, Returning the application overview map marked with the judgment active block and the decision tree as the judgment basis to the application detection request side for display includes: Obtain the image self-attention feature used when the application detection model makes a judgment, and divide the image self-attention feature according to the multi-layer image blocks of the application overview map; Take the image self-attention feature corresponding to the topmost layer of image blocks as the current feature, calculate the score value corresponding to each current feature, and select the image block corresponding to the maximum score value as the active block of the current layer; Take the image self-attention feature of the next layer of image blocks corresponding to the active block of the current layer as the current feature, return to execute the operation of calculating the score value corresponding to each current feature, and select the image block corresponding to the maximum score value as the active block of the current layer until the active block of the bottom layer is determined; Generate a decision tree corresponding to each layer of active blocks, and return the decision tree and the application overview map marked with each layer of active blocks to the application detection request side for display.

7. An application detection device, wherein, it includes: A picture splicing module, configured to obtain key application screenshots of the application to be tested, and splice the key application screenshots according to a preset splicing structure to generate an application overview map; An application detection module, configured to input the application overview map into the trained application detection model and return the application detection result output by the application detection model; wherein, the application detection model is a Transformer model based on inter-block feature aggregation and shallow feature short-circuiting; Among them, the picture splicing module is specifically configured to obtain the application screenshot sequence of the application to be tested, and obtain the image gray-scale features of each application screenshot in the application screenshot sequence; cluster the application screenshot sequence into an application loading picture cluster and an application content picture cluster according to the image gray-scale features; select a preset number of key application screenshots from the core area of the application loading picture cluster and the core area of the application content picture cluster respectively; splice each key application screenshot according to a preset splicing structure to generate an application overview map.

8. A computer device, characterized in that, the device includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the application detection method according to any one of claims 1-6.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the application detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Weak supervision electric power drawing OCR identification method based on deep learning

    CN111860348A

  • Webpage content characterization method and device, webpage classification method and device and equipment

    CN112214707A