Malicious motion graph recognition method and device based on feature pyramid and storage medium

CN120014517APending Publication Date: 2025-05-16SHENHUA HOLLYSYS INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510101494.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

虽然3D卷积可以充分利用视频的信息,但是该结构的信息量巨大,计算复杂

Benefits of technology

[0016] The method provided by the present invention first extracts the key information of the animation by extracting key frames, and then uses the attention mechanism to guide the CNN of the multi-layer structure to focus the key features in the key frame, ignores the edge features in the figure and the background that is useless to the target, and improves the recognition ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014517A_ABST
    Figure CN120014517A_ABST
Patent Text Reader

Abstract

The invention relates to the field of malicious information identification, in particular to a malicious dynamic graph identification method and device based on a feature pyramid, a processor and a storage medium. The method comprises the following steps: acquiring a target dynamic graph, and extracting a key frame from the target dynamic graph; extracting a first feature map from the key frame; inputting the first feature map into a multi-scale feature pyramid module to form a plurality of multi-dimensional second feature maps; performing global average pooling and fusion expansion on the plurality of multi-dimensional second feature maps to form a third feature map with the same size as the first feature map; and judging whether the target dynamic graph is a malicious dynamic graph based on the third feature graph. According to the method provided by the invention, firstly, the key information of the motion picture is extracted by extracting the key frame, so that the calculation amount of a feature recognition process is greatly reduced, and the recognition efficiency of the motion picture and the video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of malicious information identification, and specifically to a malicious animated image identification method based on a feature pyramid, a malicious animated image identification device based on a feature pyramid, a machine-readable storage medium, and a processor. Background Art

[0002] In order to avoid malicious information being included in the transmitted information, the review agency needs to conduct intelligent review of the transmitted information before it is released or during transmission, and classify the information to distinguish whether the corresponding information is malicious or non-malicious.

[0003] Video classification methods are represented by the two-stream method and 3D convolution: The main idea of ​​3D convolution is to make full use of the extra dimensions in the video and use the structure of 3D convolution to convolve it. The design principle of this structure is that the closer the feature map is to the output layer, the more features it should have, that is, it can produce more types of features. The main idea of ​​the two-stream method is to train two classifiers, one for identifying RGB images and the other for identifying optical flow images, and mix the two results to get the recognition result of the entire video. Although 3D convolution can make full use of video information, the amount of information in this structure is huge and the calculation is complex. For the two-stream method, the model uses two networks of optical flow information and channel information, and the storage space required for the optical flow extraction algorithm is large and difficult to reproduce. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method, device, storage medium and processor for identifying malicious animated images based on a feature pyramid to improve the recognition efficiency of target animated images and videos.

[0005] In order to achieve the above-mentioned objectives, the first aspect of the present application provides a method for identifying malicious animated images based on a feature pyramid, comprising the following steps: obtaining a target animated image and extracting key frames from the target animated image; extracting a first feature map from the key frames; inputting the first feature map into a multi-scale feature pyramid module to form multiple multi-dimensional second feature maps; performing global average pooling and fusion expansion on the multiple multi-dimensional second feature maps to form a third feature map of the same size as the first feature map; and judging whether the target animated image is a malicious animated image based on the third feature map.

[0006] Based on the first aspect, in some embodiments of the present invention, extracting key frames from the target dynamic image includes: splitting the target dynamic image into multiple frames; evenly dividing the multiple frames into multiple pictures; selecting the most representative frame from each picture as the key frame of the picture, and using the key frames of the multiple pictures as the key frames of the target dynamic image.

[0007] Based on the first aspect, in some embodiments of the present invention, selecting the most representative frame of an image from each image as the key frame of the image includes: selecting an image from an image that has the highest similarity with other images in the image as the key frame of the image.

[0008] Based on the first aspect, in some embodiments of the present invention, the number of the third feature maps is equal to the number of the second feature maps, and both are greater than 1; judging whether the target animated image is a malicious animated image based on the third feature maps includes: adding multiple third feature maps to an Alexnet classifier for classification to obtain three groups of logits, and averaging the three groups of logits to obtain the classification of the target animated image.

[0009] Based on the first aspect, in some embodiments of the present invention, extracting the first feature map from the key frame includes: extracting the first feature map from the key frame using a residual block.

[0010] In a first aspect, the present application provides a malicious animation recognition device based on a feature pyramid, comprising: a first extraction unit, used to obtain a target animation and extract key frames from the target animation; a second extraction unit, used to extract a first feature map from the key frames; a processing unit, used to input the first feature map into a multi-scale feature pyramid module to form multiple multi-dimensional second feature maps; and after global average pooling and fusion expansion of the multiple multi-dimensional second feature maps, form a third feature map of the same size as the first feature map; a judgment unit, used to judge whether the target animation is a malicious animation based on the third feature map.

[0011] Based on the second aspect, in some embodiments, the first extraction unit extracts key frames from the target dynamic image using the following method: splitting the target dynamic image into multiple frames; evenly dividing the multiple frames into multiple pictures; selecting the most representative frame from each picture as the key frame of the picture, and using the key frames of the multiple pictures as the key frames of the target dynamic image.

[0012] Based on the second aspect, in some embodiments, the judgment unit judges whether the target animated image is a malicious animated image by the following method: adding multiple third feature maps to the Alexnet classifier for classification to obtain three groups of logits, and averaging the three groups of logits to obtain the classification of the target animated image; the number of the third feature maps is equal to the number of the second feature maps, and both are greater than 1.

[0013] In a third aspect, the present invention provides a processor configured to execute the above-mentioned malicious animated image recognition method based on feature pyramid.

[0014] In a fourth aspect, the present invention provides a machine-readable storage medium having instructions stored thereon, which, when executed by a processor, configure the processor to execute the above-mentioned malicious animated image identification method based on feature pyramid.

[0015] The method and device provided by the present invention have at least the following beneficial effects:

[0016] The method provided by the present invention first extracts the key information of the dynamic image by extracting key frames, and then uses the attention mechanism to guide the multi-layer CNN to focus on the key features in the key frames, ignoring the edge features in the image and the background that is useless to the target, thereby improving the recognition ability of the model.

[0017] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:

[0019] Figure 1 A schematic diagram of an application environment of a malicious animated image recognition method based on a feature pyramid according to an embodiment of the present application is schematically shown;

[0020] Figure 2 The flowchart of the method for identifying malicious animated images based on a feature pyramid according to an embodiment of the present application is schematically shown;

[0021] Figure 3 The structure block diagram of the malicious animated image recognition device based on feature pyramid according to an embodiment of the present application is schematically shown;

[0022] Figure 4 The internal structure diagram of the computer device according to the embodiment of the present application is schematically shown.

[0023] Description of Reference Numerals

[0024] 1-first extraction unit; 2-second extraction unit; 3-processing unit; 4-judgment unit. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0026] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), such directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0027] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0028] The malicious animated image recognition method based on feature pyramid provided in this application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0029] Figure 2 The following schematically shows a flow chart of a method for identifying malicious animated images based on a feature pyramid according to an embodiment of the present application. Figure 2 As shown, in one embodiment of the present application, a method for identifying malicious animated images based on a feature pyramid is provided. This embodiment mainly applies this method to the above Figure 1 Taking the terminal 110 (or server 120) in FIG. 1 as an example, the following steps are included:

[0030] S1, obtaining a target dynamic image, and extracting key frames from the target dynamic image;

[0031] Specifically, in this embodiment, the key frame selection process is mainly carried out by similarity comparison: first, the target dynamic image is split into multiple frames; then the multiple frames are evenly divided into multiple parts; finally, the most representative picture is selected from any one of the pictures as the key frame.

[0032] Exemplarily, the animated short video is split into frames, and then the total number of frames is evenly divided into three parts. Each frame in the three video frame sets is input into a pre-trained classification network for category score output. In each video, an image with the highest similarity to other images is selected as the key frame, and the three frames with the highest similarity in the three video frames are used as the key frames of the short video as input for subsequent recognition.

[0033] S2. Extracting a first feature map from the key frame;

[0034] Specifically, the three key frames obtained in step S1 are input into a residual block, the structure of the residual block is 1*1, 3*3, 1*1, and the first feature map is obtained and then input into the attention mechanism for the next step of operation. The plug-in using the attention mechanism mentioned in this embodiment is a non-dimensionality reduction channel attention module based on the feature pyramid.

[0035] S3, inputting the first feature map into a multi-scale feature pyramid module to form a plurality of multi-dimensional second feature maps;

[0036] In this step, the first feature map obtained in S2 is passed through a multi-scale feature pyramid module, which contains three convolution kernels of different scales, 1*1, 2*2 and 3*3. Three feature maps (i.e., second feature maps of different scales (dimensions)) are obtained by using three convolution kernels of different scales for one image.

[0037] S4, performing global average pooling and fusion expansion on the multiple multi-dimensional second feature maps to form a third feature map of the same size as the first feature map;

[0038] Specifically, in this step, the view function is used to expand the channel length of the 2*2 and 3*3 outputs to a size of 1*1, and then a 1*1 two-dimensional convolution kernel is used for dimensionality reduction so that the channel length of the three groups of outputs is equal to the original channel length. Finally, the channel weights are reshaped, the three groups of reshaped channels are fused, and expanded to the size of the first feature map.

[0039] This method enhances the recognition ability of text. On the one hand, it introduces feature pyramid and performs global average pooling. In addition, the 1*1 scale has stronger regularization and structural information. In addition, the model avoids dimensionality reduction and improves the shortcomings of large number of parameters and unnecessary information. At the same time, this method can enhance the correlation between similar features and weaken the correlation between different feature channels without introducing additional parameters.

[0040] S5. Determine whether the target animated image is a malicious animated image based on the third feature image.

[0041] Specifically, in this step, the three third feature maps obtained in step S4 are added to the Alexnet classifier for classification to obtain three groups of logits, which are finally averaged to obtain the classification of the target animated image.

[0042] Exemplarily, the method can be deployed in a browser for implementation. The implementation process mainly utilizes the QT plug-in mechanism to flexibly add an identification plug-in that can be used to identify malicious animated images.

[0043] The identification plug-in referred to in this embodiment is an application plug-in in QT. QT has a built-in plug-in mechanism, through which the software can support the plug-ins set by the user. Among them, QT has two APIs related to plug-ins. The first is used to expand the QT library itself, called the high-level API. The other is to expand the application developed by the QT library. The two APIs are different, and the latter is based on the former. In the present invention, what is used is the low-level API used to expand the application. The process of the QT plug-in is divided into two parts, including application support plug-ins and plug-in development. The specific steps are prior art and will not be repeated in this embodiment.

[0044] In the program, we write the recognition results into a structure, make it into a recognition plug-in, i.e., a dll library (dynamic link library), and then deploy it to the browser for recognition.

[0045] In addition, the present invention packages the video algorithm into a plug-in form and imports it into a dynamic link library for plug-and-play convenience.

[0046] Example 2

[0047] In this embodiment, if Figure 3 As shown, a malicious animated image recognition device based on a feature pyramid is provided, comprising:

[0048] A first extraction unit 1, used to obtain a target dynamic image and extract key frames from the target dynamic image;

[0049] A second extraction unit 2, used to extract a first feature map from the key frame;

[0050] Processing unit 3 is used to input the first feature map into a multi-scale feature pyramid module to form multiple multi-dimensional second feature maps; and perform global average pooling and fusion expansion on the multiple multi-dimensional second feature maps to form a third feature map of the same size as the first feature map;

[0051] The judging unit 4 is configured to judge whether the target animated image is a malicious animated image based on the third feature map.

[0052] Furthermore, the first extraction unit 1 extracts key frames by the following method: splitting the target animated image into multiple frames; evenly dividing the multiple frames into multiple parts; and selecting the most representative picture from any one of the pictures as the key frame.

[0053] Furthermore, the judgment unit 4 judges whether the target animated image is a malicious animated image by the following method: adding multiple third feature maps to the Alexnet classifier for classification, obtaining three groups of logits, and finally taking the average value to obtain the classification of this animated image; the number of the third feature maps is equal to the number of the second feature maps, and both are greater than 1.

[0054] The malicious animated image recognition device based on feature pyramid includes a processor and a memory. The first extraction unit, the second extraction unit, the processing unit and the judgment unit are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions.

[0055] The processor includes a kernel, which calls the corresponding program unit from the memory. One or more kernels can be set, and the malicious dynamic image recognition method based on the feature pyramid can be implemented by adjusting the kernel parameters.

[0056] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0057] An embodiment of the present application provides a storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned malicious dynamic image recognition method based on feature pyramid is implemented.

[0058] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor A01, a network interface A02, a memory (not shown in the figure) and a database (not shown in the figure) connected via a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02 and a database (not shown in the figure). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A04. The database of the computer device is used to store malicious animated image recognition data based on a feature pyramid. The network interface A02 of the computer device is used to communicate with an external terminal through a network connection. When the computer program B02 is executed by the processor A01, a malicious animated image recognition method based on a feature pyramid is implemented.

[0059] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0060] In one embodiment, the malicious dynamic image recognition device based on feature pyramid provided by the present application can be implemented in the form of a computer program, and the computer program can be used in Figure 4 The computer device is run on the computer device shown. The memory of the computer device can store various program units that constitute the malicious dynamic image recognition device based on feature pyramid, such as: Figure 3 The first extraction unit, the second extraction unit and the processing unit shown in the figure. The computer program composed of each program unit enables the processor to execute the steps of the malicious dynamic image recognition method based on feature pyramid in each embodiment of the present application described in this specification.

[0061] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0062] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0063] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0065] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0066] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0067] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0068] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0069] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for identifying malicious animated images based on feature pyramid, characterized in that: The following steps are involved: Obtain a target dynamic image, and extract key frames from the target dynamic image; Extracting a first feature map from the key frame; Inputting the first feature map into a multi-scale feature pyramid module to form multiple multi-dimensional second feature maps; Performing global average pooling and fusing and expanding the plurality of multi-dimensional second feature maps to form a third feature map of the same size as the first feature map; Based on the third feature map, it is determined whether the target animated image is a malicious animated image.

2. The method for identifying malicious animated images based on feature pyramid according to claim 1, characterized in that: The step of extracting key frames from the target dynamic image comprises: Splitting the target moving picture into multiple frames; Evenly divide the multiple frames of pictures into multiple pictures; The most representative frame of each picture is selected as the key frame of the picture, and the key frames of multiple pictures are used as the key frames of the target animated picture.

3. The method for identifying malicious animated images based on feature pyramid according to claim 2, characterized in that: The step of selecting the most representative frame of each image as the key frame of the image includes: Select a picture from a set of pictures that has the highest similarity with other pictures in the set of pictures as the key frame of the picture.

4. The method for identifying malicious animated images based on feature pyramid according to claim 3 is characterized in that: The number of the third characteristic graphs is equal to the number of the second characteristic graphs, and both are greater than 1; The determining whether the target animated image is a malicious animated image based on the third feature image includes: Add multiple third feature maps to the Alexnet classifier for classification to obtain three groups of logits, and then average the three groups of logits to obtain the classification of the target animation.

5. The method for identifying malicious animated images based on feature pyramid according to claim 1, characterized in that: The extracting the first feature map from the key frame includes: extracting the first feature map from the key frame using a residual block.

6. A malicious animated image recognition device based on feature pyramid, characterized in that: include: A first extraction unit, configured to obtain a target dynamic image and extract key frames from the target dynamic image; A second extraction unit, used to extract a first feature map from the key frame; A processing unit, configured to input the first feature map into a multi-scale feature pyramid module to form a plurality of multi-dimensional second feature maps; and perform global average pooling and fusion expansion on the plurality of multi-dimensional second feature maps to form a third feature map of the same size as the first feature map; A judging unit is used to judge whether the target animated image is a malicious animated image based on the third feature map.

7. The malicious animated image recognition device based on feature pyramid according to claim 6 is characterized in that: The first extraction unit extracts key frames from the target dynamic image using the following method: Splitting the target moving picture into multiple frames; Evenly divide the multiple frames of pictures into multiple pictures; The most representative frame of each picture is selected as the key frame of the picture, and the key frames of multiple pictures are used as the key frames of the target animated picture.

8. The malicious animated image recognition device based on feature pyramid according to claim 6, characterized in that: The judging unit judges whether the target animated image is a malicious animated image by the following method: Adding multiple third feature maps to the AlexNet classifier for classification to obtain three groups of logits, and averaging the three groups of logits to obtain the classification of the target animated image; The number of the third feature maps is equal to the number of the second feature maps, and both are greater than 1.

9. A processor, characterized in that: The method is configured to execute the malicious animated image recognition method based on feature pyramid according to any one of claims 1 to 5.

10. A machine-readable storage medium having instructions stored thereon, characterized in that: When the instruction is executed by a processor, the processor is configured to execute the malicious dynamic image recognition method based on feature pyramid according to any one of claims 1 to 5.