Method for recognizing injection attack video stream, and face recognition method and device
By converting image information in the frequency domain and using a deep learning model to identify video streams that are prone to injection attacks, the problem of needing to modify the permissions of mobile terminals in existing technologies is solved, and the effect of accurately identifying injection attacks is achieved without the aid of external devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DIDI INFINITY TECH & DEV CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods for identifying injection attacks require modifications to the underlying permissions or hardware of mobile terminals, which are complex and costly to implement, making it difficult to accurately identify injection attack video streams without the aid of external devices.
By extracting image information from the target video stream and converting it to the frequency domain, the frequency domain information is used to identify injection attacks. A deep learning model is then used for classification to identify video streams that have been injected with the attack.
It can accurately identify video streams that are injected with vulnerabilities without requiring additional mobile terminal permissions or external devices, thus ensuring the authentication security of real-time video streams from cameras.
Smart Images

Figure CN122116515A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile internet security, specifically to a method for identifying injected attack video streams, as well as a face recognition method and apparatus. Background Technology
[0002] Real-time image recognition based on mobile devices is widely used in the authentication and authorization processes of mobile internet applications. For example, for non-face-to-face authentication requiring personal confirmation, a mobile application can typically request the camera to capture a real-time facial image, and facial recognition can be performed based on the camera's preview video stream. Some hacking techniques modify the underlying software or hardware of mobile devices to provide pre-recorded video to the application when it requests the camera's real-time video stream, hoping to bypass the authentication process. This type of attack is known as injection attack.
[0003] Existing methods for identifying injection attacks require modifications to the underlying permissions of the operating system or requests for permissions regarding the mobile terminal's posture, and may even require modifications to the mobile terminal hardware. The implementation process is complex and costly. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method for identifying injected attack video streams, as well as a face recognition method and corresponding apparatus, which are expected to accurately identify injected attack video streams without relying on external devices or more mobile terminal permissions, thereby ensuring the security of the authentication process based on real-time camera video streams.
[0005] Firstly, a method for identifying injection attack video streams is provided, wherein the method includes:
[0006] The request is to obtain a target video stream, wherein the target video stream is expected to be a real-time video stream;
[0007] Obtain image information of at least a portion of the image frames from the target video stream;
[0008] Determine the frequency domain information corresponding to the image information;
[0009] Injection attacks are identified based on the frequency domain information.
[0010] Secondly, a face recognition method is provided, wherein the method includes:
[0011] The request is to obtain a target video stream, wherein the target video stream is expected to be a real-time video stream;
[0012] Obtain image information of at least a portion of the image frames from the target video stream;
[0013] Determine the frequency domain information corresponding to the image information;
[0014] Injection attacks can be identified based on the frequency domain information;
[0015] In response to the absence of an injection attack, face recognition is performed based on the target video stream.
[0016] Thirdly, a device for identifying injected attack video streams is provided, wherein the device includes:
[0017] A request unit is used to request the acquisition of a target video stream, wherein the target video stream is expected to be a real-time video stream;
[0018] An image information acquisition unit is used to acquire image information of at least a portion of the image frames of the target video stream;
[0019] A frequency domain conversion unit is used to determine the frequency domain information corresponding to the image information;
[0020] An attack identification unit is used to identify injection attacks based on the frequency domain information.
[0021] Fourthly, a facial recognition device is provided, wherein the device includes:
[0022] A request unit is used to request the acquisition of a target video stream, wherein the target video stream is expected to be a real-time video stream;
[0023] An image information acquisition unit is used to acquire image information of at least a portion of the image frames of the target video stream;
[0024] A frequency domain conversion unit is used to determine the frequency domain information corresponding to the image information;
[0025] An attack identification unit is used to identify injection attacks based on the frequency domain information.
[0026] A face recognition unit is used to perform face recognition based on the target video stream in response to an undetected injection attack.
[0027] Fifthly, an electronic device is provided, the controller including a memory and a processor, the memory for storing one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first or second aspect.
[0028] A sixth aspect provides a computer-readable storage medium having stored computer program instructions thereon, which, when executed by a processor, implement the method as described in the first or second aspect.
[0029] This invention extracts image information from the requested target video stream and converts the image information to the frequency domain to obtain the corresponding frequency domain information. Since image frames stored in local video files are inevitably compressed after image encoding and decoding, there is a significant difference in the frequency domain compared to the video stream captured in real time by the camera. Therefore, injection attacks can be identified by checking whether there are compression traces in the frequency domain information. Thus, the method and apparatus of this invention do not require additional mobile terminal permissions or external devices or hardware modifications to accurately identify injection attack video streams, ensuring the security of the authentication process based on real-time camera video streams. Attached Figure Description
[0030] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0031] Figure 1 This is a flowchart of a method for identifying injection attack video streams according to an embodiment of the present invention;
[0032] Figure 2 This is a flowchart illustrating how an embodiment of the present invention determines the corresponding frequency domain information based on image information;
[0033] Figure 3 This is a schematic diagram illustrating the process of acquiring frequency domain information according to an embodiment of the present invention;
[0034] Figure 4 This is a flowchart illustrating how an embodiment of the present invention uses a server to identify video streams subjected to injection attacks.
[0035] Figure 5 This is a flowchart of the face recognition method according to an embodiment of the present invention;
[0036] Figure 6 This is a schematic diagram of a device for identifying injected attack video streams according to an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of a face recognition device according to an embodiment of the present invention;
[0038] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0039] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0040] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0041] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0042] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0043] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only under the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0044] In the following description, facial recognition authentication is used as an example to illustrate the methods and apparatus of the embodiments of the present invention. It should be understood that these exemplary descriptions are not intended to limit the embodiments of the present invention, and the methods and apparatus of the embodiments of the present invention can be applied to any scenario that requires the identification of injected attack video streams.
[0045] Figure 1 This is a flowchart of a method for identifying injection attack video streams according to an embodiment of the present invention. Figure 1 As shown, the method for identifying injection attack video streams in this embodiment includes the following steps:
[0046] Step S100: Request to obtain the target video stream. The target video stream is expected to be a real-time captured video stream.
[0047] In this step, the client program on the mobile terminal requests access to camera data from the underlying operating system or platform program during the face recognition authentication process and obtains a handle to the video stream to acquire the video stream data. However, if subjected to an injection attack, the video stream handle provided by the operating system or platform program is not provided by the camera, but rather a video stream from a local or remote file. These two video streams are visually indistinguishable from each other at the image level. However, in the frequency domain, because the images in the video stream obtained from the local file are compressed, the high-frequency information typically processed during compression differs from the high-frequency information in the real-time preview video stream (uncompressed) captured by the camera. This application utilizes this difference to identify video streams subjected to injection attacks.
[0048] Step S200: Obtain image information of at least a portion of the image frames of the target video stream.
[0049] Specifically, the live video stream captured by the camera by the client application or its add-ons is uncompressed. Therefore, information is extracted at the pixel level from the image frames in the video stream to obtain the corresponding image information. This processing method ensures that the image is not compressed during processing, thus preventing it from being indistinguishable from a video stream used in an injection attack.
[0050] Meanwhile, since all the information in the complete target video stream consumes a lot of computing resources or network bandwidth and is unnecessary, it is necessary to process the information in the target video stream and obtain only the representative image information.
[0051] In one implementation, at least one image frame can be extracted from the target video stream. This image frame is, for example, one of several image frames with good clarity. Then, face recognition is performed on the image frame to extract the face image. Further, the acquired video stream is typically a color video stream, and the image frame can include multiple color channels. Selecting one color channel can effectively detect whether the high-frequency components of the image frame are compressed. In images using the YUV or YCrCb color space, the Y channel information from which the face information is extracted can be selected as the image information. The Y channel, short for Luminance Channel, represents the brightness or intensity information of the image. It is the most important part of the image because the human eye is more sensitive to changes in brightness than to changes in color. The Y channel contains the grayscale information of the image; that is, if only the Y signal component is present without U, V (or Cb, Cr) components, the displayed image is a black and white grayscale image. The Y channel plays a crucial role in image processing and video compression. Because the Y channel contains the main brightness information of the image, it has a decisive influence on the visual quality of the image. In video compression standards such as MPEG and JPEG, the Y channel is typically encoded preferentially to ensure basic image visibility. Because the Y channel contains a significant amount of image information, frequency domain detection of the Y channel information can more effectively identify video streams susceptible to injection attacks. It should be understood that, if saving computational resources is not a concern, all color channel data can be treated as image information and converted subsequently. Alternatively, information from other color channels besides the Y channel, such as the Cr or Cb channels, can also be used to effectively identify video streams susceptible to injection attacks. It should also be understood that if the video stream uses the RGB color space, which does not include Y channel information, the Y channel information can be obtained by converting the three element values of the RGB color space.
[0052] In this process, whether it is the operation of capturing an image frame, the operation of capturing a portion of an image frame, or the operation of extracting color channel information from a portion of an image, all operations are performed one by one based on pixel-level information to ensure that the image is not compressed during processing, thus making it indistinguishable from the video stream of an injection attack.
[0053] Step S300: Determine the frequency domain information corresponding to the image information.
[0054] In this step, image information with color or brightness information (such as Y channel information) is converted to the frequency domain to obtain frequency domain information that can be recognized by machine learning models.
[0055] Specifically, such as Figure 2 As shown, step S300 may include the following steps:
[0056] Step S310: Perform DCT transformation on the image information to determine the DCT coefficient matrix of the image information.
[0057] The Discrete Cosine Transform (DCT) is a transform technique used in signal processing and image compression. DCT transforms signal or image data into a series of frequency components that can be quantized and compressed while preserving data integrity and recoverability. The core idea of DCT is to represent a signal in the time (or spatial) domain as a superposition of cosine functions of different frequencies.
[0058] For example, such as Figure 3 As shown, image information 31 can be divided into multiple 8*8 pixel blocks 32. Then, DCT transformation is performed on these 8*8 pixel blocks 32-1, 32-2, etc., to obtain multiple corresponding sub-coefficient matrices 33. The sub-coefficient matrices 33 are represented as 8*8 matrices, and the elements of the matrix can be called DCT coefficients. These sub-coefficient matrices 33 can be labeled as 33-1, 33-2, etc., in sequence. Then, according to the position of pixel block 32 in image information 31, these sub-coefficient matrices 33 are arranged accordingly to obtain the DCT coefficient matrix 34 corresponding to image information 31. In the DCT coefficient matrix 34, the frequency domain information of each 8*8 represents the frequency domain distribution of the brightness or color of the pixel block at the corresponding position in the image information. The closer the element is to the upper left position in each 8*8 sub-coefficient matrix, the lower the corresponding frequency; the closer it is to the lower right position, the higher the corresponding frequency. Thus, the DCT coefficient matrix carries both spatial and frequency domain information.
[0059] Step S320: Determine the feature matrices corresponding to DCT coefficients with different absolute values based on the coefficient value distribution of the DCT coefficient matrix, so as to determine multiple feature matrices as the frequency domain information.
[0060] Based on step S310, step S320 further extracts the feature matrix as frequency domain information and converts it into a form suitable for deep learning model reading. In step S320, since the DCT coefficient values can be distributed within a predetermined range, for example, -20 to 20, through normalization techniques, this range can be divided into N coefficient value segments based on their absolute values, where N is an integer. Correspondingly, N feature matrices are set, with the same size as the DCT coefficient matrix. That is, if the DCT coefficient matrix is a 128*128 matrix, then each feature matrix is also a 128*128 matrix. Each feature matrix corresponds to a coefficient value segment and is used to record the intensity distribution of that coefficient value segment. To facilitate deep learning matrix reading, the feature matrix is set as a Boolean matrix, that is, a matrix whose elements only include 0 and 1.
[0061] For a target feature matrix, iterate through the coefficient values at each position in the DCT coefficient matrix. If the coefficient value at that position falls within a segment of the target feature matrix's coefficient values, set the element at that position to 1; otherwise, set the element to 0. This yields N feature matrices. These feature matrices characterize the frequency distribution of image information and contain spatial information, making them suitable for further processing by deep learning models.
[0062] Taking a 32x32 image as an example, assume that in a given scene, the coefficient values are integers, ranging from -20 to 20. The DCT coefficient matrix corresponding to the image information is a 32x32 matrix. Based on the absolute value, the coefficient range of -20 to 20 can be divided into 21 segments; that is, each possible absolute value (0-20) corresponds to one segment. Thus, 21 32x32 feature matrices can be set. Each feature matrix corresponds to a possible absolute value. Then, each feature matrix is assigned a value based on the DCT coefficient matrix. For feature matrix M18 corresponding to the coefficient value 18, for each of its 32x32 elements, if the element value (i.e., the coefficient value) at the corresponding matrix position in the DCT coefficient matrix is equal to 18, then the element value at that matrix position in feature matrix M18 is set to 1; otherwise, it is set to zero, until all positions are traversed.
[0063] Return to reference Figure 3 For the DCT coefficient matrix 34, for different coefficient values, segment 1 to segment N are used to set corresponding feature matrices 35-1, 35-2, etc. Based on whether the coefficient value at each position in the DCT coefficient matrix 34 falls into the corresponding coefficient value segment, values are assigned to the corresponding positions in the different feature matrices 35-1, 35-2, etc.
[0064] Step S400: Identify injection attacks based on the frequency domain information.
[0065] Specifically, injection attacks are identified by recognizing compression traces in the frequency domain information.
[0066] In this embodiment, a deep learning model is used to identify compression traces in the frequency domain information in order to determine whether the image frames in the target video stream come from a video file that has been compressed or encoded and stored locally.
[0067] In an optional implementation, in step S400, the multiple feature matrices are input into a pre-trained compressed video stream detection model to determine whether the target video stream is an injection attack video stream. The compressed video stream detection model can be a classification model based on a deep neural network, used to classify the input features. It can be obtained by training an initialized compressed video stream detection model using pre-collected positive and negative samples. For positive samples, video streams collected in real-time by the terminal device can be collected, and the video stream can be processed in the same way as in steps S100-S300 to obtain the corresponding frequency domain information as the input information of the positive sample, and its category as the output information. For negative samples, compressed and saved video streams can be collected, and the video stream can be processed in the same way as in steps S100-S300 to obtain the corresponding frequency domain information as the input information of the negative sample, and its category as the output information. Based on the sample set composed of positive and negative samples, an optimization algorithm can be used to train the model with the optimization of the model loss function as the objective. The optimization algorithm can be a gradient descent algorithm or an accelerated gradient descent algorithm, etc. The loss function can be mean squared error, cross-entropy loss, or Kullback-Leibler divergence, etc.
[0068] In one implementation, a deep neural network model with an HRNet (High-Resolution Net) architecture is used to construct the compressed video stream detection model. HRNet is a deep learning model architecture suitable for solving tasks requiring high-precision localization. The core idea of HRNet is to maintain high-resolution feature representations throughout the network while utilizing multi-scale information to enhance feature representation capabilities. This approach differs from the traditional strategy of first reducing resolution and then increasing it. It maintains high-resolution feature maps at all stages and processes information at different scales through parallel multiple branches, ensuring the flow and fusion of high-resolution information throughout the network. This design allows HRNet to better capture detailed information in images, which is particularly important for tasks requiring precise localization. Furthermore, HRNet further enhances the interaction and complementarity between features at different scales through a cross-stage multi-scale fusion mechanism. Simultaneously, to better adapt to the feature matrix comprising multiple 8*8 blocks, information at the same frequency positions in the feature matrix can be extracted by the deep learning network. The compressed video stream detection model uses convolutional layers with 8*8 convolutional kernels that dilate the input feature matrix, thereby convolving each 8*8 submatrix in the feature matrix to obtain features at the same frequency positions. This allows HRNet to more accurately identify compressed video streams.
[0069] In some implementations, steps S100-S400 of this embodiment can be performed entirely on the mobile terminal side, thereby identifying attack behavior from the terminal side when acquiring the video stream used for authentication, reporting the result to the server, or directly rejecting further operations on the mobile terminal side.
[0070] In some alternative implementations, due to the large amount of computational resources required for frequency domain information conversion and deep learning model-based recognition, these steps (steps S300-S400) can be performed by a server to achieve better real-time performance. In other words, steps S100-S200 in this embodiment are executed on the mobile terminal side. After the mobile terminal acquires the image information (e.g., Y-channel information) of the target video stream without any compression or encoding / decoding, it uploads the image information to the server. The server then converts the image information to frequency domain information and further identifies any compression artifacts in the frequency domain information. Since the image information (e.g., Y-channel information) is uploaded to the server uncompressed, there is no issue of compression during storage. This ensures that the server can effectively identify both the video stream captured in real-time by the camera and the video stream retrieved from locally stored video files.
[0071] Figure 4This is a flowchart illustrating how an embodiment of the present invention uses a server to identify video streams susceptible to injection attacks. For example... Figure 4 As shown, in step S410, the mobile terminal's application requests the camera's video stream according to authentication requirements. That is, the camera captures a video stream in real-time for preview. If no injection attack is received, the image frames of the video stream are not compressed during the conversion to a video file, and high-frequency information in the image frames is preserved. If the obtained video stream is an injection attack video stream, the image frames of the video stream must have been encoded or compressed, resulting in the loss of some high-frequency information.
[0072] In step S420, the mobile terminal acquires image information from at least a portion of the image frames of the target video stream. Specifically, when applied in a face recognition scenario, the image information is the Y-channel information of a face image cropped from one or more selected image frames.
[0073] In step S430, the mobile terminal uploads the image information to the server.
[0074] In step S440, the server performs frequency domain transformation on the image to determine the corresponding frequency domain information.
[0075] In step S450, the server identifies injection attacks based on frequency domain information.
[0076] Specifically, the server can use frequency domain information as part of the input information of a pre-trained compressed video stream detection model, and then use the compressed video stream detection model to classify the frequency domain information to identify whether the corresponding video stream is an injection attack video stream.
[0077] This invention extracts image information from the requested target video stream and converts the image information to the frequency domain. Since images stored locally are inevitably compressed after image encoding and decoding, there is a significant difference in the frequency domain compared to video streams captured in real time by the camera. Therefore, injection attacks can be identified by checking for compression traces in the frequency domain information. Thus, the method and apparatus of this invention can accurately identify injection attack video streams without obtaining additional mobile terminal permissions or relying on external devices, thereby ensuring information security.
[0078] Figure 5 This is a flowchart of a face recognition method according to an embodiment of the present invention. Figure 5 As shown, the face recognition method in this embodiment includes:
[0079] Step S510: Request to obtain the target video stream. The target video stream is expected to be a real-time captured video stream.
[0080] Step S520: Obtain image information of at least a portion of the image frames of the target video stream.
[0081] Specifically, in face recognition scenarios, the image information refers to the image information of a face image extracted from one or more selected image frames. Since the Y channel information retains more detailed image information, high-frequency information in the Y channel is usually processed to compress the data volume during video saving. Therefore, in this embodiment, the Y channel information is used as the image information. It should be understood that if saving computational resources is not required, all color channel data can be used as the overall image information, and subsequently converted to frequency domain information either as a whole or separately. Optionally, information from other color channels besides the Y channel, such as the Cr or Cb channels, can also be used, which can also effectively identify video streams susceptible to injection attacks. It should also be understood that if the video stream uses the RGB color space, which does not include Y channel information, the Y channel information can be obtained by converting the three element values of the RGB color space.
[0082] Using information from some or all color channels in a face image as image information can also make it easier to avoid repeatedly cropping the face image during subsequent face recognition processes, reducing repetitive operations, saving computing resources, and improving response speed.
[0083] Step S530: Determine the frequency domain information corresponding to the image information.
[0084] In this step, image information containing color or brightness information (e.g., Y channel information) is converted to the frequency domain to obtain frequency domain information that can be recognized by deep learning models. Specifically, the frequency domain information can be a discretized representation of the DCT coefficient matrix corresponding to the image information. That is, the range can be divided into N coefficient value segments based on absolute values, where N is an integer. Correspondingly, N feature matrices are set, with the same size as the DCT coefficient matrix. Each feature matrix corresponds to a coefficient value segment and is used to record the intensity distribution of that segment. For a target feature matrix, the coefficient values at each matrix position in the DCT coefficient matrix are traversed. If the coefficient value at that matrix position falls into a segment of the target feature matrix, the element at that matrix position in the target feature matrix is set to 1, and the elements at that matrix position in other feature matrices are set to 0; otherwise, the element at that matrix position is set to 0. Thus, N feature matrices with only 0 and 1 elements are obtained.
[0085] Step S540: Identify injection attacks based on the frequency domain information.
[0086] In one alternative implementation, frequency domain information is used as input to a compressed video stream detection model, which then classifies the frequency domain information to determine the identification result of injection attacks.
[0087] In step S550, it is determined whether the injection attack identification result indicates that the frequency domain information originates from a video file. If so, step S560 is executed; otherwise, step S570 is executed.
[0088] In step S560, in response to the detection of an injection attack, the face recognition process is terminated.
[0089] Furthermore, after terminating the facial recognition process, an alert can be sent to the server, indicating that an injection attack has been received. Simultaneously, relevant user information can be recorded to alert users of the risks.
[0090] In step S570, in response to the absence of an injection attack, face recognition is performed based on the target video stream.
[0091] Specifically, facial recognition can be performed directly on the server side using image information uploaded by the user, specifically one or more color channels of the captured facial information. Existing services or algorithms can be invoked for facial recognition.
[0092] Therefore, before facial recognition, it is possible to determine whether the video stream provided by the mobile terminal is compressed, and thus determine whether the video stream is a real-time video stream captured by the camera. This effectively identifies injection attacks before facial recognition is initiated, preventing the use of local video for facial recognition authentication by modifying the underlying software or operating system of the mobile terminal, thereby ensuring the security of user information.
[0093] Figure 6 This is a schematic diagram of a device for identifying injection attack video streams according to an embodiment of the present invention. Figure 6 As shown, the injection attack video stream identification device of this embodiment includes a request unit 61, an image information acquisition unit 62, a frequency domain conversion unit 63, and an attack identification unit 64. The request unit 61 is used to request the acquisition of a target video stream, wherein the target video stream is expected to be a real-time captured video stream; the image information acquisition unit 62 is used to acquire image information of at least a portion of the image frames of the target video stream; the frequency domain conversion unit 63 is used to determine the frequency domain information corresponding to the image information; and the attack identification unit 64 is used to identify the injection attack based on the frequency domain information.
[0094] This invention extracts image information from the requested target video stream and converts the image information to the frequency domain. Since the image stored locally is inevitably compressed after image encoding and decoding, it is significantly different from the video stream captured by the camera in real time in the frequency domain. Therefore, injection attacks can be identified by whether there are compression traces in the frequency domain information. Thus, the method and device of this invention can accurately identify injection attack video streams without obtaining additional mobile terminal permissions, thereby ensuring information security.
[0095] Figure 7 This is a schematic diagram of a face recognition device according to an embodiment of the present invention. Figure 7 As shown, the face recognition device in this embodiment includes a request unit 71, an image information acquisition unit 72, a frequency domain conversion unit 73, an attack identification unit 74, and a face recognition unit 75. The request unit 71 requests a target video stream, which is expected to be a real-time captured video stream. The image information acquisition unit 72 acquires image information of at least a portion of the image frames of the target video stream. The frequency domain conversion unit 73 determines the frequency domain information corresponding to the image information. The attack identification unit 74 identifies injection attacks based on the frequency domain information. The face recognition unit 75 performs face recognition based on the target video stream in response to the absence of detected injection attacks.
[0096] The device in this embodiment can determine whether the video stream provided by the mobile terminal is compressed before face recognition, and then determine whether the video stream is a video stream captured by the camera in real time. In this way, injection attacks can be effectively identified before face recognition is started, avoiding the use of local video for face recognition authentication by modifying the underlying software or operating system of the mobile terminal, thus ensuring the information security of users.
[0097] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. (For example...) Figure 8 As shown, the electronic device is used to constitute the mobile terminal or server described above. The electronic device includes a general computer hardware architecture, comprising at least a processor 81 and a memory 82.
[0098] Processor 81 and memory 82 are connected via bus 83. Memory 82 is adapted to store instructions or programs executable by processor 81. Processor 81 may be a standalone microprocessor or a collection of one or more microprocessors. Thus, processor 81 executes the instructions stored in memory 82 to perform the method flow of the embodiments of the present invention as described above, thereby realizing data processing and control of other devices. Bus 83 connects the aforementioned components together, and also connects the aforementioned components to display controller 84, display device, and input / output (I / O) device 85. Input / output (I / O) device 85 may be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, input / output device 85 is connected to the system via input / output (I / O) controller 86.
[0099] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.
[0101] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.
[0102] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.
[0103] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.
[0104] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0105] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for identifying video streams subjected to injection attacks, characterized in that, The method includes: The request is to obtain a target video stream, wherein the target video stream is expected to be a real-time video stream; Obtain image information of at least a portion of the image frames from the target video stream; Determine the frequency domain information corresponding to the image information; Injection attacks are identified based on the frequency domain information.
2. The method according to claim 1, characterized in that, The method includes: The image information is the Y-channel information of the target region of at least one image frame.
3. The method according to claim 2, characterized in that, The target video stream is a video stream used for face recognition; the target region is the face region.
4. The method according to claim 2, characterized in that, The frequency domain information consists of multiple feature matrices; Determining the frequency domain information corresponding to the image information includes: Perform DCT transformation on the image information to determine the DCT coefficient matrix of the image information; The feature matrix corresponding to DCT coefficients with different absolute values is determined based on the coefficient value distribution of the DCT coefficient matrix.
5. The method according to claim 4, characterized in that, The feature matrix corresponding to DCT coefficients with different absolute values is determined based on the coefficient value distribution of the DCT coefficient matrix, including: Determine the coefficient value segments corresponding to the coefficient values, and each coefficient value segment corresponds to a feature matrix; For each matrix position of each feature matrix, iterate and perform the following operation: If the coefficient value at the corresponding position in the DCT coefficient matrix falls into the corresponding coefficient value segment of the target feature matrix, the element at that position in the target feature matrix is set to 1; otherwise, the element at that position is set to 0.
6. The method according to claim 1, characterized in that, Identifying injection attacks based on the frequency domain information includes: Based on the frequency domain information, it can be determined whether the corresponding image information has been compressed in order to identify injection attacks.
7. The method according to claim 4, characterized in that, Identifying injection attacks based on the frequency domain information includes: Multiple feature matrices are input into a pre-trained compressed video stream detection model to determine whether the target video stream is an injection attack video stream.
8. The method according to claim 7, characterized in that, The compressed video stream detection model is trained as follows: The frequency domain information corresponding to the video stream collected in real time by the acquisition terminal device is used as a positive sample; The frequency domain information corresponding to the compressed and saved video stream is collected as negative samples. The initialized compressed video stream detection model is trained based on the positive and negative samples to determine the trained compressed video stream detection model.
9. A face recognition method, characterized in that, The method includes: The request is to obtain a target video stream, wherein the target video stream is expected to be a real-time video stream; Obtain image information of at least a portion of the image frames from the target video stream; Determine the frequency domain information corresponding to the image information; Injection attacks can be identified based on the frequency domain information; In response to the absence of an injection attack, face recognition is performed based on the target video stream.
10. A device for identifying injected attack video streams, characterized in that, The device includes: A request unit is used to request the acquisition of a target video stream, wherein the target video stream is expected to be a real-time video stream; An image information acquisition unit is used to acquire image information of at least a portion of the image frames of the target video stream; A frequency domain conversion unit is used to determine the frequency domain information corresponding to the image information; An attack identification unit is used to identify injection attacks based on the frequency domain information.
11. A face recognition device, characterized in that, The device includes: A request unit is used to request the acquisition of a target video stream, wherein the target video stream is expected to be a real-time video stream; An image information acquisition unit is used to acquire image information of at least a portion of the image frames of the target video stream; A frequency domain conversion unit is used to determine the frequency domain information corresponding to the image information; An attack identification unit is used to identify injection attacks based on the frequency domain information. A face recognition unit is used to perform face recognition based on the target video stream in response to an undetected injection attack.
12. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-8.
13. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method as described in any one of claims 1-8.