A biological detection method, device and equipment
By encoding multiple facial frames into a single frame for transmission and decoding at the server, the method addresses privacy and security concerns in facial recognition systems, reducing data transmission and enhancing detection performance.
Patent Information
- Application Number
- CN202210225193.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-07
AI Technical Summary
The existing facial recognition technology poses security risks in live attacks, and the frequent collection and storage of user facial information leads to privacy protection issues, which requires technical solutions that take into account both data volume and biological detection performance.
By using pre-trained encoder and decoder between the terminal device and the server, multi-frame images are encoded into single-frame images and decoded on the server side, the reconstruction of biological detection results is achieved, data transmission is reduced and user privacy is protected.
It reduces the data transmission bandwidth requirements, takes into account data bandwidth and attack prevention performance, and effectively prevents user privacy information leakage, achieving a balance between data volume and biological detection performance.
Smart Images

Figure CN114662144B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular, to a biological detection method, device, and equipment. Background Art
[0002] In recent years, face recognition technology has been greatly popularized. For example, face swiping payment in supermarkets, face unlocking of mobile phones, face access control in buildings, etc. The above applications have greatly facilitated people's daily life and work. However, face recognition also brings unprecedented security risks - live attacks (that is, an attacker makes tools such as images or masks of a certain user and attempts to use the above tools to pass through the face recognition system to steal the user's account information or resources, etc.).
[0003] With the popularization of face recognition technology, users' facial information is collected, uploaded, and stored more frequently. People have begun to worry about the abuse of their facial information. For this reason, regulatory authorities have also introduced relevant biometric image privacy protection laws. Therefore, how to protect users' facial images from being leaked is an important security issue. For this reason, a technical solution that can balance both the data volume and biological detection performance and protect user privacy is needed. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide a technical solution that can balance both the data volume and biological detection performance and protect user privacy.
[0005] In order to achieve the above technical solution, the embodiments of this specification are implemented as follows:
[0006] A biological detection method provided by the embodiments of this specification is applied to a terminal device. The method includes: obtaining multiple frames of images collected during the biological detection of a target user. Encoding the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image. The encoder is used to encode multiple frames of images into a single-frame image. Sending the single-frame encoded image to a server. The single-frame encoded image is used to trigger the server to decode the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determining the biological detection result of the target user based on the multiple frames of reconstructed images. Receiving the biological detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biological detection result of the target user.
[0007] A biological detection method provided by an embodiment of this specification is applied to a server. The method includes: receiving a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device through encoding a plurality of acquired images based on a pre-trained encoder, and the plurality of images are images acquired during the biological detection of a target user. Decoding the single-frame encoded image based on a pre-trained decoder to obtain a plurality of reconstructed images corresponding to the plurality of images. Performing biological detection on the target user based on the plurality of reconstructed images to obtain a biological detection result of the target user, and sending the biological detection result of the target user to the terminal device, where the biological detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0008] A biological detection device provided by an embodiment of this specification includes: an image acquisition module that acquires a plurality of images acquired during the biological detection of a target user. An encoding module that encodes the plurality of images based on a pre-trained encoder to generate a single-frame encoded image, where the encoder is used to encode a plurality of images into a single-frame image. An image sending module that sends the single-frame encoded image to a server, where the single-frame encoded image is used to trigger the server to decode the single-frame encoded image to obtain a plurality of reconstructed images corresponding to the plurality of images, and determine a biological detection result of the target user based on the plurality of reconstructed images. A detection result receiving module that receives the biological detection result of the target user sent by the server and processes the corresponding service requested by the target user based on the biological detection result of the target user.
[0009] A biological detection device provided by an embodiment of this specification includes: an encoded image receiving module that receives a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device through encoding a plurality of acquired images based on a pre-trained encoder, and the plurality of images are images acquired during the biological detection of a target user. A decoding module that decodes the single-frame encoded image based on a pre-trained decoder to obtain a plurality of reconstructed images corresponding to the plurality of images. A detection result sending module that performs biological detection on the target user based on the plurality of reconstructed images to obtain a biological detection result of the target user, and sends the biological detection result of the target user to the terminal device, where the biological detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0010] A biological detection device provided by an embodiment of this specification, the biological detection device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions when executed cause the processor to: obtain multiple frames of images collected during the biological detection of a target user. Based on a pre-trained encoder, perform encoding processing on the multiple frames of images to generate a single-frame encoded image, the encoder being used to encode multiple frames of images into a single-frame image. Send the single-frame encoded image to a server, the single-frame encoded image being used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biological detection result of the target user based on the multiple frames of reconstructed images. Receive the biological detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biological detection result of the target user.
[0011] A biological detection device provided by an embodiment of this specification, the biological detection device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions when executed cause the processor to: receive a single-frame encoded image sent by a terminal device, the single-frame encoded image being a single-frame image generated by the terminal device based on a pre-trained encoder performing encoding processing on acquired multiple frames of images, the multiple frames of images being images collected during the biological detection of a target user. Based on a pre-trained decoder, perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images. Perform biological detection on the target user based on the multiple frames of reconstructed images to obtain the biological detection result of the target user, and send the biological detection result of the target user to the terminal device, the biological detection result being used to trigger the terminal device to process the corresponding service requested by the target user.
[0012] An embodiment of this specification further provides a storage medium, the storage medium is used to store computer-executable instructions, and the executable instructions when executed by a processor implement the following process: obtain multiple frames of images collected during the biological detection of a target user. Based on a pre-trained encoder, perform encoding processing on the multiple frames of images to generate a single-frame encoded image, the encoder being used to encode multiple frames of images into a single-frame image. Send the single-frame encoded image to a server, the single-frame encoded image being used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biological detection result of the target user based on the multiple frames of reconstructed images. Receive the biological detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biological detection result of the target user.
[0013] An embodiment of this specification also provides a storage medium for storing computer-executable instructions. When the executable instructions are executed by a processor, the following process is implemented: receiving a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device through encoding a plurality of acquired images based on a pre-trained encoder, and the plurality of images are images acquired during the biological detection of a target user. Decoding the single-frame encoded image based on a pre-trained decoder to obtain a plurality of reconstructed images corresponding to the plurality of images. Performing biological detection on the target user based on the plurality of reconstructed images to obtain a biological detection result of the target user, and sending the biological detection result of the target user to the terminal device, where the biological detection result is used to trigger the terminal device to process the corresponding service requested by the target user. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1A This is an embodiment of a biological detection method in this specification;
[0016] Figure 1B This is a schematic diagram of a biological detection process in this specification;
[0017] Figure 2 This is a schematic diagram of the structure of a biological detection system in this specification;
[0018] Figure 3 This is a schematic diagram of another biological detection process in this specification;
[0019] Figure 4A This is another embodiment of a biological detection method in this specification;
[0020] Figure 4B This is a schematic diagram of yet another biological detection process in this specification;
[0021] Figure 5 This is a schematic diagram of yet another biological detection process in this specification;
[0022] Figure 6 This is an embodiment of a biological detection device in this specification;
[0023] Figure 7 This is another embodiment of a biological detection device in this specification;
[0024] Figure 8 This is an embodiment of a biological detection device in this specification. Detailed implementation manners
[0025] The embodiments of this specification provide a biological detection method, device and equipment.
[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0027] Embodiment 1
[0028] As Figure 1A and Figure 1B shown, the embodiments of this specification provide a biological detection method. The execution subject of this method can be a terminal device. Among them, the terminal device can be a certain terminal device such as a mobile phone, a tablet computer, etc., or can also be a computer device such as a notebook computer or a desktop computer, or can also be an IoT device (specifically such as a smart watch, a vehicle-mounted device, etc.). This method can specifically include the following steps:
[0029] In step S102, obtain multiple frames of images collected during the biological detection of the target user.
[0030] Among them, the target user can be any user who needs to undergo biological detection. Biological detection can be biometric detection or detection of the user's actions and behaviors. The biological detection in this embodiment can be used to determine whether the object undergoing biological detection is a real person and the user himself (rather than an image, video, etc. of the user). The multiple frames of images can include images of a specified part of the target user's body. For example, they can include images of the face and head or can only include images of the face, etc. The multiple frames of images can be images that can perform biological detection on the target user and can obtain accurate biological detection results, which can be specifically set according to the actual situation, and this specification does not limit this.
[0031] In implementation, in recent years, face recognition technology has been greatly popularized. For example, face payment in supermarkets, face unlocking of mobile phones, face access control in buildings, etc. The above applications have greatly facilitated people's daily life and work. However, face recognition also brings security risks that have never occurred before - spoofing attacks (i.e., an attacker makes tools such as images or masks of a certain user and attempts to use the above tools to pass through the face recognition system for recognition, thereby stealing the user's account information or resources, etc.). To address the above possible security risks, currently, biometric detection algorithms have been widely integrated into face recognition systems. Biometric detection algorithms can determine whether there is a risk of spoofing attacks for the subject undergoing face recognition through single-frame images or multi-frame images. However, the biometric detection performance based on single-frame images is poor, but the amount of data to be transmitted and processed is small. For biometric detection based on multi-frame images, the performance is good, but the amount of data to be transmitted and processed is large.
[0032] With the popularization of face recognition technology, users' facial information is collected, uploaded, and stored more frequently. People have begun to worry about the abuse of their facial information. Therefore, regulatory authorities have also introduced relevant biometric image privacy protection bills. Therefore, how to protect users' facial images from being leaked is an important security issue. For this purpose, a technical solution that can balance the amount of data and biometric detection performance while protecting user privacy is needed. The embodiments of this specification provide an implementable technical solution, which can specifically include the following content:
[0033] An application program for executing a certain service can be installed in the terminal device of the user (i.e., the target user). Buttons or hyperlinks for triggering the service can be set in the application program. When the target user needs to execute the service, the application program can be opened, and the above buttons or hyperlinks can be clicked. At this time, the application program in the terminal device can determine whether biometric detection of the target user is required when executing the service. Or, during the execution of the service by the terminal device, if a certain process requires biometric detection of the target user, the terminal device can start the camera component and, with the authorization of the target user, perform biometric detection on the current target user through the camera component. At this time, multiple frames of images collected during the biometric detection of the target user can be obtained. The multiple frames of images can be images collected at every preset time interval, or can be a video of the biometric detection process taken by the camera component. The images in the video, etc., can be specifically set according to the actual situation. The embodiments of this specification do not make limitations in this regard.
[0034] In step S104, encode the multi-frame images based on a pre-trained encoder to generate a single-frame encoded image. The encoder is used to encode multi-frame images into a single-frame image.
[0035] Among them, an encoder may be a component that encodes and converts a certain signal (such as a bitstream, etc.) or data into a signal form or data form that can be used for communication, transmission, and storage. The encoder may implement the corresponding functions of the encoder through an application program (i.e., in the form of software). For example, a certain network model algorithm (such as a neural network model algorithm, etc.) can be used to construct a corresponding network model, and the corresponding functions of the encoder can be implemented through this network model. In addition, the encoder may also be implemented by combining a hardware device and an application program (i.e., in the form of hardware + software). The encoder in this embodiment can generate a single frame of image by encoding multiple different frames of images. In addition, the single frame of image generated by the encoder can be decoded by a decoder that matches the encoder to restore the multiple frames of images before encoding. The single-frame encoded image means that the corresponding encoded image is generated, and there is only one frame of this encoded image.
[0036] In implementation, to ensure the performance of biometric detection, there may be a relatively large number of images required in biometric detection. In this case, the transmission of images will consume a relatively large amount of network resources, and the amount of data to be transmitted is also large. In addition, directly transmitting multiple frames of images may lead to the leakage of private data of the target user. To balance the data volume, biometric detection performance, and the protection of the target user's private data, an encoder can be pre-constructed. Through this encoder, multiple frames of images can be encoded into a single-frame image, thereby reducing the amount of data to be transmitted. At the same time, since multiple frames of images are converted into a single-frame image, it will be very difficult to distinguish the user's private data in the obtained single-frame image. In addition, the single-frame image generated by the encoder can be decoded by a decoder that matches the encoder to restore the multiple frames of images before encoding for biometric detection processing. Thus, the purpose of balancing the data volume, biometric detection performance, and protecting the target user's private data can be achieved. Specifically, corresponding algorithms can be selected based on the above requirements. For example, algorithms that can achieve the above purpose can be selected from relevant deep learning algorithms (such as convolutional neural network algorithms, etc.), and the architecture of the encoder can be constructed through the selected algorithms. Then, multiple frames of images collected during the biometric identification process of multiple different users can be obtained as training samples. In addition, the convergence conditions for model training can be set according to the actual situation, and specific conditions or requirements can also be set for the generated single-frame image, such as generating a single-frame image with 3 channels or generating a single-frame image with 4 channels, etc., which can be specifically set according to the actual situation. The obtained training samples can be input into the architecture of the encoder to train the encoder. When the trained encoder meets the above-set convergence conditions (at the same time, the output single-frame image meets the set specific conditions or requirements), the training of the encoder can be stopped, thereby obtaining the trained encoder. If the trained encoder does not meet the above-set convergence conditions, the encoder can be continuously trained with the training samples until the trained encoder meets the above-set convergence conditions.
[0037] After obtaining the encoder through the above training method, the obtained multiple frames of images can be input into the encoder. The encoder analyzes the multiple frames of images and re-encodes the multiple frames of images based on the analysis results, outputting a single-frame encoded image. Since the features in multiple frames of images are concentrated in one frame of image, it will be very difficult to distinguish the private data of the target user contained in the multiple frames of images in the single-frame encoded image, thereby protecting the user's private data from being leaked and reducing the amount of subsequent data transmission at the same time.
[0038] In step S106, the encoded image of a single frame is sent to the server, and the encoded image of the single frame is used to trigger the server to decode the encoded image of the single frame, obtain multiple reconstructed images corresponding to multiple frames of images, and determine the biometric detection result of the target user based on the multiple reconstructed images.
[0039] Among them, the server can be a biometric detection server or the background server of the above services, etc., which can be specifically set according to the actual situation, and the embodiments of this specification do not limit this. The reconstructed image can be an image reconstructed by decoding the encoded image in order to restore the image before encoding processing. Since some information is often lost after the image is encoded, the purpose of the decoding process is to restore or recover the original image (i.e., the image before encoding processing). However, due to the loss of some information in the image, the reconstructed image is highly similar to the original image, but there may be certain differences.
[0040] In implementation, as Figure 2 shown, the terminal device can send the obtained encoded image of the single frame to the server with the authorization of the target user. After receiving the encoded image of the single frame, the server can analyze the encoded image of the single frame. If it is determined that the encoded image of the single frame is an image obtained after encoding processing, the decoding program preset in the server corresponding to the above encoder can be started. Through this decoding program, the encoded image of the single frame can be decoded to restore the multiple frames of images before encoding processing from the encoded image of the single frame. During the decoding process, the server can reconstruct the corresponding image based on the image features parsed by the decoding program, so as to obtain multiple reconstructed images corresponding to multiple frames of images. Since the multiple reconstructed images are highly similar to the multiple frames of images before encoding processing, the target user can be biometrically detected based on the multiple reconstructed images, so as to obtain the biometric detection result of the target user.
[0041] In step S108, receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user.
[0042] In implementation, after the server obtains the biometric detection result of the target user, it can send the biometric detection result to the terminal device. The terminal device can present the biometric detection result to the target user, and the target user can process the corresponding service requested by the target user based on the biometric detection result. For example, if the biometric detection result is a pass, the terminal device can continue to respond to the above service request of the target user, so as to process the corresponding service requested by the target user through the data interaction between the terminal device and the server.
[0043] An embodiment of this specification provides a biological detection method, which is applied to a terminal device. By acquiring multiple frames of images collected during the biological detection of a target user, then, based on a pre-trained encoder, encoding the multiple frames of images to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to decode the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determining the biological detection result of the target user based on the multiple frames of reconstructed images, receiving the biological detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biological detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then decoding the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth for biological detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0044] Embodiment 2
[0045] As Figure 3 shown, an embodiment of this specification provides a biological detection method. The execution subject of this method can be a terminal device. Among them, the terminal device can be a certain terminal device such as a mobile phone or a tablet computer, or can also be a computer device such as a laptop computer or a desktop computer, or can also be an IoT device (specifically such as a smart watch, a vehicle-mounted device, etc.).
[0046] This method can specifically include the following steps:
[0047] In step S302, acquire multiple frames of sample images for biological detection.
[0048] In implementation, multiple frames of sample images can be acquired in various ways. For example, multiple frames of images during the biological detection of users can be purchased from multiple different users. Among them, multiple sample images can be anonymized images, or can also be images used with the special authorization of the users, etc. In addition, multiple frames of sample images for biological detection can also be acquired during the process of a user performing a certain service, etc., which can be specifically set according to the actual situation, and this specification does not limit this.
[0049] In step S304, based on multi-frame sample images, the conditions that the single-frame image generated by the encoder needs to satisfy, and a preset first loss function, the encoder is trained to obtain a trained encoder. The conditions that the single-frame image needs to satisfy include the conditions for the number of channels of the single-frame image. The first loss function is determined based on image style loss information and / or image content loss information.
[0050] Among them, the encoder can be constructed by a variety of different machine learning models. In this embodiment, the encoder can be constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model. The U-Net model presents a structure similar to the letter "U". It consists of a contracting path on the left half and an expansive path on the right half. The contracting path can be constructed by a convolutional neural network, and the structure of 2 convolutional layers and 1 max pooling layer can be repeatedly adopted. After each pooling operation, the dimension of the image will increase. In the expansive path, first perform 1 deconvolution operation to halve the dimension of the image, and then splice and crop it corresponding to the contracting path to obtain the corresponding feature map. Based on the above feature maps, a new feature map is re-formed, and then 2 convolutional layers are used for feature extraction, and the above structure is repeated. In the final output layer, 2 convolutional layers are used to map the high-dimensional feature map into a low-dimensional output image. The U-Net model can be specifically divided into two parts: upsampling and downsampling. The downsampling part mainly uses continuous convolutional pooling layers to extract feature information in the image and gradually maps the feature information to a high dimension. There is rich feature information of the entire image at the highest dimension of the entire network. The U-Net model can directly map the high-dimensional features to a low dimension through deconvolution processing instead of directly pooling the image and directly upsampling it to an output image of the same size as the original image. During the mapping process, in order to enhance the segmentation accuracy, the image with the same dimension in the contracting network in the same dimension will be fused. Since the dimension will become twice the original dimension during the fusion process, it is necessary to perform convolution processing again to ensure that the dimension after processing is the same as the dimension before the fusion operation, so that after another deconvolution processing, it can be fused with the image in the same dimension again until the dimension is the same as the original image and then the output image is output. The structure of the encoder in this embodiment can be composed of a U-Net model with a certain number of network layers. Specifically, for example, it can be composed of a U-Net model with 8 or 10 network layers, etc., which can be specifically set according to the actual situation, and this specification embodiment does not limit this. The image style loss information can include the loss information of the image texture, etc. In practical applications, the smaller the loss value corresponding to the image style loss information, the closer the generated single-frame image is to the original image. The larger the loss value corresponding to the image content loss information, the more conducive to the protection of user privacy data in the image.
[0051] In implementation, an initial architecture of the encoder can be constructed based on the network model structure of the U-Net model. Multiple-frame sample images are input into the constructed encoder to generate a single-frame image that meets the conditions. Then, based on the generated single-frame image and a preset first loss function, corresponding loss information can be determined. Based on the determined loss information, relevant parameters in the encoder are adjusted. Then, the multiple-frame sample images are input into the constructed encoder again to generate a single-frame image that meets the conditions. Then, based on the generated single-frame image and the preset first loss function, corresponding loss information can be determined. If the determined loss information meets the convergence condition, the trained encoder is output. If the determined loss information does not meet the convergence condition, the relevant parameters in the encoder are adjusted based on the determined loss information, and the encoder is trained again in the above manner until the determined loss information meets the convergence condition.
[0052] It should be noted that, in order to more simply and effectively determine the image style loss information and the image content loss information, a VGG network model pre-trained on ImageNet (such as the VGG19 network model, etc.) can be used as an auxiliary (wherein, the parameters of the VGG network model remain unchanged during training). Based on this, the image style loss information can be determined based on minimizing the distance between the features of the single-frame image generated by the encoder and the multiple-frame sample images in the first preset network layer of the VGG network model. The image content loss information can be determined based on maximizing the distance between the features of the single-frame image generated by the encoder and the multiple-frame sample images in the second preset network layer of the VGG network model. Among them, the first preset network layer can be set according to the actual situation, specifically such as the 10th network layer. Correspondingly, the image style loss information can be determined based on minimizing the L2 distance between the features of the single-frame image generated by the encoder and the multiple-frame sample images in the 10th network layer of the VGG network model. The second preset network layer can be set according to the actual situation, specifically such as the 16th network layer. Correspondingly, the image content loss information can be determined based on maximizing the L2 distance between the features of the single-frame image generated by the encoder and the multiple-frame sample images in the 16th network layer of the VGG network model.
[0053] In addition, the training of the above encoder is completed in the terminal device. In actual applications, the training of the encoder can also be completed in the server. Based on this, the following processing can be included: receiving the trained encoder sent by the server. The trained encoder is obtained by the server based on the multiple-frame sample images for biometric detection, the conditions that the single-frame image generated by the encoder needs to meet, and a preset first loss function.
[0054] Among them, the processing process for the server to train the encoder can refer to the above relevant content and will not be elaborated here.
[0055] After obtaining the encoder in the above manner, the terminal device can apply the encoder to perform biometric detection processing, which can specifically include the processing of steps S306 to S314 below.
[0056] In step S306, multiple frames of images collected during the biometric detection of the target user are obtained.
[0057] In step S308, the multiple frames of images are encoded based on the pre-trained encoder to generate a single-frame encoded image. The encoder is used to encode multiple frames of images into a single-frame image.
[0058] In step S310, the encoded image is compressed based on a preset image compression algorithm to obtain a compressed encoded image.
[0059] Among them, there can be various image compression algorithms, such as the JPEG image compression algorithm, the image compression algorithm based on Huffman coding, etc., which can be specifically set according to the actual situation, and this specification embodiment does not limit this.
[0060] In implementation, the image compression algorithm can be pre-selected according to the actual situation. Through the image compression algorithm, the data volume of the single-frame encoded image can be further reduced, thereby obtaining the compressed encoded image.
[0061] In step S312, the compressed encoded image is sent to the server. The encoded image is used to trigger the server to decode the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and the biometric detection result of the target user is determined based on the multiple frames of reconstructed images.
[0062] In step S314, the biometric detection result of the target user sent by the server is received, and the corresponding service requested by the target user is processed based on the biometric detection result of the target user.
[0063] An embodiment of this specification provides a biological detection method, which is applied to a terminal device. By acquiring multiple frames of images collected during the biological detection of a target user, and then encoding the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to decode the single-frame encoded image to obtain multiple reconstructed images corresponding to the multiple frames of images, and determining the biological detection result of the target user based on the multiple reconstructed images, receiving the biological detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biological detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, and then decoding the single-frame encoded image on the server side to restore or reconstruct the original images, the requirement for data transmission bandwidth for biological detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0064] Embodiment 3
[0065] As Figure 4A and Figure 4B As shown, an embodiment of this specification provides a biological detection method. The execution subject of this method can be a server. Among them, the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server such as a financial service or an online shopping service, or a background server of a certain application program, etc. This method can specifically include the following steps:
[0066] In step S402, receive the single-frame encoded image sent by the terminal device. The single-frame encoded image is a single-frame image generated by the terminal device encoding the acquired multiple frames of images based on a pre-trained encoder, and the multiple frames of images are images collected during the biological detection of the target user.
[0067] In step S404, decode the single-frame encoded image based on a pre-trained decoder to obtain multiple reconstructed images corresponding to the multiple frames of images.
[0068] Among them, the decoder can be a component that restores information from an encoded form to its original form. The decoder can implement the corresponding functions of the decoder through an application program (i.e., in the form of software). For example, a corresponding network model can be constructed using a certain network model algorithm (such as a neural network model algorithm, etc.), and the corresponding functions of the decoder can be implemented through this network model. In addition, the decoder can also be implemented by combining a hardware device and an application program (i.e., in the form of hardware + software). The decoder in this embodiment can decode a single-frame encoded image generated after encoding processing to reconstruct or restore the multi-frame image before encoding processing.
[0069] In implementation, to ensure the performance of biometric detection, there may be a large number of images required in biometric detection. In this way, the transmission of images will consume a large amount of network resources, and the amount of data to be transmitted is large. In addition, directly transmitting multi-frame images may cause the leakage of private data of the target user. To balance the data volume and biometric detection performance, as well as the protection of the private data of the target user, a corresponding encoder is constructed. Through this encoder, multi-frame images can be encoded into a single-frame image, thereby reducing the amount of data to be transmitted. At the same time, the single-frame image generated by this encoder is decoded by a decoder that matches this encoder to restore the multi-frame image before encoding processing, and then biometric detection processing is performed, achieving the purpose of balancing the data volume and biometric detection performance while being able to protect the private data of the target user. Specifically, a corresponding algorithm can be selected based on the above requirements. For example, an algorithm that can achieve the above purpose can be selected from relevant algorithms in deep learning (specifically, such as a convolutional neural network algorithm, etc.), and the architecture of the decoder can be constructed through the selected algorithm. Then, a single-frame image generated by the encoding processing of the encoder can be obtained as a training sample. In addition, the convergence conditions for model training can also be set according to the actual situation, which can be set specifically according to the actual situation. The obtained training sample can be input into the architecture of the decoder to train the decoder. When the trained decoder meets the above-set convergence conditions, the training of the decoder can be stopped, thus obtaining the trained decoder. If the trained decoder does not meet the above-set convergence conditions, the decoder can be continuously trained with the training sample until the trained decoder meets the above-set convergence conditions.
[0070] After obtaining the decoder through the above training method, the obtained single-frame encoded image can be input into the decoder. The decoder analyzes the single-frame encoded image and reconstructs the corresponding multi-frame image based on the analysis result, and outputs the multi-frame reconstructed images corresponding to the above multi-frame images.
[0071] In step S406, perform biometric detection on the target user based on the multi-frame reconstructed images, obtain the biometric detection result of the target user, and send the biometric detection result of the target user to the terminal device. This biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0072] For the specific processing procedures of the above steps S402 and S406, reference can be made to the above relevant content, which will not be elaborated here.
[0073] An embodiment of this specification provides a biometric detection method, which is applied to a server. The terminal device acquires multiple frames of images collected during the biometric detection of the target user, and then encodes the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, and sends the single-frame encoded image to the server to trigger the server to decode the single-frame encoded image to obtain the multi-frame reconstructed images corresponding to the above multiple frames of images, and determine the biometric detection result of the target user based on the multi-frame reconstructed images, receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then decoding the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth for biometric detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, with a strong privacy protection function, and can effectively prevent the leakage of users' privacy information.
[0074] Embodiment Four
[0075] As Figure 5 shown, an embodiment of this specification provides a biometric detection method. The execution subject of this method can be a server. Among them, the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server such as a financial service or an online shopping service, or a background server of a certain application program, etc. This method can specifically include the following steps:
[0076] In step S502, obtain multiple frames of image samples for biometric detection.
[0077] Among them, the multiple frames of image samples can include various acquisition methods, such as specific purchase methods or obtaining after data anonymization processing with user authorization, etc. Specifically, it can be set according to the actual situation, and this specification does not limit this.
[0078] In step S504, multiple-frame image samples are input into the encoder to train the encoder, and a single-frame encoded image sample output by the encoder is obtained. The single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample.
[0079] For the specific training process of the encoder in step S504 above, reference can be made to the relevant content above, which will not be elaborated here.
[0080] In step S506, the single-frame encoded image sample is input into the decoder to train the decoder, and multiple-frame reconstructed image samples output by the decoder are obtained.
[0081] For the specific training process of the decoder in step S506 above, reference can be made to the relevant content above, which will not be elaborated here.
[0082] In step S508, if it is determined based on the multiple-frame reconstructed image samples that the encoder and the decoder meet the preset first convergence condition, the decoder is stored, and the encoder is sent to the terminal device.
[0083] Among them, the preset first convergence condition may include conditions corresponding to a first loss function and a second loss function. The first loss function is determined based on image style loss information and / or image content loss information. The second loss function is determined based on image reconstruction loss information. The image style loss information is determined based on minimizing the distance between the features of the single-frame encoded image sample and the multiple-frame image samples in the first preset network layer of the VGG network model. The image content loss information is determined based on maximizing the distance between the features of the single-frame encoded image sample and the multiple-frame image samples in the second preset network layer of the VGG network model. The image reconstruction loss information is determined based on minimizing the distance between the features of the multiple-frame reconstructed image samples and the multiple-frame image samples. In practical applications, the image reconstruction loss information can be determined based on minimizing the L2 distance between the multiple-frame reconstructed image samples and the multiple-frame image samples.
[0084] In step S510, if it is determined based on the multiple-frame reconstructed image samples that the encoder and the decoder do not meet the preset first convergence condition, the process of obtaining multiple-frame image samples and training the encoder and the decoder is re-executed until the preset first convergence condition is met.
[0085] In implementation, if it is determined based on the multiple-frame reconstructed image samples that the encoder and the decoder do not meet the preset first convergence condition, the processing of steps S502 to S510 above is re-executed, that is, the process of obtaining multiple-frame image samples and training the encoder and the decoder is re-executed until the preset first convergence condition is met.
[0086] It should be noted that the above content is the process where both the encoder and the decoder are trained on the server side. In actual applications, the encoder can also be trained by the terminal device. For specific details, please refer to the relevant content mentioned above. At the same time, the decoder can be trained by the server. Based on this, the corresponding processing can include the processing of the following step A2 and this step A4.
[0087] In step A2, a single-frame encoded image sample obtained after the encoder encodes multiple-frame image samples is acquired.
[0088] In step A4, based on the single-frame encoded image sample and a preset second loss function, the decoder is trained to obtain the trained decoder.
[0089] Among them, the decoder can be constructed by a variety of different machine learning models. In this embodiment, the decoder can be constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model. The structure of the U-Net model can be referred to the relevant content mentioned above and will not be elaborated here. The structure of the decoder in this embodiment can be composed of a U-Net model with a certain number of network layers. Specifically, for example, it can be composed of a U-Net model with 16 or 20 network layers, etc., which can be set according to the actual situation, and this specification embodiment does not make a limitation on this. The second loss function is determined based on the image reconstruction loss information. In actual applications, the image reconstruction loss information can be determined based on minimizing the L2 distance between the multiple-frame reconstructed image samples and the multiple-frame image samples.
[0090] In implementation, the initial architecture of the decoder can be constructed based on the network model structure of the U-Net model. The single-frame encoded image sample is input into the constructed decoder to reconstruct the multiple-frame image samples before encoding processing. Then, based on the reconstructed multiple-frame image samples before encoding processing and the preset second loss function, the corresponding loss information can be determined. Based on the determined loss information, the relevant parameters in the decoder are adjusted. Then, the single-frame encoded image sample is input into the constructed decoder again to reconstruct the multiple-frame image samples before encoding processing. Then, based on the reconstructed multiple-frame image samples before encoding processing and the preset second loss function, the corresponding loss information can be determined. If the determined loss information meets the convergence condition, the trained decoder is output. If the determined loss information does not meet the convergence condition, the relevant parameters in the decoder are adjusted based on the determined loss information, and the decoder is trained again in the above manner until the determined loss information meets the convergence condition.
[0091] After obtaining the encoder and decoder through the above method, the encoder and decoder can also be jointly trained with the biological detection model, so that the biological detection model and the encoder and decoder adapt to each other, achieving a better performance matching effect. For specific details, please refer to the processing steps S512 to S522 described below.
[0092] In step S512, obtain multiple frame image samples for biological detection.
[0093] In step S514, input the multiple frame image samples into the encoder to train the encoder, and obtain a single frame encoded image sample output by the encoder. The single frame encoded image sample meets the preset conditions, and the preset conditions include the conditions for the number of channels of the single frame encoded image sample.
[0094] In step S516, input the single frame encoded image sample into the decoder to train the decoder, and obtain multiple frame reconstructed image samples output by the decoder.
[0095] In step S518, input the multiple frame reconstructed image samples into the biological detection model to train the biological detection model and obtain sample detection results.
[0096] In step S520, if it is determined based on the sample detection results that the encoder, decoder, and biological detection model meet the preset second convergence condition, then obtain the trained biological detection model.
[0097] In step S522, if it is determined based on the sample detection results that the encoder, decoder, and biological detection model do not meet the preset second convergence condition, then re - execute the process of obtaining multiple frame image samples and training the encoder and decoder until the preset convergence condition is met.
[0098] For the specific processing process of the above steps S512 to S522, please refer to the relevant content mentioned above. In addition, the above training process can be implemented using the Stochastic Gradient Descent (SGD) algorithm. Specifically, the jointly trained model can be trained for multiple Epochs using a pre - collected training sample set, and finally, the encoder, decoder, and biological detection model are obtained. In addition, it should be noted that during the training process, the corresponding parameters in the encoder and decoder can be changed through training, and the model parameters in the biological detection model can be the parameters of the specified network layers, such as the parameters of the last 2 network layers in the biological detection model can be changed, etc. Specifically, it can be set according to the actual situation, and the embodiments of this specification do not make limitations on this.
[0099] Through the above processing, a trained encoder, decoder, and biometric detection model can be obtained. The terminal device and the server can perform biometric detection processing on the target user through the above models. For the specific processing, reference can be made to the processing in steps S524 to S532 below.
[0100] In step S524, a single-frame encoded image sent by the terminal device is received. The single-frame encoded image is a single-frame image generated by the terminal device through encoding the acquired multi-frame images based on a pre-trained encoder. The multi-frame images are images collected during the biometric detection of the target user.
[0101] In step S526, the single-frame encoded image is decoded based on a pre-trained decoder to obtain multi-frame reconstructed images corresponding to the multi-frame images.
[0102] In step S528, the multi-frame reconstructed images are input into a pre-trained biometric detection model to perform biometric detection on the target user, and a biometric detection result of the target user is obtained. The biometric detection model is trained based on the images collected during the biometric detection of multiple different users.
[0103] In step S530, the biometric detection result of the target user is sent to the terminal device, and this biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0104] In step S532, the multi-frame reconstructed images are deleted, and the single-frame encoded image is stored.
[0105] Among them, since it is difficult to visually recognize the content in the original image from the single-frame encoded image, storing the single-frame encoded image by the server will not pose a risk of user privacy leakage. In another embodiment, the single-frame encoded image can also be deleted from the server, which can be specifically set according to the actual situation, and the embodiments of this specification do not limit this.
[0106] For the specific processing process of the above steps S502 to S532, reference can be made to the relevant content above, and details are not described herein again.
[0107] An embodiment of this specification provides a biological detection method, which is applied to a server. The terminal device acquires multiple frames of images collected during the biological detection of a target user, and then encodes the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image. The single-frame encoded image is sent to the server to trigger the server to decode the single-frame encoded image, obtain multiple reconstructed images corresponding to the multiple frames of images, and determine the biological detection result of the target user based on the multiple reconstructed images. The terminal device receives the biological detection result of the target user sent by the server and processes the corresponding service requested by the target user based on the biological detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, and then decoding the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth in biological detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of user privacy information.
[0108] Embodiment 5
[0109] The above is the biological detection method provided by the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a biological detection device, as Figure 6 shown.
[0110] The biological detection device includes: an image acquisition module 601, an encoding module 602, an image sending module 603, and a detection result receiving module 604, where:
[0111] The image acquisition module 601 acquires multiple frames of images collected during the biological detection of a target user;
[0112] The encoding module 602 encodes the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, and the encoder is used to encode multiple frames of images into a single frame of image;
[0113] The image sending module 603 sends the single-frame encoded image to the server, and the single-frame encoded image is used to trigger the server to decode the single-frame encoded image, obtain multiple reconstructed images corresponding to the multiple frames of images, and determine the biological detection result of the target user based on the multiple reconstructed images;
[0114] The detection result receiving module 604 receives the biological detection result of the target user sent by the server and processes the corresponding service requested by the target user based on the biological detection result of the target user.
[0115] In the embodiments of this specification, the image sending module 603 includes:
[0116] An image compression unit that compresses the encoded image based on a preset image compression algorithm to obtain a compressed encoded image;
[0117] An image sending unit that sends the compressed encoded image to the server.
[0118] In the embodiments of this specification, the device further includes:
[0119] A sample image acquisition module that acquires multiple frames of sample images for biometric detection;
[0120] An encoder training module that trains the encoder based on the multiple frames of sample images, the conditions that the single-frame image generated by the encoder needs to meet, and a preset first loss function to obtain a trained encoder. The conditions that the single-frame image needs to meet include the conditions of the number of channels of the single-frame image, and the first loss function is determined based on image style loss information and / or image content loss information.
[0121] In the embodiments of this specification, the encoder is constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model.
[0122] In the embodiments of this specification, the device further includes:
[0123] An encoder deployment module that receives the trained encoder sent by the server. The trained encoder is obtained by the server based on multiple frames of sample images for biometric detection, the conditions that the single-frame image generated by the encoder needs to meet, and a preset first loss function.
[0124] An embodiment of this specification provides a biological detection device. By acquiring multiple frames of images collected during the biological detection of a target user, then, based on a pre-trained encoder, encoding the multiple frames of images to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to decode the single-frame encoded image to obtain multiple reconstructed images corresponding to the multiple frames of images, and determining the biological detection result of the target user based on the multiple reconstructed images, receiving the biological detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biological detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then decoding the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth for biological detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0125] Embodiment Six
[0126] Based on the same idea, an embodiment of this specification also provides a biological detection device, as Figure 7 shown.
[0127] The biological detection device includes: an encoded image receiving module 701, a decoding module 702, and a detection result sending module 703, where:
[0128] The encoded image receiving module 701 receives a single-frame encoded image sent by a terminal device. The single-frame encoded image is a single-frame image generated by the terminal device based on a pre-trained encoder to encode multiple frames of images obtained. The multiple frames of images are images collected during the biological detection of a target user;
[0129] The decoding module 702 decodes the single-frame encoded image based on a pre-trained decoder to obtain multiple reconstructed images corresponding to the multiple frames of images;
[0130] The detection result sending module 703 performs biological detection on the target user based on the multiple reconstructed images to obtain the biological detection result of the target user, and sends the biological detection result of the target user to the terminal device. The biological detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0131] In an embodiment of this specification, the device further includes:
[0132] An image processing module deletes the multiple reconstructed images and stores the single-frame encoded image.
[0133] In the embodiments of this specification, the apparatus further includes:
[0134] An encoded image sample acquisition module, which acquires a single-frame encoded image sample obtained after the encoder encodes and processes multiple frames of image samples;
[0135] A first decoder training module, which trains the decoder based on the single-frame encoded image sample and a preset second loss function to obtain a trained decoder, and the second loss function is determined based on image reconstruction loss information.
[0136] In the embodiments of this specification, the apparatus further includes:
[0137] A first sample acquisition module, which acquires multiple frames of image samples for biological detection;
[0138] A first encoder training module, which inputs the multiple frames of image samples into the encoder to train the encoder, and acquires a single-frame encoded image sample output by the encoder, and the single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample;
[0139] A second decoder training module, which inputs the single-frame encoded image sample into the decoder to train the decoder, and acquires multiple frames of reconstructed image samples output by the decoder;
[0140] A processing module, if it is determined based on the multiple frames of reconstructed image samples that the encoder and the decoder meet a preset first convergence condition, stores the decoder, and sends the encoder to the terminal device;
[0141] A first continuous training module, if it is determined based on the multiple frames of reconstructed image samples that the encoder and the decoder do not meet the preset first convergence condition, re-executes acquiring multiple frames of image samples and training the encoder and the decoder until the preset first convergence condition is met.
[0142] In the embodiments of this specification, the preset first convergence condition includes conditions corresponding to a first loss function and the second loss function, and the first loss function is determined based on image style loss information and / or image content loss information.
[0143] In the embodiments of this specification, the image style loss information is determined based on minimizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in the first preset network layer of the VGG network model. The image content loss information is determined based on maximizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in the second preset network layer of the VGG network model. The image reconstruction loss information is determined based on minimizing the distance between the multi-frame reconstructed image sample and the image features corresponding to the multi-frame image sample.
[0144] In the embodiments of this specification, the detection result sending module 703 inputs the multi-frame reconstructed images into a pre-trained biometric detection model to perform biometric detection on the target user and obtain the biometric detection result of the target user. The biometric detection model is trained based on the images collected during the biometric detection of multiple different users.
[0145] In the embodiments of this specification, the apparatus further includes:
[0146] A second sample acquisition module, which acquires multi-frame image samples for biometric detection;
[0147] A second encoder training module, which inputs the multi-frame image samples into the encoder to train the encoder and obtains the single-frame encoded image sample output by the encoder. The single-frame encoded image sample meets the preset conditions, and the preset conditions include the conditions for the number of channels of the single-frame encoded image sample;
[0148] A third decoder training module, which inputs the single-frame encoded image sample into the decoder to train the decoder and obtains the multi-frame reconstructed image sample output by the decoder;
[0149] A model training module, which inputs the multi-frame reconstructed image samples into the biometric detection model to train the biometric detection model and obtain the sample detection result;
[0150] An output module, which obtains the trained biometric detection model if it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model meet the preset second convergence condition;
[0151] A second continuous training module, which, if it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model do not meet the preset second convergence condition, re-executes the acquisition of multi-frame image samples and trains the encoder and the decoder until the preset convergence condition is met.
[0152] An embodiment of this specification provides a biological detection device. By acquiring multiple frames of images collected during the biological detection of a target user, then, based on a pre-trained encoder, encoding the multiple frames of images to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to decode the single-frame encoded image to obtain multiple reconstructed images corresponding to the multiple frames of images, and determining the biological detection result of the target user based on the multiple reconstructed images, receiving the biological detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biological detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then decoding the single-frame encoded image on the server side to restore or reconstruct the original images, the requirement for data transmission bandwidth for biological detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0153] Embodiment Seven
[0154] The above is the biological detection device provided by the embodiments of this specification. Based on the same concept, the embodiments of this specification also provide a biological detection device, as Figure 8 shown.
[0155] The biological detection device may be the terminal device or server provided in the above embodiments, etc.
[0156] The biological detection device may vary greatly due to configuration or performance differences, and may include one or more processors 801 and a memory 802. One or more application programs or data may be stored in the memory 802. Among them, the memory 802 may be transient storage or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the biological detection device. Further, the processor 801 may be set to communicate with the memory 802 and execute a series of computer-executable instructions in the memory 802 on the biological detection device. The biological detection device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, and one or more keyboards 806.
[0157] Specifically, in this embodiment, the biometric detection device includes a memory and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the biometric detection device, and is configured to execute the one or more programs by one or more processors, and the one or more programs include computer-executable instructions for performing the following:
[0158] Obtain multiple frames of images collected during the biometric detection of the target user;
[0159] Based on a pre-trained encoder, encode the multiple frames of images to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single-frame image;
[0160] Send the single-frame encoded image to the server, where the single-frame encoded image is used to trigger the server to decode the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biometric detection result of the target user based on the multiple frames of reconstructed images;
[0161] Receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user.
[0162] In the embodiment of this specification, the sending the single-frame encoded image to the server includes:
[0163] Based on a preset image compression algorithm, compress the encoded image to obtain a compressed encoded image;
[0164] Send the compressed encoded image to the server.
[0165] In the embodiment of this specification, it further includes:
[0166] Obtain multiple frames of sample images for biometric detection;
[0167] Based on the multiple frames of sample images, the conditions that the single-frame image generated by the encoder needs to meet, and a preset first loss function, train the encoder to obtain a trained encoder, where the conditions that the single-frame image needs to meet include the conditions of the number of channels of the single-frame image, and the first loss function is determined based on image style loss information and / or image content loss information.
[0168] In the embodiment of this specification, the encoder is constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model.
[0169] In the embodiments of this specification, it further includes:
[0170] Receiving the trained encoder sent by the server, where the trained encoder is obtained by the server based on the acquired multi-frame sample images for biometric detection, the conditions that the single-frame image generated by the encoder needs to meet, and a preset first loss function.
[0171] In addition, specifically in this embodiment, the biometric detection device includes a memory and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the biometric detection device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions:
[0172] Receiving the single-frame encoded image sent by the terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device after encoding the acquired multi-frame images based on a pre-trained encoder, and the multi-frame images are images collected during the biometric detection of the target user;
[0173] Performing decoding processing on the single-frame encoded image based on a pre-trained decoder to obtain multi-frame reconstructed images corresponding to the multi-frame images;
[0174] Performing biometric detection on the target user based on the multi-frame reconstructed images to obtain the biometric detection result of the target user, and sending the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0175] In the embodiments of this specification, it further includes:
[0176] Deleting the multi-frame reconstructed images and storing the single-frame encoded image.
[0177] In the embodiments of this specification, it further includes:
[0178] Obtaining a single-frame encoded image sample obtained by encoding a multi-frame image sample through the encoder;
[0179] Training the decoder based on the single-frame encoded image sample and a preset second loss function to obtain a trained decoder, where the second loss function is determined based on image reconstruction loss information.
[0180] In the embodiments of this specification, it further includes:
[0181] Obtaining multi-frame image samples for biometric detection;
[0182] Input the multi-frame image samples into the encoder to train the encoder and obtain a single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample;
[0183] Input the single-frame encoded image sample into the decoder to train the decoder and obtain a multi-frame reconstructed image sample output by the decoder;
[0184] If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder meet a preset first convergence condition, store the decoder and send the encoder to the terminal device;
[0185] If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder do not meet the preset first convergence condition, re-execute obtaining the multi-frame image sample and training the encoder and the decoder until the preset first convergence condition is met.
[0186] In the embodiments of this specification, the preset first convergence condition includes conditions corresponding to a first loss function and a second loss function, and the first loss function is determined based on image style loss information and / or image content loss information.
[0187] In the embodiments of this specification, the image style loss information is determined based on minimizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in a first preset network layer of the VGG network model, the image content loss information is determined based on maximizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in a second preset network layer of the VGG network model, and the image reconstruction loss information is determined based on minimizing the distance between the multi-frame reconstructed image sample and the image features corresponding to the multi-frame image sample.
[0188] In the embodiments of this specification, performing biometric detection on the target user based on the multi-frame reconstructed image to obtain a biometric detection result of the target user includes:
[0189] Input the multi-frame reconstructed image into a pre-trained biometric detection model to perform biometric detection on the target user and obtain a biometric detection result of the target user, where the biometric detection model is trained based on images collected during the biometric detection of multiple different users.
[0190] The embodiments of this specification further include:
[0191] Obtain multi-frame image samples for biometric detection;
[0192] Input the multi-frame image samples into the encoder to train the encoder, and obtain a single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample;
[0193] Input the single-frame encoded image sample into the decoder to train the decoder, and obtain a multi-frame reconstructed image sample output by the decoder;
[0194] Input the multi-frame reconstructed image samples into the biometric detection model to train the biometric detection model and obtain a sample detection result;
[0195] If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model meet a preset second convergence condition, then obtain the trained biometric detection model;
[0196] If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model do not meet the preset second convergence condition, then re-execute the process of obtaining multi-frame image samples and training the encoder and the decoder until the preset convergence condition is met.
[0197] An embodiment of this specification provides a biometric detection device. By acquiring multiple frames of images collected during the biometric detection of a target user, then, based on a pre-trained encoder, encoding the multiple frames of images to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to perform decoding processing on the single-frame encoded image to obtain the multiple-frame reconstructed images corresponding to the above multiple frames of images, and determining the biometric detection result of the target user based on the multiple-frame reconstructed images, receiving the biometric detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biometric detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then performing decoding processing on the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth for biometric detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0198] Embodiment Eight
[0199] Further, based on the above Figures 1A to 5For the method shown, one or more embodiments of this specification also provide a storage medium for storing computer-executable instruction information. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following processes can be achieved:
[0200] Obtain multiple frames of images collected during the biometric detection of the target user;
[0201] Perform encoding processing on the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single-frame image;
[0202] Send the single-frame encoded image to the server, where the single-frame encoded image is used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biometric detection result of the target user based on the multiple frames of reconstructed images;
[0203] Receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user.
[0204] In the embodiments of this specification, the sending the single-frame encoded image to the server includes:
[0205] Perform compression processing on the encoded image based on a preset image compression algorithm to obtain a compressed encoded image;
[0206] Send the compressed encoded image to the server.
[0207] In the embodiments of this specification, it further includes:
[0208] Obtain multiple frames of sample images for biometric detection;
[0209] Train the encoder based on the multiple frames of sample images, the conditions that the single-frame image generated by the encoder needs to satisfy, and a preset first loss function. The conditions that the single-frame image needs to satisfy include the conditions of the number of channels of the single-frame image, and the first loss function is determined based on image style loss information and / or image content loss information.
[0210] In the embodiments of this specification, the encoder is constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model.
[0211] In the embodiments of this specification, it further includes:
[0212] Receive the trained encoder sent by the server, where the trained encoder is obtained by the server based on the acquired multi-frame sample images for biometric detection, the conditions that the single-frame images generated by the encoder need to meet, and a preset first loss function.
[0213] In addition, in another specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented:
[0214] Receive a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device after encoding a multi-frame image acquired during the biometric detection of a target user based on a pre-trained encoder;
[0215] Perform decoding processing on the single-frame encoded image based on a pre-trained decoder to obtain multiple reconstructed images corresponding to the multi-frame image;
[0216] Perform biometric detection on the target user based on the multiple reconstructed images to obtain a biometric detection result of the target user, and send the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
[0217] In the embodiments of this specification, it further includes:
[0218] Delete the multiple reconstructed images and store the single-frame encoded image.
[0219] In the embodiments of this specification, it further includes:
[0220] Obtain a single-frame encoded image sample obtained by encoding a multi-frame image sample through the encoder;
[0221] Train the decoder based on the single-frame encoded image sample and a preset second loss function to obtain a trained decoder, where the second loss function is determined based on image reconstruction loss information.
[0222] In the embodiments of this specification, it further includes:
[0223] Obtain multi-frame image samples for biometric detection;
[0224] Input the multi-frame image samples into the encoder to train the encoder, and obtain a single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets preset conditions, and the preset conditions include conditions on the number of channels of the single-frame encoded image sample;
[0225] Input the single-frame encoded image sample into the decoder to train the decoder and obtain the multi-frame reconstructed image sample output by the decoder;
[0226] If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder meet the preset first convergence condition, store the decoder and send the encoder to the terminal device;
[0227] If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder do not meet the preset first convergence condition, re-perform obtaining the multi-frame image sample and training the encoder and the decoder until the preset first convergence condition is met.
[0228] In the embodiments of this specification, the preset first convergence condition includes conditions corresponding to the first loss function and the second loss function, and the first loss function is determined based on image style loss information and / or image content loss information.
[0229] In the embodiments of this specification, the image style loss information is determined based on minimizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in the first preset network layer of the VGG network model, the image content loss information is determined based on maximizing the distance between the features of the single-frame encoded image sample and the multi-frame image sample in the second preset network layer of the VGG network model, and the image reconstruction loss information is determined based on minimizing the distance between the multi-frame reconstructed image sample and the image features corresponding to the multi-frame image sample.
[0230] In the embodiments of this specification, performing biometric detection on the target user based on the multi-frame reconstructed image to obtain the biometric detection result of the target user includes:
[0231] Input the multi-frame reconstructed image into a pre-trained biometric detection model to perform biometric detection on the target user and obtain the biometric detection result of the target user, where the biometric detection model is trained based on images collected during the biometric detection of multiple different users.
[0232] In the embodiments of this specification, it further includes:
[0233] Obtain multi-frame image samples for biometric detection;
[0234] Input the multi-frame image samples into the encoder to train the encoder and obtain the single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets the preset conditions, and the preset conditions include conditions on the number of channels of the single-frame encoded image sample;
[0235] Input the single-frame encoded image sample into the decoder to train the decoder and obtain the multi-frame reconstructed image sample output by the decoder;
[0236] Input the multi-frame reconstructed image sample into the biometric detection model to train the biometric detection model and obtain a sample detection result;
[0237] If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model meet a preset second convergence condition, then obtain the trained biometric detection model;
[0238] If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model do not meet the preset second convergence condition, then re-perform obtaining the multi-frame image sample and training the encoder and the decoder until the preset convergence condition is met.
[0239] An embodiment of this specification provides a storage medium. By acquiring multiple frames of images collected during the biometric detection of a target user, then, based on a pre-trained encoder, encoding the multiple frames of images to generate a single-frame encoded image, and sending the single-frame encoded image to a server to trigger the server to perform decoding processing on the single-frame encoded image to obtain the multi-frame reconstructed images corresponding to the above-mentioned multiple frames of images, and determining the biometric detection result of the target user based on the multi-frame reconstructed images, receiving the biometric detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biometric detection result of the target user. In this way, by encoding multiple frames of images into a single-frame encoded image with a smaller data volume on the terminal device side, but then performing decoding processing on the single-frame encoded image on the server side to restore or reconstruct the original image, the requirement for data transmission bandwidth for biometric detection processing through images is greatly reduced, taking into account both data bandwidth and anti-attack performance. In addition, the data transmitted and stored by this method are all data after encoding processing, which has a strong privacy protection function and can effectively prevent the leakage of users' privacy information.
[0240] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0241] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0242] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0243] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0244] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0245] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0246] Embodiments of the present specification are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable serial-parallel devices for fraud cases to generate a machine, such that the instructions executed by the processor of the computer or other programmable serial-parallel devices for fraud cases generate a device for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0247] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable serial-parallel devices for fraud cases to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0248] These computer program instructions can also be loaded onto a computer or other programmable serial-parallel devices for fraud cases, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0249] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0250] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0251] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0252] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0253] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0254] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0255] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.
[0256] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A biometric detection method applied to a terminal device, the method comprising: Obtaining multiple frames of images collected during the biometric detection of a target user; Performing encoding processing on the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single-frame image; Sending the single-frame encoded image to a server, where the single-frame encoded image is used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determining a biometric detection result of the target user based on the multiple frames of reconstructed images; Receiving the biometric detection result of the target user sent by the server, and processing the corresponding service requested by the target user based on the biometric detection result of the target user.
2. The method according to claim 1, where the sending the single-frame encoded image to the server includes: Performing compression processing on the encoded image based on a preset image compression algorithm to obtain a compressed encoded image; Sending the compressed encoded image to the server.
3. The method according to claim 1, the method further comprising: Obtaining multiple frames of sample images for biometric detection; Training the encoder based on the multiple frames of sample images, the conditions that the single-frame image generated by the encoder needs to satisfy, and a preset first loss function, where the conditions that the single-frame image needs to satisfy include the conditions of the number of channels of the single-frame image, and the first loss function is determined based on image style loss information and / or image content loss information.
4. The method according to claim 1 or 2, where the encoder is constructed by a preset number of U-Net models, and the U-Net model is constructed by a fully connected network model.
5. The method according to claim 1, the method further comprising: Receiving the trained encoder sent by the server, where the trained encoder is obtained by the server based on the multiple frames of sample images for biometric detection, the conditions that the single-frame image generated by the encoder needs to satisfy, and a preset first loss function.
6. A biometric detection method applied to a server, the method comprising: Receiving a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device based on a pre-trained encoder performing encoding processing on multiple frames of images obtained during the biometric detection of a target user; Performing decoding processing on the single-frame encoded image based on a pre-trained decoder to obtain multiple frames of reconstructed images corresponding to the multiple frames of images; Performing biometric detection on the target user based on the multiple frames of reconstructed images to obtain a biometric detection result of the target user, and sending the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
7. The method according to claim 6, the method further comprising: Delete the multi-frame reconstructed image and store the encoded image of the single frame.
8. The method according to claim 6 or 7, wherein the method further comprises: Obtain a single-frame encoded image sample obtained after the encoder encodes a multi-frame image sample; Based on the single-frame encoded image sample and a preset second loss function, train the decoder to obtain a trained decoder, where the second loss function is determined based on image reconstruction loss information.
9. The method according to claim 6 or 7, wherein the method further comprises: Obtain multi-frame image samples for biometric detection; Input the multi-frame image samples into the encoder to train the encoder and obtain a single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample; Input the single-frame encoded image sample into the decoder to train the decoder and obtain a multi-frame reconstructed image sample output by the decoder; If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder meet a preset first convergence condition, store the decoder and send the encoder to the terminal device; If it is determined based on the multi-frame reconstructed image sample that the encoder and the decoder do not meet the preset first convergence condition, re-execute obtaining multi-frame image samples and training the encoder and the decoder until the preset first convergence condition is met.
10. The method according to claim 8, wherein the preset first convergence condition includes conditions corresponding to a first loss function and the second loss function, and the first loss function is determined based on image style loss information and / or image content loss information.
11. The method according to claim 10, wherein the image style loss information is determined based on minimizing the distance between the single-frame encoded image sample and the features of the multi-frame image sample in a first preset network layer of the VGG network model, the image content loss information is determined based on maximizing the distance between the single-frame encoded image sample and the features of the multi-frame image sample in a second preset network layer of the VGG network model, and the image reconstruction loss information is determined based on minimizing the distance between the multi-frame reconstructed image sample and the image features corresponding to the multi-frame image sample.
12. The method according to claim 9, wherein the biometric detection of the target user based on the multi-frame reconstructed image to obtain a biometric detection result of the target user comprises: Input the multi-frame reconstructed image into a pre-trained biometric detection model to perform biometric detection on the target user to obtain a biometric detection result of the target user, where the biometric detection model is trained based on images collected during the biometric detection of multiple different users.
13. The method according to claim 12, wherein the method further comprises: Obtain multi-frame image samples for biometric detection; Input the multi-frame image samples into the encoder to train the encoder and obtain a single-frame encoded image sample output by the encoder, where the single-frame encoded image sample meets a preset condition, and the preset condition includes a condition on the number of channels of the single-frame encoded image sample; Input the single-frame encoded image sample into the decoder to train the decoder and obtain a multi-frame reconstructed image sample output by the decoder; Input the multi-frame reconstructed image samples into the biometric detection model to train the biometric detection model and obtain a sample detection result; If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model meet a preset second convergence condition, obtain the trained biometric detection model; If it is determined based on the sample detection result that the encoder, the decoder, and the biometric detection model do not meet the preset second convergence condition, re-perform the operation of obtaining multi-frame image samples and training the encoder and the decoder until the preset convergence condition is met.
14. A biometric detection device, the device includes: An image acquisition module that acquires multiple frames of images collected during the biometric detection of a target user; An encoding module that performs encoding processing on the multiple frames of images based on a pre-trained encoder to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single frame of image; An image sending module that sends the single-frame encoded image to a server, where the single-frame encoded image is used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biometric detection result of the target user based on the multiple frames of reconstructed images; A detection result receiving module that receives the biometric detection result of the target user sent by the server and processes the corresponding service requested by the target user based on the biometric detection result of the target user.
15. A biometric detection device, the device includes: An encoded image receiving module that receives a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device based on a pre-trained encoder performing encoding processing on acquired multiple frames of images, and the multiple frames of images are images collected during the biometric detection of a target user; A decoding module that performs decoding processing on the single-frame encoded image based on a pre-trained decoder to obtain multiple frames of reconstructed images corresponding to the multiple frames of images; A detection result sending module that performs biometric detection on the target user based on the multiple frames of reconstructed images to obtain the biometric detection result of the target user, and sends the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
16. A biometric detection device, the device includes a trusted execution environment, and the biometric detection device includes: A processor; And A memory arranged to store computer-executable instructions that, when executed, cause the processor: Obtain multiple frames of images collected during the biometric detection of a target user; Based on a pre-trained encoder, perform encoding processing on the multiple frames of images to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single-frame image; Send the single-frame encoded image to a server, where the single-frame encoded image is used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biometric detection result of the target user based on the multiple frames of reconstructed images; Receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user.
17. A biometric detection device, the device includes a trusted execution environment, and the biometric detection device includes: A processor; And A memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to: Receive a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device based on a pre-trained encoder to perform encoding processing on acquired multiple frames of images, and the multiple frames of images are images collected during the biometric detection of a target user; Based on a pre-trained decoder, perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images; Perform biometric detection on the target user based on the multiple frames of reconstructed images to obtain the biometric detection result of the target user, and send the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
18. A storage medium, the storage medium is used to store computer-executable instructions, and the executable instructions, when executed by a processor, implement the following process: Obtain multiple frames of images collected during the biometric detection of a target user; Based on a pre-trained encoder, perform encoding processing on the multiple frames of images to generate a single-frame encoded image, where the encoder is used to encode multiple frames of images into a single-frame image; Send the single-frame encoded image to a server, where the single-frame encoded image is used to trigger the server to perform decoding processing on the single-frame encoded image to obtain multiple frames of reconstructed images corresponding to the multiple frames of images, and determine the biometric detection result of the target user based on the multiple frames of reconstructed images; Receive the biometric detection result of the target user sent by the server, and process the corresponding service requested by the target user based on the biometric detection result of the target user.
19. A storage medium, the storage medium is used to store computer-executable instructions, and the executable instructions, when executed by a processor, implement the following process: Receive a single-frame encoded image sent by a terminal device, where the single-frame encoded image is a single-frame image generated by the terminal device based on a pre-trained encoder to perform encoding processing on acquired multiple frames of images, and the multiple frames of images are images collected during the biometric detection of a target user; Decode the encoded image of the single frame based on a pre-trained decoder to obtain multiple reconstructed images corresponding to the multiple frames; Perform biometric detection on the target user based on the multiple reconstructed images to obtain a biometric detection result of the target user, and send the biometric detection result of the target user to the terminal device, where the biometric detection result is used to trigger the terminal device to process the corresponding service requested by the target user.
Citation Information
Patent Citations
Face feature extraction method, device and equipment
CN111368795A
Image processing method, device and equipment based on privacy protection
CN113223101A