Liveness detection method and apparatus, electronic device, storage medium, and product
By acquiring acoustic signals and video data, and using time-frequency analysis and feature extraction models to determine key frames, the problem of identity forgery in face authentication was solved, achieving higher detection accuracy and efficiency.
Patent Information
- Application Number
- PCT/CN2024/114314
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2024-08-23
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, facial recognition authentication is easily forged by means of photos or electronic screen playback, resulting in inaccurate authentication and reduced security.
By acquiring acoustic signals and video data, key frames in the acoustic signals and video data are determined using time-frequency analysis, and liveness detection is performed by combining a feature extraction model to screen out effective acoustic and video features, thereby improving detection accuracy and efficiency.
It improves the accuracy and robustness of liveness detection, reduces the amount of data processed, and enhances the security of identity verification.
Smart Images

Figure CN2024114314_04122025_PF_FP_ABST
Abstract
Description
Live body detection method and device, electronic equipment, storage medium and product
[0001] The present disclosure claims priority to the Chinese patent application No. 202410684281.0, filed on May 29, 2024, and entitled "Live body detection method and device, electronic equipment, storage medium and product", the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of computer, and particularly relates to a live body detection method and device, electronic equipment, storage medium and product. BACKGROUND
[0003] At present, identity verification is required in various applications of user terminals. With the continuous development of face recognition technology, face identity verification applications are more widely used. However, in the identity verification process, there may be a situation of forging identity features by means of photo, paper or electronic screen playback, which leads to inaccurate identity verification and reduces security. Therefore, live body detection is very important in face identity verification.
[0004] SUMMARY
[0005] The present disclosure provides a live body detection method and device, electronic equipment, storage medium and product.
[0006] In a first aspect, the present disclosure provides a live body detection method, which comprises: in response to a live body detection request, acquiring a first sound wave signal and video data; determining a first key frame in the first sound wave signal according to time-frequency information of the first sound wave signal; determining a second key frame in the video data according to the first key frame; and determining a live body detection result according to the first key frame and the second key frame.
[0007] In a second aspect, the present disclosure provides a live body detection device, which comprises: an acquisition module, configured to acquire a first sound wave signal and video data in response to a live body detection request; a first processing module, configured to determine a first key frame in the first sound wave signal according to time-frequency information of the first sound wave signal; a second processing module, configured to determine a second key frame in the video data according to the first key frame; and a detection module, configured to determine a live body detection result according to the first key frame and the second key frame.
[0008] In a third aspect, the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the living body detection method described above.
[0009] In a fourth aspect, the present disclosure provides a computer-readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the living body detection method described above.
[0010] In a fifth aspect, the present disclosure provides a computer program product comprising computer readable code or a non-transitory computer-readable storage medium carrying computer readable code, which, when run in a processor of an electronic device, causes the processor in the electronic device to perform the living body detection method described above.
[0011] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, which together with the embodiments of the present disclosure serve to explain the present disclosure, and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent from the detailed description of the specific example embodiments, taken in conjunction with the accompanying drawings, in which:
[0013] FIG. 1 is an application scenario provided by an embodiment of the present disclosure;
[0014] FIG. 2 is a flowchart of a living body detection method provided by an embodiment of the present disclosure;
[0015] FIG. 3 is a logic block diagram of a living body detection method in an embodiment of the present disclosure;
[0016] FIG. 4 is a block diagram of a living body detection apparatus provided by an embodiment of the present disclosure;
[0017] FIG. 5 is a block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] For those skilled in the art to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below in conjunction with the drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered only as exemplary. Therefore, those skilled in the art should realize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0019] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.
[0020] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0021] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise" and / or "consist of", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms "connected" or "coupled" and / or similar terms are not limited to a physical or mechanical connection, but can include an electrical connection, whether direct or indirect.
[0022] Unless otherwise defined, all terms used herein, including technical and scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in a commonly used dictionary should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0023] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs. The use of user data in the technical solutions complies with relevant national laws and regulations (for example, "Information Security Technology Personal Information Security Specification" and the like). For example, appropriate measures are taken for personal information access control; restrictions are given for the display of personal information; the use purpose of personal information does not exceed the direct or reasonably related range; the identity of the specific individual is eliminated when personal information is used to avoid precise positioning to a specific individual.
[0024] In the related art, the living body detection is mainly based on two-dimensional image pixel texture analysis to detect the non-living body forgery, but this method is limited by light, pixel texture and other factors, and has low accuracy.
[0025] Therefore, the present disclosure provides a living body detection method, which acquires a first sound wave signal and collected video data, and then fuses the sound wave and the video for living body detection, thereby improving the accuracy and detection capability of the living body detection. The first sound wave signal and the video data are processed to determine the first key frame and the second key frame for living body detection, which can filter out more effective sound wave and video related features, further improve the detection accuracy, reduce the data amount of the detection processing, and improve the detection efficiency.
[0026] FIG. 1 schematically shows an application scenario of a living body detection method and device provided by an embodiment of the present disclosure.
[0027] As shown in FIG. 1, the application scenario of the embodiment of the present disclosure can include a terminal device 101, a network 103 and a server 102. The network 103 is a medium for providing a communication link between the terminal device 101 and the server 102. The network 103 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0028] A user can use the terminal device 101 to interact with the server 102 through the network 103 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0029] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.
[0030] The server 102 can be a server providing various services, such as a background management server supporting the website browsed by the user using the terminal device 101 (only as an example). The background management server can analyze and process the received user request data, etc., and feed back the processing result (such as a webpage, information or data generated according to the user request, etc.) to the terminal device.
[0031] It should be noted that the living body detection method and device provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the living body detection method and device provided in the embodiments of the present disclosure can be arranged in the server 102. The living body detection method and device provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102. Accordingly, the living body detection method and device provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102.
[0032] In addition, the living body detection method and device in the embodiments of the present disclosure can also be executed by the terminal device 101, and the embodiments of the present disclosure do not limit this.
[0033] It should be understood that the number of terminal devices, networks and servers in FIG. 1 is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks and servers.
[0034] FIG. 2 is a flowchart of a living body detection method provided in the embodiments of the present disclosure. Referring to FIG. 2, the method comprises:
[0035] S201, in response to a living body detection request, obtaining a first sound wave signal, and obtaining video data.
[0036] Living body detection means detection of a real physiological feature of an object in some identity verification scenarios, for example, in a face identity recognition application, living body detection is mainly used to determine whether an object is a real living body user before face identity recognition. In the embodiments of the present disclosure, information collection can be performed on a user after obtaining authorization consent of the user, for example, when the user operates payment, personal information query and other applications.
[0037] A possible implementation for the above step S201 is provided, a sound wave signal is emitted to a target object at a preset emission frequency, and a first sound wave signal reflected back after the sound wave signal reaches the target object is received, and video data for the target object is collected synchronously; in response to recognizing that the target object performs a preset behavior, the video data collection and the first sound wave signal reception are stopped, and the first sound wave signal and the video data of the target object are obtained.
[0038] The emitted acoustic wave signal may be, for example, an ultrasonic signal or other audio signal, and the present embodiment is not limited in this regard. For example, when the target object operates an application in the terminal device and triggers identity verification, after authorization by the target object, the application may prompt the target object to perform a nodding, shaking or other preset action, and simultaneously turn on the loudspeaker of the terminal device to emit an ultrasonic signal to the target object at a preset emission frequency, while simultaneously recording the first acoustic wave signal and video data. When the correct action is recognized, the recording is stopped, and the synchronously recorded first acoustic wave signal and video data are obtained. The first acoustic wave signal may be received by the microphone in the terminal device, and the video data may be captured by the camera in the terminal device. In this way, the first acoustic wave signal and video data required for living body detection can be obtained based on the existing loudspeaker, microphone and camera in the terminal device, without the need for additional hardware devices or reliance on special hardware. This saves resources and is simple to implement.
[0039] A possible implementation of the above step S201 is provided, in which an acoustic wave signal is sent to the target object, and a reflected signal emitted back by the target object after the acoustic wave signal reaches the target object is received. The reflected signal is filtered based on characteristic information of the acoustic wave signal sent to the target object, and a first acoustic wave signal is obtained.
[0040] The characteristic information of the acoustic wave signal may include, but is not limited to, the wavelength of the acoustic wave signal, the signal amplitude of the acoustic wave signal, the frequency of the acoustic wave signal, and the like. The first acoustic wave signal may be obtained by filtering out signals in the reflected signal that are not characteristic information, for example, by filtering out reflected signals whose wavelengths are not the wavelength of the acoustic wave signal.
[0041] The characteristic information of the acoustic wave signal may include, but is not limited to, the wavelength of the acoustic wave signal, the signal amplitude of the acoustic wave signal, the frequency of the acoustic wave signal, and the like. The first acoustic wave signal may be obtained by filtering out signals in the reflected signal that are not characteristic information, for example, by filtering out reflected signals whose wavelengths are not the wavelength of the acoustic wave signal.
[0042] It should be noted that in the present embodiment, the collection of the first acoustic wave signal and the video data is performed synchronously, which is to ensure that the first acoustic wave signal and the video data are time-aligned, so as to facilitate the determination of the second key frame, improve the processing efficiency. However, the collection may not be strictly synchronous, and the first acoustic wave signal and the video data may be time-aligned after being obtained, and the present embodiment is not limited in this regard.
[0043] In addition, for example, the living body detection method in the present embodiment may be performed by the terminal device, and the first acoustic wave signal and the video data may be collected by the loudspeaker, microphone and camera in the terminal device, and then the living body detection may be performed by the terminal device based on the first acoustic wave signal and the video data.
[0044] For example, the living body detection method in the embodiments of the present disclosure can be performed by a server, and after the terminal device collects the first sound wave signal and the video data, the terminal device sends the first sound wave signal and the video data to the server, and then the server performs living body detection according to the first sound wave signal and the video data. For another example, in the embodiments of the present disclosure, after the terminal device collects the first sound wave signal and the video data, the terminal device can first perform simple data processing and analysis, such as time-frequency analysis, determination of the first key frame and the second key frame, and then send the first key frame and the second key frame to the server, and then the server performs living body detection according to the first key frame and the second key frame.
[0045] In S202, a first key frame in the first sound wave signal is determined according to time-frequency information of the first sound wave signal.
[0046] For this step S202, the present disclosure provides possible implementation manners, including:
[0047] 1) The time-frequency information of the first sound wave signal is divided into a plurality of time-frequency slices.
[0048] Specifically, in this step, the target time-frequency information can be divided into a plurality of time-frequency slices according to a preset time interval. The preset time interval can be set according to the accuracy requirement of the first key frame. For example, if the accuracy requirement of the first key frame is higher, the preset time interval is shorter.
[0049] For example, the target time-frequency information is a time-frequency graph, which can be represented by a two-dimensional array. According to the preset time interval, the length of the time slice corresponding to the time-frequency graph can be calculated. For example, if the first sound wave signal is 2.4s and the preset time interval is 0.1s, the time dimension corresponding to each time-frequency slice is 0.1s, and the total length of the time-frequency graph in the horizontal time dimension is 80, then the length of each time-frequency slice can be determined as 3, and then the time-frequency graph can be divided into a plurality of time-frequency slices according to the length 3 (corresponding to the time dimension of 0.3s), and a time-frequency slice sequence including a plurality of time-frequency slices sorted in time order is obtained.
[0050] 2) The signal energy value of each time-frequency slice is determined.
[0051] For example, for the time-frequency graph, each row corresponds to the spectral energy value of a frequency, that is, the signal energy value, and for each time-frequency slice, the spectral energy value of all frequency elements in the time-frequency slice can be calculated.
[0052] 3) The first key frame in the first sound wave signal is determined according to the signal energy value of each time-frequency slice.
[0053] In this step, in one possible implementation, according to the signal energy value of each time-frequency slice, time-frequency slices with a signal energy value greater than or equal to a preset threshold value can be screened out, and a first key frame of the first sound wave signal can be obtained according to the screened time-frequency slices.
[0054] The preset threshold value can be set according to actual requirements and experience, and is not limited in the embodiments of the present disclosure. It can be understood that a higher signal energy can represent the degree of Doppler shift caused by facial motion, and the greater the degree of frequency shift, the more obvious the corresponding facial motion feature, and the more effective features representing the target object. Therefore, in the embodiments of the present disclosure, time-frequency slices with an energy value greater than or equal to the preset threshold value can be screened out, which can be considered as key time-frequency slices, and a first key frame in the first sound wave signal can be determined according to the time corresponding to the key time-frequency slices.
[0055] In addition, in the embodiments of the present disclosure, the time-frequency information of the first sound wave signal can be obtained by performing time-frequency analysis on the first sound wave signal, and possible implementation manners are further provided, including:
[0056] If the first signal length of the first sound wave signal is less than a preset minimum signal length, a second signal length is determined according to the first signal length and the preset minimum signal length, and the second signal length is equal to the difference between the preset minimum signal length and the first signal length. If the second signal length is less than or equal to the first signal length, time-frequency information with a length equal to the second signal length is cut from the time-frequency information of the first sound wave signal as extended time-frequency information, and the time-frequency information of the first sound wave signal and the extended time-frequency information are sequentially merged to obtain updated time-frequency information of the first sound wave signal. The preset minimum signal length can be set according to the minimum length required by time-frequency analysis.
[0057] For example, the first signal length of the first sound wave signal is 0.85s, and the preset minimum signal length is 1s, so the second signal length is 0.15s. The second signal length 0.15s is less than the first signal length 0.85s, and time-frequency information with a length of 0.15s is cut from the time-frequency information (0-0.85s corresponding time-frequency information) of the first sound wave signal as extended time-frequency information (for example, 0.7s-0.85s corresponding time-frequency information), and the time-frequency information (0-0.85s corresponding time-frequency information) of the first sound wave signal and the extended time-frequency information (0.7s-0.85s corresponding time-frequency information) are sequentially merged to obtain the updated time-frequency information of the first sound wave signal.
[0058] The embodiment can ensure that the signal length of the updated first sound wave signal meets the minimum signal length required by time-frequency analysis, and since the expanded time-frequency information is part of the time-frequency information of the first sound wave signal, the rationality of the expansion of the first sound wave signal is improved.
[0059] In addition, in the embodiment of the present disclosure, the time-frequency information of the first sound wave signal can be obtained by performing time-frequency analysis on the first sound wave signal, and possible implementation manners are further provided, including:
[0060] If the second signal length is greater than the first signal length, the sound wave signal is retransmitted, and the first sound wave signal transmitted by the retransmitted sound wave signal is received. If the re-received first sound wave signal is greater than or equal to the preset minimum signal length, the reception of the first sound wave signal is stopped, and the time-frequency information of the re-received first sound wave signal is taken as the time-frequency information of the first sound wave signal.
[0061] In the embodiment, when the second signal length is greater than the first signal length, the first sound wave signal is reacquired, which can avoid that the first sound wave signal includes multiple repeated sound wave signals, and thus the authenticity of the first sound wave signal can be improved by reacquiring the first sound wave signal.
[0062] In addition, in the embodiment of the present disclosure, the time-frequency information of the first sound wave signal can be obtained by performing time-frequency analysis on the first sound wave signal, and possible implementation manners are further provided, including:
[0063] 1) The first sound wave signal is expanded to obtain a second sound wave signal.
[0064] In this step, specifically, the first sound wave signal can be expanded according to the corresponding transmission frequency of the first sound wave signal and the preset sampling rate to obtain the second sound wave signal.
[0065] In the embodiment of the present disclosure, in the face identity verification scene, the recognition of the target object performing a preset behavior can require a relatively short time, and since the first sound wave signal and the video data are synchronously collected, the obtained first sound wave signal can be relatively short. Therefore, in the embodiment of the present disclosure, the first sound wave signal can be expanded to achieve a higher frequency resolution and improve the accuracy and effectiveness of time-frequency analysis.
[0066] For example, the first sound wave signal initially acquired is a one-dimensional array S[1:n], and the minimum signal length required by time-frequency analysis is N, and then the signal can be extended based on S[1:n], and the second sound wave signal S[n+1:N] extended can be: S[t]=mean(S[1:n])*sin(2*PI*f*t) / sample_rate
[0067] wherein t takes a value from n+1 to N, f is the transmission frequency, and sample_rate is the preset sampling rate.
[0068] In this way, by extending the first sound wave signal, the length of the first sound wave signal can be extended from n to N, and the accuracy of subsequent time-frequency analysis can be improved.
[0069] 2) From the time-frequency information of the second sound wave signal, a first time-frequency information segment is intercepted, and a second time-frequency information segment in the time-frequency information of the second sound wave signal is filled according to the first time-frequency information segment, to obtain the time-frequency information of the first sound wave signal.
[0070] For example, based on short-time Fourier transform, the time-frequency information of the second sound wave signal, such as a time-frequency graph, is calculated, and in the time-frequency graph determined directly based on the second sound wave signal, short-time strong pulses can appear, which affect the extraction of the first key frame, and therefore, in the embodiment of the present disclosure, to eliminate the strong pulses, the time-frequency information of the second sound wave signal can also be processed, the first time-frequency information segment is intercepted from the time-frequency graph of the second sound wave signal, and the second time-frequency information segment in the time-frequency graph of the second sound wave signal is filled, for example, the first sound wave signal is 0.85s, and the part of 0-0.85s in the initial time-frequency graph is intercepted to obtain the first time-frequency information segment, and then the first time-frequency information segment is mirrored and filled after 0.85s in the initial time-frequency graph, to replace the original part after 0.85s in the second sound wave signal.
[0071] In this way, by extending and intercepting the first sound wave signal acquired, the time-frequency information of the first sound wave signal is obtained, which not only can have more original effective features of the first sound wave signal, but also can reduce the influence of strong pulse noise, thereby improving the accuracy.
[0072] S203, determining a second key frame in the video data according to the first key frame.
[0073] In the embodiments of the present disclosure, the second key frame can be extracted in combination with the first key frame. The present disclosure provides possible implementation manners, including: determining a video segment that is aligned in time with the first key frame from the video data; and obtaining the second key frame from the video segment; wherein the first sound wave signal is aligned in time with the video data.
[0074] In the embodiments of the present disclosure, the time-frequency slice whose signal energy value is greater than or equal to the preset threshold value, for example, the time interval of each time-frequency slice when being divided is 0.1s, can correspond to the first key frame in the 0.1s time period. According to the time information corresponding to the first key frame, the corresponding video segment can be determined, or the corresponding video segment can also be directly determined according to the time information of the screened time-frequency slice.
[0075] It should be noted that in the embodiments of the present disclosure, when determining the corresponding video segment, the time of the first sound wave signal and the video data needs to be aligned, that is, synchronized. The synchronization can be achieved by controlling the synchronization during data acquisition, or the alignment can be achieved after acquisition by using the acquisition time, and the like, which is not limited. The first key frame corresponding video segment can be determined by using the aligned time information, so as to improve the accuracy. Furthermore, since the screened first key frame can represent the frame with object activity, the video segment determined according to the first key frame can also contain more effective object features.
[0076] In the embodiments of the present disclosure, the time-frequency slice whose signal energy value is greater than or equal to the preset threshold value, for example, the time interval of each time-frequency slice when being divided is 0.1s, can correspond to the first key frame in the 0.1s time period. According to the time information corresponding to the first key frame, the corresponding video segment can be determined, or the corresponding video segment can also be directly determined according to the time information of the screened time-frequency slice.
[0077] In one possible implementation manner, the video data is the data collected for the face of the target object, and obtaining the second key frame from the video segment includes:
[0078] 1) determining the face posture presented by each video frame in the video segment.
[0079] In some possible implementation manners, this step determines the face key points included in the video frame, and determines the face posture presented by the video frame according to the face key points.
[0080] 2) selecting the video frame that meets the posture condition from the video segment based on the face posture presented by each video frame in the video segment, to obtain the second key frame.
[0081] In the embodiments of the present disclosure, the face key points in the video frame are detected, which represent the key region positions of the face, for example, including eyebrows, eyes, nose, mouth, etc. According to the position information of these key points, the face posture can be determined. Furthermore, in the embodiments of the present disclosure, some video frames that meet the posture condition are screened from the face posture of each video frame in the video segment, to be used as the second key frame.
[0082] The pose condition is, for example, that the face poses corresponding to the selected video frames are different. For example, the face key points are identified as eyes and a mouth, and then the face pose is estimated by calculating the distance from the eyes to the mouth. Different video frames are selected from the plurality of video frames included in the video segment according to the distance from the eyes to the mouth, as the second key frames. In this way, the selected second key frames can cover different face poses as much as possible, and the accuracy of the living body detection can be improved.
[0083] In another possible implementation, the second key frames are obtained from the video segment, including: selecting a preset number of video frames from the video segment to obtain the second key frames.
[0084] For example, one or more video frames can be randomly selected from the determined video segment in the embodiment of the present disclosure, or all the video frames included in the video segment can be used as the second key frames. The preset number is not limited in the embodiment of the present disclosure, and can be set according to experience and requirements.
[0085] In this way, in the embodiment of the present disclosure, the second key frames in the video data are determined in combination with the first key frames, which can improve the accuracy of the determination of the second key frames, and the second key frames are used for subsequent living body detection, which can not only improve the accuracy, but also reduce the amount of processing data and improve the efficiency.
[0086] S204, determining a living body detection result according to the first key frames and the second key frames.
[0087] When the step S204 is performed, the following steps are specifically included:
[0088] 1) determining a first feature of the first key frames, and determining a second feature of the second key frames.
[0089] In the embodiment of the present disclosure, the first feature can be obtained by performing feature extraction on the first key frames based on a trained first feature extraction model, for example, a long short-term memory (LSTM). Alternatively, the first feature can be obtained by directly performing feature extraction on the selected key time-frequency slices.
[0090] The second feature can be obtained by performing feature extraction on the second key frames based on a trained second feature extraction model, for example, a 3D neural network, to improve the living body detection capability.
[0091] 2) fusing the first feature and the second feature to obtain a fused feature.
[0092] In an embodiment, the first feature and the second feature are spliced to obtain the fused feature.
[0093] In another embodiment, the first feature and the second feature are weighted based on pre-trained weights to obtain a fusion feature. The pre-trained weights can be determined in a training process of the living body detection model.
[0094] 3) determining a living body detection result of the target object based on the fusion feature.
[0095] In the embodiments of the present disclosure, the living body detection result of the target object can be obtained by analyzing the fusion feature based on the trained living body detection model.
[0096] For example, the first feature and the second feature are spliced and input into the living body detection model, and through classification and recognition, the detection result that the target object is a living body or a non-living body can be output, wherein the living body detection model is a multi-layer linear classifier, for example.
[0097] In the embodiments of the present disclosure, when performing living body detection, the first sound wave signal and the video data are acquired, the first key frame is determined according to the time-frequency information of the first sound wave signal, the second key frame in the video data is determined according to the first key frame, and the living body detection result is determined according to the first key frame and the second key frame. In this way, the features of two modalities of video and sound wave can be fused to improve the accuracy and robustness of living body detection, and the screening of the first key frame and the second key frame can reduce the data processing amount, the living body detection is performed based on more effective features, and the accuracy of the living body detection can be further improved.
[0098] The training process of the model used for the extraction of the first feature and the second feature and the living body detection in the embodiments of the present disclosure will be described below.
[0099] Taking the extraction of the first feature based on the first feature extraction model, the extraction of the second feature based on the second feature extraction model, and the living body detection based on the living body detection model as an example, and taking a living body as a real person as an example, in the embodiments of the present disclosure, in order to improve the accuracy, the three models can be trained at the same time. In one possible embodiment, the training process can be as follows:
[0100] 1) acquiring a positive sample set and a negative sample set.
[0101] The positive sample set includes the first sound wave signal positive sample of the object with the classification label of living body and the video data positive sample, and the negative sample set includes the first sound wave signal negative sample of the object with the classification label of non-living body and the video data negative sample. The non-living body can include various non-real person attack scenes, such as photos, electronic screen playback, etc.
[0102] For example, based on the application of face identity verification, the first sound wave signal positive sample and the first sound wave signal negative sample can be obtained through the reflection and reception of the sound wave signal, and the video data positive sample and the video data negative sample can be obtained through synchronous video acquisition.
[0103] 2) Determine the first key frame and the second key frame for the positive sample set and the negative sample set, respectively.
[0104] The specific manner of the first key frame and the second key frame is the same as the above embodiment, which will not be described here.
[0105] 3) Based on the first feature extraction model, the first feature is extracted for the first sound wave signal positive sample and the first sound wave signal negative sample, and based on the second feature extraction model, the second feature is extracted for the video data positive sample and the video data negative sample.
[0106] For each object in the positive sample set and the negative sample set, the first feature and the second feature corresponding to the object are spliced and input into the living body detection model to obtain a predicted living body detection result, and a cross-entropy loss function is calculated according to the predicted living body detection result and the corresponding classification label. Based on the cross-entropy loss function, the network model weight of the living body detection model, the first feature extraction model and the second feature extraction model is updated to continuously train the model until the loss function converges or reaches a preset iteration number.
[0107] In this way, in the embodiment of the disclosure, the model is trained through the positive sample set and the negative sample set, the scene is covered more, the training accuracy is improved, and then in actual application, the trained living body detection model, the first feature extraction model and the second feature extraction model can be used for living body detection to improve the living body detection accuracy.
[0108] Based on the above embodiment, the overall implementation logic of the living body detection method in the embodiment of the disclosure is described, and referring to FIG. 3, a logic block diagram of the living body detection method in the embodiment of the disclosure is shown.
[0109] 1) As shown in FIG. 3, when face identity verification is performed, the terminal device can prompt to perform a preset action, and simultaneously turn on the loudspeaker of the terminal device to emit sound wave signals to the target object, and simultaneously record the first sound wave signal received by the microphone and the video data collected by the camera.
[0110] The first sound wave signal is subjected to time-frequency analysis, and the first key frame in the first sound wave signal is determined. According to the first key frame, the second key frame in the video data is determined.
[0111] In addition, in the embodiment of the present disclosure, simple live body pre-judgment can be performed in the terminal device according to the first sound wave signal and / or the video data, and many simple attack non-live body scenes can be filtered out, such as no face motion in front of the screen, simple shaking of the terminal device to simulate face motion, and the like. Thus, some simple live body detection can be initially performed, and for the detected non-live body, no further detection through subsequent feature extraction and model is required, the cost of sending simple invalid data to the server can be saved, and the processing efficiency can be improved.
[0112] 2) performing feature extraction on the first key frame to obtain first features, and performing feature extraction on the second key frame to obtain second features.
[0113] 3) splicing the first features and the second features to obtain fused features, and inputting the fused features into a live body detection model to obtain a live body detection result.
[0114] Further, when the live body detection result indicates that the object is a live body, face identity recognition can be performed according to the video data and the like to determine whether the object identity verification is passed.
[0115] In the embodiment of the present disclosure, the sound wave and the video are fused to enhance the feature representation of live body detection, improve the accuracy and capability of live body detection, and perform screening on the first key frame and the second key frame for subsequent live body detection. The detection accuracy can be further improved, and the detection efficiency can be improved.
[0116] It can be understood that the above-mentioned various method embodiments of the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Limited by the length, the present disclosure will not be repeated. Those skilled in the art can understand that the specific execution order of each step in the above-mentioned method should be determined according to its function and possible internal logic.
[0117] In addition, the present disclosure also provides a live body detection device, an electronic device, and a computer readable storage medium, which can be used to implement any one of the live body detection methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated.
[0118] FIG. 4 is a block diagram of a live body detection device provided by an embodiment of the present disclosure.
[0119] Referring to FIG. 4, the present disclosure provides a live body detection device, which comprises:
[0120] The acquisition module 41 is configured to acquire a first sound wave signal and video data in response to a live body detection request.
[0121] The first processing module 42 is configured to determine a first key frame in the first sound wave signal according to time-frequency information of the first sound wave signal.
[0122] The second processing module 43 is configured to determine a second key frame in the video data according to the first key frame.
[0123] The detection module 44 is configured to determine a living body detection result according to the first key frame and the second key frame.
[0124] In a possible implementation, when determining the first key frame in the first sound wave signal according to the time-frequency information of the first sound wave signal, the first processing module 42 is configured to:
[0125] divide the time-frequency information of the first sound wave signal into a plurality of time-frequency slices;
[0126] determine a signal energy value of each of the time-frequency slices;
[0127] determine the first key frame in the first sound wave signal according to the signal energy value of each of the time-frequency slices.
[0128] In a possible implementation, the acquisition module 41 is further configured to:
[0129] transmit a sound wave signal and receive a reflection signal corresponding to the sound wave signal;
[0130] filter the reflection signal according to characteristic information of the sound wave signal to obtain the first sound wave signal.
[0131] In a possible implementation, the first processing module 42 is further configured to:
[0132] if a first signal length of the first sound wave signal is less than a preset minimum signal length, determine a second signal length according to the first signal length and the preset minimum signal length;
[0133] if the second signal length is less than or equal to the first signal length, cut time-frequency information with a length equal to the second signal length from the time-frequency information of the first sound wave signal as extended time-frequency information;
[0134] merge the time-frequency information of the first sound wave signal and the extended time-frequency information to obtain the time-frequency information of the first sound wave signal.
[0135] In a possible implementation, the first processing module 42 is further configured to:
[0136] if the second signal length is greater than the first signal length, retransmit a sound wave signal and receive a first sound wave signal emitted by the retransmitted sound wave signal;
[0137] If the re-received first acoustic wave signal is greater than or equal to the preset minimum signal length, time-frequency information of the re-received first acoustic wave signal is taken as the time-frequency information of the first acoustic wave signal.
[0138] In a possible embodiment, the first processing module 42 is further configured to:
[0139] extend the first acoustic wave signal to obtain a second acoustic wave signal;
[0140] cut a first time-frequency information segment from the time-frequency information of the second acoustic wave signal, and fill a second time-frequency information segment in the time-frequency information of the second acoustic wave signal according to the first time-frequency information segment to obtain the time-frequency information of the first acoustic wave signal.
[0141] In a possible embodiment, the first acoustic wave signal is aligned with the video data in time, and when the second key frame in the video data is determined according to the first key frame, the second processing module 43 is configured to:
[0142] determine a video segment aligned with the first key frame in time from the video data;
[0143] obtain the second key frame from the video segment.
[0144] In a possible embodiment, the video data is data collected for a face, and when the second key frame is obtained from the video segment, the second processing module 43 is configured to:
[0145] determine a face posture presented by each video frame in the video segment;
[0146] select a video frame meeting a posture condition from the video segment based on the face posture presented by each video frame in the video segment to obtain the second key frame.
[0147] In a possible embodiment, when the face posture presented by each video frame in the video segment is determined, the second processing module 43 is configured to:
[0148] determine a face key point included in the video frame;
[0149] determine the face posture presented by the video frame according to the face key point.
[0150] In a possible embodiment, when the living body detection result is determined according to the first key frame and the second key frame, the detection module 44 is configured to:
[0151] determine a first feature of the first key frame, and determine a second feature of the second key frame;
[0152] fuse the first feature and the second feature to obtain a fused feature;
[0153] determine a living body detection result based on the fused feature.
[0154] In a possible embodiment, when the first feature and the second feature are fused to obtain the fused feature, the detection module 44 is configured to:
[0155] splice the first feature and the second feature to obtain the fused feature; or
[0156] perform weighted operation on the first feature and the second feature based on pre-trained weights to obtain the fused feature.
[0157] The modules in the living body detection apparatus can be implemented in whole or in part by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.
[0158] FIG. 5 is a block diagram of an electronic device according to an embodiment of the present disclosure.
[0159] Referring to FIG. 5, the electronic device according to an embodiment of the present disclosure includes at least one processor 501, at least one memory 502, and one or more I / O interfaces 503 connected between the processor 501 and the memory 502. The memory 502 stores one or more computer programs executable by the at least one processor 501. The one or more computer programs are executed by the at least one processor 501 to enable the at least one processor 501 to perform the living body detection method described above.
[0160] The modules in the electronic device can be implemented in whole or in part by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.
[0161] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the living body detection method described above. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0162] The embodiment of the present disclosure further provides a computer program product, comprising computer readable code or a nonvolatile computer readable storage medium carrying computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device performs the living body detection method.
[0163] Those of ordinary skill in the art understand that all or some of the steps in the method disclosed above, the functions of the modules / units in the system and the device can be implemented as software, firmware, hardware and appropriate combinations thereof. In the hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable storage medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media).
[0164] As known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable program instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Furthermore, it is known to those of ordinary skill in the art that communication media typically includes computer readable program instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery medium.
[0165] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0166] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0167] The computer program product described herein can be embodied in a tangible computer readable storage medium, or embodied as a software product, such as a software development kit (SDK), and the like.
[0168] The computer program product described herein can be embodied in a tangible computer readable storage medium, or embodied as a software product, such as a software development kit (SDK), and the like.
[0169] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage viruses or other code or instructions implementing a functionally equivalent process, such that the instructions, defining functions described by the flowchart and / or block diagram block or blocks are stored in the computer readable storage medium.
[0170] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0171] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0172] Example embodiments have been disclosed herein and, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that features, characteristics or elements described with reference to one particular embodiment can be used, combined or modified for use with other embodiments, unless specifically noted otherwise. Accordingly, it will be understood by those skilled in the art that various changes in form and details can be made without departing from the scope of the disclosure as set forth in the appended claims.
Claims
1. A method for live body detection, comprising: obtaining a first acoustic signal and video data in response to a live body detection request; determining a first key frame in the first acoustic signal according to time-frequency information of the first acoustic signal; determining a second key frame in the video data according to the first key frame; determining a live body detection result according to the first key frame and the second key frame. 2.The method of claim 1, wherein the determining the first key frame in the first acoustic signal according to time-frequency information of the first acoustic signal comprises: segmenting the time-frequency information of the first acoustic signal into a plurality of time-frequency slices; determining a signal energy value of each of the time-frequency slices; determining the first key frame in the first acoustic signal according to the signal energy value of each of the time-frequency slices. 3.The method of claim 1, further comprising: sending an acoustic signal and receiving a reflection signal corresponding to the acoustic signal; filtering the reflection signal according to characteristic information of the acoustic signal to obtain the first acoustic signal. 4.The method of claim 1, further comprising: if a first signal length of the first acoustic signal is less than a preset minimum signal length, determining a second signal length according to the first signal length and the preset minimum signal length; if the second signal length is less than or equal to the first signal length, cutting time-frequency information with a length equal to the second signal length from the time-frequency information of the first acoustic signal as extended time-frequency information; merging the time-frequency information of the first acoustic signal and the extended time-frequency information to obtain the time-frequency information of the first acoustic signal. 5.The method of claim 4, further comprising: if the second signal length is greater than the first signal length, re-sending an acoustic signal and receiving a first acoustic signal emitted by the re-sent acoustic signal; if the re-received first acoustic signal is greater than or equal to the preset minimum signal length, taking the time-frequency information of the re-received first acoustic signal as the time-frequency information of the first acoustic signal. 6.The method of claim 1, further comprising: extending the first acoustic signal to obtain a second acoustic signal; cutting a first time-frequency information segment from the time-frequency information of the second acoustic signal, and complementing a second time-frequency information segment in the time-frequency information of the second acoustic signal according to the first time-frequency information segment to obtain the time-frequency information of the first acoustic signal. 7.The method of claim 1, wherein the first acoustic signal is aligned in time with the video data, and the determining the second key frame in the video data according to the first key frame comprises: determining a video segment aligned in time with the first key frame from the video data; obtaining the second key frame from the video segment.
8. The method of claim 7, wherein the video data is face data, and the obtaining the second key frame from the video clip comprises: determining a face pose presented in each video frame of the video clip; and selecting a video frame that meets a pose condition from the video clip based on the face pose presented in each video frame of the video clip, to obtain the second key frame.
9. The method of claim 8, wherein the determining the face pose presented in each video frame of the video clip comprises: determining face key points included in the video frame; and determining the face pose presented in the video frame based on the face key points.
10. The method of any one of claims 1-9, wherein the determining the live body detection result based on the first key frame and the second key frame comprises: determining a first feature of the first key frame, and determining a second feature of the second key frame; fusing the first feature and the second feature to obtain a fused feature; and determining the live body detection result based on the fused feature.
11. The method of claim 10, wherein the fusing the first feature and the second feature to obtain the fused feature comprises: splicing the first feature and the second feature to obtain the fused feature; or performing weighted operation on the first feature and the second feature based on pre-trained weights to obtain the fused feature.
12. A live body detection apparatus, comprising: an obtaining module configured to obtain a first sound wave signal and video data in response to a live body detection request; a first processing module configured to determine a first key frame in the first sound wave signal based on time-frequency information of the first sound wave signal; a second processing module configured to determine a second key frame in the video data based on the first key frame; and a detection module configured to determine a live body detection result based on the first key frame and the second key frame.
13. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the live body detection method of any one of claims 1-11.
14. A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, implementing the live body detection method of any one of claims 1-11.
15. A computer program product comprising computer-readable code or a non-transitory computer-readable storage medium carrying computer-readable code, the computer-readable code, when run in a processor of an electronic device, causing the processor in the electronic device to perform the live body detection method of any one of claims 1-11.
Citation Information
Patent Citations
Living body detection method, device, electronic equipment and storage medium
CN113505652A
Living body detection method and device, computer equipment and storage medium
CN114821820A
Face multi-modal detection method and system fusing ultrasonic wave and image information
CN116453233A
Techniques for performing video-based verification
US11736455B1