Face data processing method and device, terminal equipment and storage medium
By collecting three-dimensional image data through a depth camera, identifying the facial area and generating a motion trajectory, and adjusting the acquisition mode according to the duration of the trajectory, the problem of low accuracy caused by the short appearance time of the face is solved, and efficient face matching is achieved in different scenarios.
Patent Information
- Application Number
- CN202510894673.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing face acquisition technology cannot obtain accurate features in scenarios where faces appear for a short time in real-time video streams, resulting in low accuracy.
The depth camera is used to collect three-dimensional image data, identify the face area, generate and predict the motion trajectory, adjust the acquisition mode according to the trajectory duration, generate feature vector matching under short duration, perform multi-frame continuous shooting and image fusion under long duration, and obtain super-resolution images for matching.
Improves the accuracy of face recognition, is applicable to scenarios where faces appear at various times, and improves matching accuracy.
Smart Images

Figure CN120808413A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of face collection, and in particular to a face data processing method and device, a terminal device and a storage medium. BACKGROUND
[0002] Face collection technology is widely used in public security, identity authentication and face recognition, etc. In the field of public security, a large number of video probes installed at city road intersections, buses, etc. collect face images to assist in tracking criminal suspects, and provide identity verification basis for access control systems, financial transactions, etc.
[0003] However, in the real-time video stream face collection scene, the existing face collection technology needs to perform super-resolution face detection operation on all faces appearing in the video, but in the application scene where the face appears for a short time, the face obtained by super-resolution still cannot obtain specific features, so the existing face collection technology cannot be applied to the application scene where the face appears for a short time, resulting in the problem of low accuracy of face collection technology. SUMMARY
[0004] The present application provides a face data processing method, device, terminal device and storage medium, which can solve the problem of low accuracy of face collection technology in the prior art.
[0005] The face data processing method provided by the present application comprises:
[0006] Obtaining a plurality of frames of three-dimensional image data collected by a depth camera within a preset time, and after obtaining each frame of three-dimensional image data, identifying each frame of three-dimensional image data according to a preset face three-dimensional structure sample to obtain a face region corresponding to each frame of three-dimensional image data;
[0007] Based on the face region of each frame, generating a face motion trajectory within the preset time, and predicting the face motion trajectory to obtain a face prediction trajectory;
[0008] Determining a first duration of face appearance according to the face motion trajectory; determining a second duration of face appearance according to the face prediction trajectory; and calculating a total duration of face appearance according to the first duration and the second duration;
[0009] If the total duration is less than a first time threshold, generating a face feature vector corresponding to each frame of three-dimensional image data based on the plurality of frames of three-dimensional image data, so that a user performs face matching based on the face feature vector;
[0010] If the total time length is greater than the second time threshold, the depth camera is controlled to start a multi-frame continuous shooting mode, and a plurality of frames of three-dimensional continuous shooting image data obtained after starting the multi-frame continuous shooting mode is subjected to image fusion with the plurality of frames of three-dimensional image data to obtain a super-resolution image, so that the user performs face matching based on the super-resolution image; wherein the second time threshold is greater than the first time threshold.
[0011] Further, when obtaining each frame of three-dimensional image data, a light intensity collected by a light sensor and a device vibration amplitude collected by a gyroscope are synchronously obtained to obtain a light intensity and a device vibration amplitude corresponding to the current frame of three-dimensional image data;
[0012] According to the face region corresponding to the current frame of three-dimensional image data and the face region corresponding to the previous frame of three-dimensional image data, a face motion speed corresponding to the current frame of three-dimensional image data is calculated.
[0013] According to the light intensity, the device vibration amplitude, and the face motion speed corresponding to the current frame of three-dimensional image data, the depth camera is adjusted so that the adjusted depth camera collects the next frame of three-dimensional image data.
[0014] Further, the adjusting the depth camera according to the light intensity, the device vibration amplitude, and the face motion speed corresponding to the current frame of three-dimensional image data so that the adjusted depth camera collects the next frame of three-dimensional image data comprises:
[0015] The face motion speed, the light intensity, and the device vibration amplitude corresponding to the current frame of three-dimensional image data are judged.
[0016] If the face speed corresponding to the current frame of three-dimensional image data is greater than or equal to a speed threshold, the depth camera is adjusted in parameters based on a first frame rate threshold and a first resolution threshold so that the depth camera adjusted in parameters collects the next frame of three-dimensional image data.
[0017] If the face speed corresponding to the current frame of three-dimensional image data is less than the speed threshold, the depth camera is adjusted in parameters based on a second frame rate threshold and a second resolution threshold so that the depth camera adjusted in parameters collects the next frame of three-dimensional image data.
[0018] If the light intensity corresponding to the current frame of three-dimensional image data is less than a light intensity threshold, the depth camera is adjusted in parameters based on a preset exposure time threshold so that the depth camera adjusted in parameters collects the next frame of three-dimensional image data.
[0019] If the device vibration amplitude corresponding to the current frame of three-dimensional image data is greater than or equal to the vibration amplitude threshold, the depth camera is controlled to start the anti-shake algorithm, so that the depth camera starting the anti-shake algorithm collects the next frame of three-dimensional image data.
[0020] Further, the face motion trajectory in the preset time is generated based on the face region of each frame, including:
[0021] For each frame of three-dimensional image data, each frame of three-dimensional image data is converted into a corresponding two-dimensional image.
[0022] Based on the position of the face region of each frame, the optical flow vector of each group of adjacent two frames of face region corresponding two-dimensional images is calculated by an optical flow algorithm, and the face region position offset of each group of adjacent two frames of two-dimensional images is determined based on the optical flow vector.
[0023] In the preset time, each frame of two-dimensional image is arranged in time sequence, and based on the arranged two-dimensional image and the face region position offset of each group of adjacent two frames of two-dimensional images, the face motion trajectory is generated.
[0024] Further, the face motion trajectory is predicted to obtain a face prediction trajectory, including:
[0025] The face motion trajectory is predicted by a Kalman filtering algorithm to obtain a face prediction trajectory.
[0026] Further, the face feature vector corresponding to each frame of three-dimensional image data is generated based on a plurality of frames of three-dimensional image data, including:
[0027] The two-dimensional image corresponding to each frame of three-dimensional image data is input into a neural network to generate a weighted feature map corresponding to each frame of three-dimensional image data; the weighted feature map corresponding to each frame of three-dimensional image data is spatially dimensionally compressed to obtain a face feature vector corresponding to each frame of three-dimensional image data.
[0028] Further, the embodiment also includes:
[0029] The face matching result is judged.
[0030] If the matching is successful, the depth camera is controlled to follow the movement based on the face motion trajectory corresponding to the matching successful face region, so that the matching successful face region remains in the collection range of the depth camera.
[0031] If the matching fails, no operation is performed.
[0032] Another embodiment of the present application also provides a face data processing device, which comprises a data acquisition module, a trajectory prediction module, a data judgment module, a first matching module and a second matching module.
[0033] The data acquisition module is configured to acquire a plurality of frames of three-dimensional image data collected by the depth camera within a preset time, and identify each frame of three-dimensional image data according to a preset face three-dimensional structure sample after acquiring each frame of three-dimensional image data, to obtain a face region corresponding to each frame of three-dimensional image data.
[0034] The trajectory prediction module is configured to generate a face motion trajectory within the preset time based on each frame of face region, and predict the face motion trajectory to obtain a face predicted trajectory.
[0035] The data judgment module is configured to determine a first time length of face appearance according to the face motion trajectory, determine a second time length of face appearance according to the face predicted trajectory, and calculate a total time length of face appearance according to the first time length and the second time length.
[0036] The first matching module is configured to generate a face feature vector corresponding to each frame of three-dimensional image data based on the plurality of frames of three-dimensional image data, if the total time length is less than a first time threshold, so that a user performs face matching based on the face feature vector.
[0037] The second matching module is configured to control the depth camera to start a multi-frame continuous shooting mode if the total time length is greater than a second time threshold, and perform image fusion on a plurality of frames of three-dimensional continuous shooting image data obtained after starting the multi-frame continuous shooting mode and the plurality of frames of three-dimensional image data, to obtain a super-resolution image, so that the user performs face matching based on the super-resolution image, wherein the second time threshold is greater than the first time threshold.
[0038] Another embodiment of the present application further provides a terminal device, comprising a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, when the processor executes the computer program, the steps of the face data processing method provided by the present application are implemented.
[0039] Another embodiment of the present application further provides a computer readable storage medium item, comprising a stored computer program, when the computer program runs, the device where the computer readable storage medium is located is controlled to execute the steps of the face data processing method provided by the present application.
[0040] The present application has the following beneficial effects:
[0041] The application discloses a face data processing method, after a plurality of three-dimensional image data of a depth camera in a preset time is acquired, face recognition is performed on each three-dimensional image data to obtain a face region corresponding to each three-dimensional image data; face motion track generation and face prediction track prediction are performed based on each face region, and finally, face matching operation judgment is performed according to the sum of a first time length corresponding to the face motion track and a second time length corresponding to the face prediction track; if the total time length is less than a first time threshold, fast face matching is performed through the generated face feature vector; if the total time length is greater than a second time threshold, more accurate face matching is performed through a multi-frame continuous shooting mode and image fusion to obtain a more fine super-resolution image. Compared with the prior art, the application can switch the data used for face matching based on the length of the face appearance time, can be applied to various scenes of face appearance, and improves the accuracy of face recognition. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0043] Figure 1 is a flowchart of the face data processing method provided by an embodiment of the present application;
[0044] Figure 2 is a structural schematic diagram of the face data processing device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application more clear, the technical solutions in the present application will be described clearly and completely in the following with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "include" and "have" and any variations thereof in the specification and claims of the present application and the above description of drawings are intended to cover non-exclusive inclusion.
[0047] In the description of the embodiments of the present application, the technical terms "first", "second" and the like are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise explicitly and specifically limited.
[0048] Reference herein to "embodiments" means that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily independent or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0049] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it.
[0050] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two), and similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0051] In the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the technical terms "mounting", "connection", "connection", "fixing" and the like should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanical connection, or it can be electrical connection; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the embodiments of the present application can be understood according to the specific circumstances.
[0052] Reference Figure 1 To solve the problem of low accuracy of face collection technology in the prior art, an embodiment of the present application provides a face data processing method, which comprises:
[0053] 101、Obtain a plurality of frames of three-dimensional image data collected by the depth camera within a preset time, and after obtaining each frame of three-dimensional image data, identify each frame of three-dimensional image data according to a preset face three-dimensional structure sample to obtain a face region corresponding to each frame of three-dimensional image data.
[0054] It should be noted that any face information collected by the present patent application is applied to public security, and the use of face information is within the scope of legal operation, and the collection of face information is based on user authorization.
[0055] In the present embodiment, when acquiring each frame of three-dimensional image data, the light intensity collected by the light sensor and the device vibration amplitude collected by the gyroscope are synchronously acquired to obtain the light intensity and the device vibration amplitude corresponding to the current frame of three-dimensional image data;
[0056] According to the face region corresponding to the current frame of three-dimensional image data and the face region corresponding to the previous frame of three-dimensional image data, the face motion speed corresponding to the current frame of three-dimensional image data is calculated;
[0057] According to the light intensity, the device vibration amplitude and the face motion speed corresponding to the current frame of three-dimensional image data, the depth camera is adjusted so that the adjusted depth camera collects the next frame of three-dimensional image data.
[0058] In a specific embodiment, the face motion speed v is calculated by an optical flow algorithm, and the formula is wherein dx / dt and dy / dt are the change rates of the optical flow in the x and y directions in the three-dimensional image data, which are calculated from the optical flow of the face region corresponding to the previous frame of three-dimensional image data and the optical flow of the face region corresponding to the current frame of three-dimensional image data.
[0059] In the present embodiment, the adjustment of the depth camera according to the light intensity, the device vibration amplitude and the face motion speed corresponding to the current frame of three-dimensional image data so that the adjusted depth camera collects the next frame of three-dimensional image data comprises:
[0060] The face motion speed, the light intensity and the device vibration amplitude corresponding to the current frame of three-dimensional image data are judged;
[0061] If the face speed corresponding to the current frame of three-dimensional image data is greater than or equal to the speed threshold, the depth camera is adjusted in parameters based on the first frame rate threshold and the first resolution threshold so that the depth camera after the parameter adjustment collects the next frame of three-dimensional image data;
[0062] If the face speed corresponding to the current frame of three-dimensional image data is less than the speed threshold, the depth camera is adjusted in parameters based on the second frame rate threshold and the second resolution threshold so that the depth camera after the parameter adjustment collects the next frame of three-dimensional image data;
[0063] If the light intensity corresponding to the current frame of three-dimensional image data is less than the light intensity threshold, the depth camera is adjusted in parameters based on a preset exposure time threshold, so that the depth camera after the parameter adjustment collects the next frame of three-dimensional image data.
[0064] If the device vibration amplitude corresponding to the current frame of three-dimensional image data is greater than or equal to the vibration amplitude threshold, the depth camera is controlled to start a anti-shake algorithm, so that the depth camera after starting the anti-shake algorithm collects the next frame of three-dimensional image data.
[0065] In a specific embodiment, if the face speed corresponding to the current frame of three-dimensional image data is less than the speed threshold, the depth camera is not adjusted in parameters.
[0066] If the face speed corresponding to the current frame of three-dimensional image data is greater than or equal to the speed threshold, the depth camera is not adjusted in parameters.
[0067] If the light intensity corresponding to the current frame of three-dimensional image data is greater than or equal to the light intensity threshold, the depth camera is not adjusted in parameters.
[0068] If the device vibration amplitude corresponding to the current frame of three-dimensional image data is less than the vibration amplitude threshold, the depth camera is not adjusted in parameters.
[0069] In a specific embodiment, when the face motion speed is detected to be greater than or equal to the speed threshold of 50 pixels / s, the frame rate is increased to 60 frames / s (i.e., the first frame rate threshold in the present application), and the resolution is reduced to 640x480 (i.e., the first resolution threshold in the present application);
[0070] When the face motion speed is less than the speed threshold, the frame rate is maintained at 30 frames / s (i.e., the second frame rate threshold in the present application), and the resolution is increased to 1920x1080 (i.e., the second resolution threshold in the present application);
[0071] When the light sensor detects that the light intensity is lower than 5 lux (i.e., the light intensity threshold in the present application), the camera exposure time is automatically increased to 50 ms (i.e., the preset exposure time threshold in the present application);
[0072] When the gyroscope detects that the device vibration amplitude exceeds 0.5g (i.e., the vibration amplitude threshold in the present application), the anti-shake algorithm is used to compensate the image based on the depth camera data, so as to reduce the dynamic blur.
[0073] In a specific embodiment, the depth camera is adjusted in parameters based on the dynamic resolution adaptive adjustment mechanism based on the face speed, the light intensity, and the device vibration amplitude.
[0074] 102. generating a face motion trajectory within a preset time based on the face region of each frame, and predicting the face motion trajectory to obtain a face predicted trajectory.
[0075] In the embodiment, the generating a face motion trajectory within a preset time based on the face region of each frame comprises:
[0076] For each frame of three-dimensional image data, converting each frame of three-dimensional image data into corresponding two-dimensional image data;
[0077] Based on the position of the face region of each frame, calculating the optical flow vector of the corresponding two-dimensional image of each group of adjacent two frames of face regions by an optical flow algorithm, and determining the position offset of the face region of each group of adjacent two frames of two-dimensional images based on the optical flow vector.
[0078] Within a preset time, arranging each frame of two-dimensional image in time sequence, and based on the arranged two-dimensional image and the position offset of the face region of each group of adjacent two frames of two-dimensional images, generating a face motion trajectory.
[0079] In a specific embodiment, when the optical flow algorithm generates the face motion trajectory, the input is the two-dimensional image data corresponding to the position of each frame of face region obtained by recognizing the preset face three-dimensional structure sample.
[0080] First, the two-dimensional image is grayed and preprocessed, and then the optical flow vector between adjacent frames is calculated by the pyramid Lucas-Kanade optical flow algorithm to determine the position offset of the face region, and the coordinates of each frame of face region are mapped to a preset frame to form a trajectory point sequence in time sequence, and the face motion trajectory within a preset time is obtained after smoothing processing.
[0081] In the embodiment, the predicting the face motion trajectory to obtain a face predicted trajectory comprises:
[0082] The face predicted trajectory is obtained by predicting the face motion trajectory by a Kalman filtering algorithm.
[0083] In a specific embodiment, the formula for predicting the next frame position by the Kalman filtering algorithm is wherein, is the predicted state, A is the state transition matrix, B is the control matrix, and uk is the control input.
[0084] By taking the predicted next frame position as the new input of the Kalman filtering algorithm, and inputting the next next frame position, the face predicted trajectory is generated by predicting a plurality of frame positions.
[0085] In a specific embodiment, the training logic of the state transition matrix A, the control matrix B and the control input uk of the Kalman filtering algorithm is as follows:
[0086] The state transition matrix A is determined by analyzing the position and speed change rules in the historical face motion trajectory data, for example, the mapping relationship between the adjacent frame face coordinate displacement and speed is counted within a preset time, and a matrix reflecting the face motion dynamics characteristics is constructed.
[0087] The control matrix B is trained according to the influence coefficient of the camera acquisition parameters (such as frame rate, resolution) on the face trajectory, and is used to quantify the effect of external control on the trajectory.
[0088] The control input uk is generated based on the real-time environmental parameters such as the face motion speed calculated by the optical flow algorithm and the device vibration amplitude detected by the gyroscope, without additional training, and is obtained and input into the algorithm by real-time measurement.
[0089] The combination of the three enables the Kalman filter to predict the future trajectory based on the historical trajectory and the current state, ensuring the prediction accuracy.
[0090] 103. Determine the first duration of the face appearance according to the face motion trajectory; determine the second duration of the face appearance according to the face prediction trajectory; and calculate the total duration of the face appearance according to the first duration and the second duration.
[0091] In a specific embodiment, the first face appearance time and the second face appearance time are used to determine the time of the target corresponding to the face region in the depth camera field of view, and the face region is divided into two categories: "instantaneous passing" and "continuous staying". When the predicted face stays in the current region for 2 seconds (i.e. the first time threshold), it is determined as an "instantaneous passing" face; when the predicted face stays in the current region for more than 5 seconds (i.e. the second time threshold), it is determined as a "continuous staying" face.
[0092] 104. If the total duration is less than the first time threshold, generate a face feature vector corresponding to each frame of three-dimensional image data based on a plurality of frames of three-dimensional image data, so that the user can perform face matching based on the face feature vector.
[0093] In a specific embodiment, for the "instantaneous passing" face, the feature vector is hashed and stored in a fast retrieval database. The fast retrieval database uses Redis to achieve second-level retrieval, so that the regulatory user responsible for public security (i.e. the user described in the present application) can perform face matching based on the collected face feature vector, so that the regulatory user can maintain public security according to the face matching result.
[0094] In this embodiment, the face feature vector corresponding to each frame of three-dimensional image data is generated based on a plurality of frames of three-dimensional image data, which includes:
[0095] The two-dimensional image corresponding to each frame of three-dimensional image data is input into the neural network to generate a weighted feature map corresponding to each frame of three-dimensional image data; and the weighted feature map corresponding to each frame of three-dimensional image data is subjected to spatial dimension compression to obtain a face feature vector corresponding to each frame of three-dimensional image data.
[0096] In a specific embodiment, the neural network adopts a lightweight neural network based on an attention mechanism, and the attention weight of the neural network is updated according to real-time feedback of the face feature. The attention weight is updated once every 30 frames of images to adapt to changes in face features of different races and ages.
[0097] 105. If the total duration is greater than the second time threshold, the depth camera is controlled to start a multi-frame continuous shooting mode, and a plurality of frames of three-dimensional continuous shooting image data obtained after the multi-frame continuous shooting mode is started are subjected to image fusion with the plurality of frames of three-dimensional image data to obtain a super-resolution image, so that the user performs face matching based on the super-resolution image; wherein the second time threshold is greater than the first time threshold.
[0098] In a specific embodiment, for a face that "continuously stays", a multi-frame continuous shooting mode of the depth camera is triggered and a super-resolution image is generated by fusion and stored in a long-term storage. The long-term storage adopts a distributed file system Ceph for long-term storage of high-quality face image data.
[0099] In this embodiment, the embodiment also includes:
[0100] The face matching result is judged.
[0101] If the matching is successful, the depth camera is controlled to follow the movement based on the face motion trajectory corresponding to the face region for which the matching is successful, so that the face region for which the matching is successful remains within the capture range of the depth camera.
[0102] If the matching fails, no operation is performed.
[0103] In a specific embodiment, the face feature vector collected is compared and matched with a feature library. If the matching is successful, a tracking program is started. During the tracking process, the face position and motion trajectory changes are continuously monitored. When the face is out of the current camera field of view, the position information of the face in other cameras is obtained through data interaction with the surrounding cameras to realize cross-camera tracking. At the same time, the motion trajectory of the face during the tracking process is analyzed. If an abnormal motion trajectory is found, a warning information is issued. The abnormalities include sudden change of direction or too long stay.
[0104] The specific personnel face feature library is constructed by extracting face texture features by using a local binary pattern algorithm, and the feature library is updated once every 50 face images of different angles and expressions of the specific personnel are successfully collected;
[0105] In the abnormal motion trajectory judgment, the direction change angle threshold is set to 30 degrees, and the stay time threshold is 60 seconds.
[0106] In a specific embodiment, if the total duration is greater than or equal to the first time threshold and less than or equal to the second time threshold, the three-dimensional image data of each frame is stored in a preset database, so that the user directly performs face matching of the target personnel based on the three-dimensional image data stored in the preset database.
[0107] As shown in Figure 2 Based on the above method embodiment, a corresponding device embodiment is provided;
[0108] An embodiment of the present application provides a face data processing device, which comprises a data acquisition module 201, a trajectory prediction module 202, a data judgment module 203, a first matching module 204 and a second matching module 205.
[0109] The data acquisition module is used for acquiring a plurality of frames of three-dimensional image data collected by a depth camera within a preset time, and after acquiring each frame of three-dimensional image data, identifying each frame of three-dimensional image data according to a preset face three-dimensional structure sample to obtain a face region corresponding to each frame of three-dimensional image data.
[0110] The trajectory prediction module is used for generating a face motion trajectory within the preset time based on the face region of each frame, and predicting the face motion trajectory to obtain a face predicted trajectory.
[0111] The data judgment module is used for determining a first duration of face appearance according to the face motion trajectory, determining a second duration of face appearance according to the face predicted trajectory, and calculating a total duration of face appearance according to the first duration and the second duration.
[0112] The first matching module is used for generating a face feature vector corresponding to each frame of three-dimensional image data based on the plurality of frames of three-dimensional image data if the total duration is less than the first time threshold, so that the user performs face matching based on the face feature vector.
[0113] The second matching module is configured to control the depth camera to start a multi-frame continuous shooting mode if the total time length is greater than the second time threshold, and to perform image fusion on a plurality of frames of three-dimensional continuous shooting image data obtained after the multi-frame continuous shooting mode is started and the plurality of frames of three-dimensional image data, to obtain a super-resolution image, so that the user performs face matching based on the super-resolution image; and the second time threshold is greater than the first time threshold.
[0114] It can be understood that the above device embodiments correspond to the method embodiments of the present application, and can realize the face data processing method provided by any one of the method embodiments of the present application.
[0115] It should be noted that the device embodiments described above are only schematic, and part or all of the modules can be selected to achieve the purpose of the present embodiment. In addition, in the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be realized as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.
[0116] On the basis of the above-mentioned embodiments of the face data processing method, another embodiment of the present application provides a terminal device, which comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the face data processing method of any one of the embodiments of the present application is realized.
[0117] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the terminal device.
[0118] The terminal device can be a desktop computer, a notebook computer, a palm computer and a cloud server, etc. The terminal device can include, but is not limited to, a processor and a memory.
[0119] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like, which is the control center of the terminal device and connects all parts of the terminal device through various interfaces and lines.
[0120] On the basis of the above-mentioned method embodiment, another embodiment of the present application provides a computer readable storage medium, including a stored computer program, wherein when the computer program runs, the device where the computer readable storage medium is located executes the face data processing method in any one of the above-mentioned method embodiments of the present application.
[0121] The modules / units integrated in the device / terminal equipment, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0122] The above-mentioned is the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered to be within the protection scope of the present application.
Claims
1. A facial data processing method, characterized in that: include: Acquire several frames of three-dimensional image data collected by the depth camera within a preset time, and after acquiring each frame of three-dimensional image data, identify each frame of three-dimensional image data based on a preset three-dimensional facial structure sample to obtain the face area corresponding to each frame of three-dimensional image data; Based on the face area of each frame, a face motion trajectory within a preset time is generated, and the face motion trajectory is predicted to obtain a face prediction trajectory; Determining a first duration of face appearance based on the face motion trajectory; determining a second duration of face appearance based on the face prediction trajectory; and calculating a total duration of face appearance based on the first duration and the second duration; If the total duration is less than the first time threshold, generating a facial feature vector corresponding to each frame of three-dimensional image data based on the plurality of frames of three-dimensional image data, so that the user can perform face matching based on the facial feature vector; If the total duration is greater than a second time threshold, the depth camera is controlled to start a multi-frame continuous shooting mode, and several frames of three-dimensional continuous shooting image data obtained after starting the multi-frame continuous shooting mode are fused with several frames of the three-dimensional image data to obtain a super-resolution image, so that the user can perform face matching based on the super-resolution image; wherein the second time threshold is greater than the first time threshold.
2. The facial data processing method according to claim 1, wherein: When acquiring each frame of three-dimensional image data, a light intensity acquired by the light sensor and a device vibration amplitude acquired by the gyroscope are synchronously acquired to obtain the light intensity and device vibration amplitude corresponding to the current frame of three-dimensional image data; Calculating a face motion speed corresponding to the current frame of 3D image data based on a face region corresponding to the current frame of 3D image data and a face region corresponding to the previous frame of 3D image data; The depth camera is adjusted according to the light intensity, device vibration amplitude and face movement speed corresponding to the current frame of three-dimensional image data, so that the adjusted depth camera collects the next frame of three-dimensional image data.
3. The facial data processing method according to claim 2, wherein: The adjusting the depth camera according to the light intensity, device vibration amplitude, and face movement speed corresponding to the current frame of three-dimensional image data, so that the adjusted depth camera collects the next frame of three-dimensional image data, includes: Determine the facial motion speed, light intensity, and device vibration amplitude corresponding to the current frame of 3D image data; If the face speed corresponding to the current frame of three-dimensional image data is greater than or equal to the speed threshold, adjusting the parameters of the depth camera based on the first frame rate threshold and the first resolution threshold so that the depth camera with the adjusted parameters collects the next frame of three-dimensional image data; If the face speed corresponding to the current frame of three-dimensional image data is less than the speed threshold, adjusting the parameters of the depth camera based on the second frame rate threshold and the second resolution threshold so that the depth camera with the adjusted parameters collects the next frame of three-dimensional image data; If the light intensity corresponding to the current frame of 3D image data is less than the light intensity threshold, the depth camera parameters are adjusted based on the preset exposure time threshold so that the depth camera with the adjusted parameters collects the next frame of 3D image data; If the device vibration amplitude corresponding to the current frame of three-dimensional image data is greater than or equal to the vibration amplitude threshold, the depth camera is controlled to start the anti-shake algorithm, so that the depth camera after starting the anti-shake algorithm collects the next frame of three-dimensional image data.
4. The facial data processing method according to claim 3, wherein: The method of generating a face motion trajectory within a preset time based on the face area of each frame includes: For each frame of three-dimensional image data, convert each frame of three-dimensional image data into a corresponding two-dimensional image; Based on the position of the face area in each frame, the optical flow algorithm is used to calculate the optical flow vector of the two-dimensional images corresponding to the face areas of each set of two adjacent frames, and the position offset of the face areas of the two-dimensional images of each set of two adjacent frames is determined based on the optical flow vector; Within a preset time, each frame of two-dimensional images is arranged in chronological order, and based on the arranged two-dimensional images and the face area position offset of each group of two adjacent two-dimensional images, a face motion trajectory is generated.
5. The facial data processing method according to claim 4, wherein: The step of predicting the face motion trajectory to obtain the predicted face trajectory includes: The face motion trajectory is predicted using the Kalman filter algorithm to obtain the predicted face trajectory.
6. The facial data processing method according to claim 5, wherein: The step of generating a facial feature vector corresponding to each frame of three-dimensional image data based on a plurality of frames of three-dimensional image data includes: The two-dimensional image corresponding to each frame of three-dimensional image data is input into the neural network to generate a weighted feature map corresponding to each frame of three-dimensional image data; the weighted feature map corresponding to each frame of three-dimensional image data is spatially compressed to obtain a facial feature vector corresponding to each frame of three-dimensional image data.
7. The facial data processing method according to claim 6, wherein: Also includes: Judge the face matching results; If the match is successful, the depth camera is controlled to follow the facial motion trajectory corresponding to the matched face area so that the matched face area remains within the acquisition range of the depth camera; If the match fails, no action is taken.
8. A facial data processing device, characterized in that: include: Data acquisition module, trajectory prediction module, data judgment module, first matching module and second matching module; The data acquisition module is used to acquire a plurality of frames of three-dimensional image data collected by the depth camera within a preset time, and after acquiring each frame of three-dimensional image data, identify each frame of three-dimensional image data based on a preset three-dimensional facial structure sample to obtain a facial region corresponding to each frame of three-dimensional image data; The trajectory prediction module is used to generate a facial motion trajectory within a preset time based on the face area of each frame, and predict the facial motion trajectory to obtain a predicted facial trajectory; The data judgment module is used to determine the first duration of the appearance of the face according to the face movement trajectory; Determining a second duration of face appearance based on the predicted face trajectory; calculating a total duration of face appearance based on the first duration and the second duration; The first matching module is configured to generate, based on a plurality of frames of three-dimensional image data, a facial feature vector corresponding to each frame of three-dimensional image data if the total duration is less than a first time threshold, so that the user can perform face matching based on the facial feature vector; The second matching module is used to control the depth camera to start a multi-frame continuous shooting mode if the total duration is greater than a second time threshold, and to fuse several frames of three-dimensional continuous shooting image data obtained after starting the multi-frame continuous shooting mode with several frames of the three-dimensional image data to obtain a super-resolution image, so that the user can perform face matching based on the super-resolution image; wherein the second time threshold is greater than the first time threshold.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for processing facial data according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that include: A stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the face data processing method according to any one of claims 1 to 7.