Target object identification method, device, and equipment, and storage medium
By performing face detection and synchronous lip shape video generation on video files, and combining the comparison model to compare lip image sequences, the problems of low efficiency and low accuracy in lip information verification are solved, thereby improving the efficiency and accuracy of target object recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA PING AN LIFE INSURANCE CO LTD
- Filing Date
- 2023-07-28
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for verifying the identity of a target object using lip information are inefficient and have low accuracy, resulting in poor identity recognition performance.
By performing face detection on the video file, the target lip image sequence is extracted, and the audio file is input into a preset synchronization model to generate a synchronized lip shape video file. The target lip image sequence and the synchronized lip image sequence are compared using a preset comparison model to determine the identity recognition result.
It improves the efficiency of using lip images in financial or insurance business and enhances the accuracy of target object recognition.
Smart Images

Figure CN117058575B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for target object recognition. Background Technology
[0002] With the continuous development of computer technology, target object identification has seen significant progress in recent years. It is being applied in an increasing number of fields, such as the expanding business volume of financial institutions like banks, securities firms, and insurance companies, which generates substantial demand for identification.
[0003] In current technologies, target object recognition typically verifies the identity of a target object through voiceprint recognition. For example, in the banking sector, when identity information needs to be identified, it is usually done by matching the target object's voiceprint information with pre-stored voiceprint information to determine the target object's identity. This requires the cooperation of a large number of personnel to pre-store a large amount of voiceprint information to confirm the target object's identity during verification. Furthermore, the application of lip information in facial images generally involves lip-reading from video to assist voiceprint recognition verification. This method has low efficiency in utilizing lip information, resulting in low accuracy for auxiliary verification. Summary of the Invention
[0004] This invention provides a target object identification method, apparatus, device, and storage medium to improve the low efficiency and low accuracy of lip information verification in the prior art.
[0005] A target object recognition method, comprising:
[0006] Obtain the video file and the corresponding audio file;
[0007] Face detection is performed on each target object in the video file to obtain a sequence of target lip images corresponding to each target object in the video file;
[0008] Each of the aforementioned audio files is input into a preset synchronization model to generate a synchronized lip-sync video file corresponding to each of the aforementioned audio files;
[0009] Extract the synchronized lip image sequence corresponding to each of the target objects from the synchronized lip-shape video file;
[0010] A preset comparison model is obtained, and the target lip image sequence and the synchronous lip image sequence are compared with the same target object through the preset comparison model to obtain the identity recognition result of the target object.
[0011] A target object recognition device, comprising:
[0012] The file acquisition module is used to acquire video files and audio files corresponding to the video files;
[0013] The face detection module is used to perform face detection on target objects in the video file and obtain a sequence of target lip images corresponding to each target object in the video file;
[0014] The synchronous video module is used to input each of the audio files into a preset synchronous model and generate a synchronous lip-sync video file corresponding to each of the audio files.
[0015] The image extraction module is used to extract the synchronized lip image sequence corresponding to each of the target objects in the synchronized lip video file;
[0016] The identification result module is used to obtain a preset comparison model, and compare the target lip image sequence and the synchronous lip image sequence with the same target object through the preset comparison model to obtain the identification result of the target object.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described target object recognition method.
[0018] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described target object recognition method.
[0019] This invention provides a method, apparatus, device, and storage medium for target object recognition. The method performs face detection on target objects in acquired video files, extracting lip images of each target object in the video file, thereby enabling the acquisition of target lip image sequences in financial or insurance applications. By inputting audio files into a preset synchronization model, the audio files are converted into synchronized lip-shape video files. Synchronized lip image sequences corresponding to each target object are extracted from the synchronized lip-shape video files, achieving the acquisition of synchronized lip image sequences. A preset comparison model compares the target lip image sequences and synchronized lip image sequences corresponding to the same target object, determining the identity of the target object and improving the utilization efficiency of lip images in financial or insurance applications. Furthermore, by comparing the target lip image sequences and synchronized lip image sequences using the preset comparison model, the target object is identified through lip image comparison, improving the accuracy of target object recognition in financial or insurance applications. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the application environment of the target object recognition method in one embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a target object identification method in one embodiment of the present invention;
[0023] Figure 3 This is a flowchart of step S30 of the target object identification method in one embodiment of the present invention;
[0024] Figure 4 This is a flowchart of step S50 of the target object identification method in one embodiment of the present invention;
[0025] Figure 5 This is a schematic block diagram of a target object recognition device according to an embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The target object identification method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment is shown. Specifically, this target object recognition method is applied in a target object recognition device, which includes, as shown in the example, a target object recognition device. Figure 1The client and server shown communicate over a network to improve the low efficiency and accuracy of lip-based identity verification in existing technologies. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The client, also known as the user terminal, refers to the program that provides categorized services to customers, corresponding to the server. The client can be installed on, but is not limited to, various computers, laptops, smartphones, tablets, and portable wearable devices.
[0029] In one embodiment, such as Figure 2 As shown, a target object recognition method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps are included:
[0030] S10: Obtain the video file and the audio file corresponding to the video file.
[0031] In essence, a video file is a recording of one or more people. For example, in the insurance industry, a video file might be a conversation between an agent and a customer during an insurance transaction. In the banking industry, a video file might be a conversation between a staff member and a user during a credit card transaction. An audio file is the audio data of the people in the video file. For example, in the insurance industry, an audio file might be a conversation between an agent and a customer. In the banking industry, an audio file might be a conversation between a staff member and a user. Both video and audio files can be pre-prepared and retrieved from a database, or sent from the client to the server. For instance, in an insurance agent's quality inspection process, video files can be pre-processed and stored in a database, or they can be directly inspected in real time.
[0032] S20: Perform face detection on the target objects in the video file to obtain a sequence of target lip images corresponding to each target object in the video file.
[0033] Understandably, a target lip image sequence is a sequence composed of multiple target lip images stitched together.
[0034] Specifically, the process involves recognizing the facial images of each target object in the video file. This is achieved through keypoint detection using computer vision technology. A trained facial keypoint recognition model is used to identify facial features in each target object's facial image, resulting in facial feature points corresponding to each frame of the face image. Then, facial feature points corresponding to the lips are selected from all facial feature points, and these lip feature points are used to extract the lip image from each frame, thus obtaining the lip image corresponding to each frame of the face image. All lip images are then scaled and enhanced, and the scaled lip images from all frames are further filtered. This involves arranging the scaled lip images in chronological order and filtering them at preset frame intervals to obtain the target lip images. Finally, all target lip images are stitched together to obtain the target lip image sequence. For example, in the insurance industry, facial images of salespeople or customers in acquired video files are detected to obtain the target lip image sequence for each target object. The identity information of the corresponding audio file is then determined through lip image quality inspection.
[0035] S30, input each of the audio files into the preset synchronization model to generate a synchronized lip-sync video file corresponding to each of the audio files.
[0036] Understandably, the synchronized lip-sync video file is a video generated by a preset synchronization model based on the audio file.
[0037] Specifically, after obtaining the audio files, speech extraction is performed on each audio file. This extraction can be done using a trained model or publicly available methods to obtain speech information corresponding to each target object. Then, the speech information corresponding to each target object and a preset image are input into a preset synchronization model. The feature extraction layer in the preset synchronization model extracts features from both the speech information and the preset image. Next, a generator transforms the extracted speech and image features to generate a synchronized video. Then, a discriminator performs synchronization judgment on the lip features of the synchronized video, and a visual quality discriminator is used to improve visual quality and synchronization accuracy, thereby obtaining synchronized lip-sync video files corresponding to each audio file. For example, in the insurance field, the acquired audio file and the image of the salesperson are input into the preset synchronization model to generate synchronized lip-sync video files corresponding to the audio files.
[0038] S40: Extract the synchronized lip image sequence corresponding to each of the target objects from the synchronized lip-shaped video file.
[0039] Understandably, a synchronized lip image sequence is a sequence composed of multiple synchronized lip images.
[0040] Specifically, facial image recognition is performed on each target object in the synchronized lip-sync video file. This involves using a trained facial landmark recognition model to identify facial features in the face images of each target object, thereby obtaining synchronized facial feature points corresponding to each frame of the face image. Based on all synchronized facial feature points, the synchronized lip image corresponding to each frame of the face image is determined. This involves selecting synchronized facial feature points corresponding to the synchronized lips from all synchronized facial feature points, and then extracting the synchronized lip image for each frame using these selected points. The synchronized lip images of all frames are then filtered, meaning that all synchronized frame lip images are arranged according to the time sequence of each frame, and then filtered and extracted at a preset frame interval to obtain the synchronized lip image sequence. Understandably, the synchronized lip image sequence corresponding to each target object in the synchronized lip-sync video file can be extracted using a different method than the target lip image sequence. In this embodiment, to improve recognition efficiency, the same method is used to extract the synchronized lip image sequence.
[0041] S50: Obtain a preset comparison model, and compare the target lip image sequence and the synchronous lip image sequence with the same target object through the preset comparison model to obtain the identity recognition result of the target object.
[0042] Understandably, the identity recognition result indicates whether the target object in the audio file and the target object in the video file are the same target object. The preset comparison model is built based on a twin target object recognition network.
[0043] Specifically, the similarity of lip images in the same sequence frames of the target lip image sequence and the synchronized lip image sequence is first compared. This involves performing deep convolutions on all lip features using deep convolutional layers to obtain convolutional features corresponding to each lip feature. Then, a Long Short-Term Memory (LSTM) network is used to perform temporal processing on all convolutional features to obtain temporal features. Similarity is calculated between two temporal features of the same sequence frame to obtain a similarity value corresponding to each sequence frame. Based on all similarity values corresponding to the same video file, the confidence level corresponding to that video file is determined. When the confidence level is greater than or equal to a preset confidence threshold, the first recognition result confirms that the target object in the audio file and the target object in the video file are the same target object. When the confidence level is less than the preset confidence threshold, the second recognition result confirms that the target object in the audio file and the target object in the video file are not the same target object. For example, in the insurance industry, the target lip image sequence of the acquired video file is compared with the generated synchronized lip image sequence to determine whether the person in the audio file is a salesperson and to identify the audio file corresponding to the customer.
[0044] This invention provides a target object recognition method. This method performs face detection on target objects in acquired video files, extracting lip images of each target object in the video file, and thus enabling the acquisition of target lip image sequences in financial or insurance businesses. By inputting each audio file into a preset synchronization model, the audio files are converted into synchronized lip-shaped video files. Synchronized lip image sequences corresponding to each target object are extracted from the synchronized lip-shaped video files, achieving the acquisition of synchronized lip image sequences. A preset comparison model compares the target lip image sequences and synchronized lip image sequences corresponding to the same target object, determining the identity of the target object and improving the utilization efficiency of lip images in financial or insurance businesses. Furthermore, by comparing the target lip image sequences and synchronized lip image sequences using the preset comparison model, the target object is identified through lip image comparison, improving the accuracy of target object recognition in financial or insurance businesses.
[0045] In one embodiment, step S20, namely performing face detection on the target objects in the video file to obtain a sequence of target lip images corresponding to each target object in the video file, includes:
[0046] S201, the facial images of each target object in the video file are identified to obtain facial feature points corresponding to each frame of the facial image.
[0047] Understandably, facial feature points are key points on the face. These key points can be organs such as the eyes and mouth, or features such as single eyelids or the corners of the lips.
[0048] Specifically, facial landmark detection is performed on the facial images of each target object in the video file. This involves acquiring a trained facial landmark detection model and inputting the video file into this model. The model's landmark detection network then performs feature recognition on the facial images of each target object in the video file, marking feature points in each frame of the face image to obtain the facial feature points corresponding to each frame. Specifically, the facial landmark detection model can identify a predetermined number of facial feature points in each frame of the video file to identify the facial contours.
[0049] S202, determine the lip image corresponding to each frame of the face image based on all the facial feature points.
[0050] S203, filter all the lip images in all frames to obtain the target lip image sequence.
[0051] Understandably, a target lip image sequence consists of lip images from different frames of the same target object.
[0052] Specifically, the lip image corresponding to each frame of a face image is determined based on all facial feature points. This involves selecting facial key points corresponding to the lips from the facial feature points of each frame, and extracting the lip image from each frame using these key points. In other words, the region corresponding to the lip's facial key points is segmented from each frame of the face image, thus obtaining the lip image corresponding to each frame. Further, each extracted lip image is enhanced by scaling it (e.g., to 32x38 pixels) and sorting the scaled lip images according to the time sequence of each frame, resulting in an image sorting result. Then, all target lip images are selected from the image sorting result at a preset frame interval (set according to actual conditions, e.g., 10 frames). All selected target lip images are then stitched together into a target lip image sequence according to the time sequence of each frame. For example, in a financial institution's security information verification scenario, to prevent fraud and scams, the speaker needs to be detected to confirm whether they are the target person. Detection typically requires the speaker to recite a passage or verification information for the extraction and recognition of lip information.
[0053] This invention, through the recognition of facial images of target objects in video files, achieves the determination of facial feature points in each frame of facial images and the selection of lip images in each frame. By filtering all lip images, a target lip image sequence is obtained.
[0054] In one embodiment, step S30, which involves inputting each of the audio files into a preset synchronization model to generate synchronized lip-sync video files corresponding to each of the audio files, includes:
[0055] S301, Speech extraction is performed on all the audio files to obtain speech information corresponding to each of the target objects.
[0056] Understandably, an audio file is the corresponding audio data from a video file. Speech information refers to the speech data of the target object.
[0057] Specifically, after obtaining the audio files, speech extraction is performed on all audio files, that is, extracting the speech information corresponding to each target object from the audio files. For example, all speech in the audio files can be clustered to obtain the speech information corresponding to each target object. Alternatively, the similarity between all speech is calculated, that is, the similarity between the voiceprint features of the speech is calculated, and speech with similarity exceeding a threshold is clustered to obtain the speech information corresponding to each target object. It should be noted that this embodiment does not limit the extraction method. For example, in an insurance claims scenario, the audio file is a dialogue between a user and a staff member, and speech extraction is used to obtain the speech information of the two target objects. Or, in a bank loan scenario, the audio file is a dialogue between a bank staff member and a customer, and speech extraction is used to obtain the speech information of the two target objects.
[0058] S302, the preset synchronization model is used to convert each of the voice information into video to obtain a synchronized lip-sync video file corresponding to each of the voice information.
[0059] Understandably, the preset synchronization model is built based on the Wav2Lip method and trained on a large amount of data. The synchronized lip-sync video file is a synchronized lip-sync video generated by the preset synchronization model based on speech information.
[0060] Specifically, speech information is input into a pre-defined synchronization model, which then performs video conversion on each speech message. Specifically, the pre-defined synchronization model uses pre-defined facial images to describe the speech information, extracting features from both the speech information and the pre-defined facial images to obtain corresponding speech and image features. These features are then input into a generator to produce synchronized lip-sync videos corresponding to each speech message. A pre-trained discriminator performs synchronization judgment on the generated lip-sync videos, and a visual quality discriminator improves visual quality and synchronization accuracy. After the generated lip-sync videos pass the discriminator's evaluation, synchronized lip-sync video files corresponding to each speech message are obtained. The pre-defined synchronization model can be trained in-house or utilize publicly available methods; no limitation is imposed here. For example, in the financial field, the pre-defined synchronization model generates videos from extracted audio files and pre-defined images, and the generated synchronized lip-sync videos are used to determine the personnel information corresponding to the audio file.
[0061] This invention extracts speech information from all audio files. By using a preset synchronization model to convert each speech information into video, it achieves the conversion of speech information into video and obtains synchronized lip-sync video files, thereby improving the accuracy of subsequent target object identification.
[0062] In one embodiment, step S50, namely comparing the target lip image sequence and the synchronous lip image sequence with the same target object using the preset comparison model to obtain the identity recognition result of the target object, includes:
[0063] S501, perform a similarity comparison on the lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence to obtain a similarity value corresponding to each sequence frame.
[0064] Specifically, after obtaining the synchronized lip image sequence, a similarity comparison is performed on the lip images of the same sequence frames in both the target lip image sequence and the synchronized lip image sequence. This involves feature recognition of the lip images in all sequence frames of both sequences, specifically feature recognition of the lip images in all sequence frames of each target lip image in the target lip image sequence, and feature recognition of the lip images in all sequence frames of each synchronized lip image in the synchronized lip image sequence. This yields at least one lip feature corresponding to each frame of the lip image. By acquiring all lip features corresponding to the lip images of the same sequence frames in both the target lip image sequence and the synchronized lip image sequence, and calculating the similarity between the acquired lip features of the same frame, a similarity value corresponding to each sequence frame can be obtained. For example, in the insurance industry, the identity of the target object can be determined by comparing the similarity between the synchronized lip image sequence of the generated synchronized video and the lip images of the same sequence frames in the target lip image sequence of the acquired video file.
[0065] S502, determine the confidence level corresponding to the video file based on all the similarity values corresponding to the same video file.
[0066] Furthermore, based on all similarity values corresponding to the same video file, the confidence level corresponding to that video file is determined. This is achieved by calculating the confidence level of each target object using the similarity values of each sequence frame corresponding to the same target object. Specifically, the confidence level of each target object is obtained by determining the number of lip images in each sequence frame whose similarity value is greater than or equal to a preset similarity threshold. Then, the confidence level of the video file is calculated using the confidence levels of all target objects corresponding to the same video file. This involves multiplying the confidence level of each target object by a preset weight (which can be the same or different) to determine the confidence level corresponding to that video file. Alternatively, the confidence level of each target object can be scored, all scores can be summed, and then the confidence level corresponding to the video file can be determined using a preset correspondence.
[0067] S503, when the confidence level is greater than or equal to a preset confidence threshold, the identity recognition result is confirmed as the first recognition result, and the first recognition result indicates that the target object in the audio file and the target object in the video file are the same target object.
[0068] S504, when the confidence level is less than the preset confidence threshold, the identity recognition result is confirmed as the second recognition result, which indicates that the target object in the audio file and the target object in the video file are not the same target object.
[0069] Specifically, a preset confidence threshold is retrieved and compared with the confidence level corresponding to the video file. When the confidence level is greater than or equal to the preset confidence threshold, the identity recognition result is confirmed as the first recognition result, indicating that the target object in the audio file and the target object in the video file are the same target object. When the confidence level is less than the preset confidence threshold, the identity recognition result is confirmed as the second recognition result, indicating that the target object in the audio file and the target object in the video file are not the same target object. For example, in the banking sector, when opening an account at Ping An Bank or Ping An Securities, or when transferring large sums of money, security verification of the operator is required. This involves comparing and recognizing the lip information of the operator while speaking to determine whether it is the same person.
[0070] This invention achieves the acquisition of a similarity value for each sequence frame by comparing the similarity of lip images from the same sequence frames in a target lip image sequence and a synchronized lip image sequence. The confidence level is calculated using all similarity values corresponding to the same video file. By comparing the confidence level with a preset confidence threshold, the identity of the target object is identified, thereby improving the accuracy of target object identification.
[0071] In one embodiment, step S501, namely comparing the similarity of lip images in the same sequence frames of the target lip image sequence and the synchronized lip image sequence to obtain a similarity value corresponding to each sequence frame, includes:
[0072] S5011, feature recognition is performed on the lip images of all sequence frames in the target lip image sequence and the synchronous lip image sequence respectively to obtain at least one lip feature corresponding to each frame of lip image.
[0073] Specifically, feature recognition is performed on the lip images of all frames in both the target lip image sequence and the synchronized lip image sequence. This involves identifying one or more of the following features in all lip images: lip opening degree, left lip drooping degree, and right lip drooping degree. First, feature points in the lip region are identified. Then, the lip opening degree is determined based on the distance between the inner center feature point of the upper lip and the inner center feature point of the lower lip. The left corner lip feature point is connected to the nearest feature point on the outer contour lines of the upper and lower lips to form a first vector. The angle between these first vectors is then calculated to obtain the left lip drooping degree feature. Similarly, the right corner lip feature point is connected to the nearest feature point on the outer contour lines of the upper and lower lips to form a second vector. The angle between these second vectors is then calculated to obtain the right lip drooping degree feature. In this way, at least one lip feature corresponding to each frame of the lip image can be obtained. For example, in the insurance field, by extracting features from lip images from different sources, we can obtain lip features such as the degree of lip opening, the degree of left lip slant, and the degree of right lip slant corresponding to the lip images when speaking.
[0074] S5012, acquire all the lip features corresponding to the lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence, and determine the similarity value based on the similarity between the acquired lip features of the same frame.
[0075] Specifically, all lip features corresponding to the same sequence frames in the target lip image sequence and the synchronized lip image sequence are acquired, resulting in lip opening degree features, left lip aversion degree features, and right lip aversion degree features corresponding to the lip images in the same sequence frames. Then, image similarity is calculated between the lip opening degree features, or between the left lip aversion degree features, or between the right lip aversion degree features corresponding to the lip images in the same sequence frames. This is achieved by first performing deep convolution on the lip opening degree features, left lip aversion degree features, and right lip aversion degree features respectively, using deep convolutional layers to obtain convolutional features corresponding to the lip features in each frame. Then, a long short-term memory network is used to perform temporal processing on the convolutional features to obtain temporal features. Similarity is then calculated between two temporal features in the same sequence frames to obtain the similarity value corresponding to the lip images in the same sequence frames. For example, in financial scenarios, during identity verification, the similarity between two sets of lip images from different sources is calculated to determine whether the verification personnel are the target individuals. Alternatively, in remote security information verification, it is necessary to identify and verify the identity of the target object in the video to ensure that it is the person performing the operation.
[0076] This invention achieves lip feature extraction by performing feature recognition on lip images in all sequence frames of both the target lip image sequence and the synchronized lip image sequence, thereby acquiring all lip features from lip images within the same sequence frame. Based on the similarity between the acquired lip features within the same frame, a similarity value is calculated, ensuring the accuracy of the recognition results.
[0077] In one embodiment, before step S50, that is, before obtaining the preset comparison model, the following steps are included:
[0078] S601, Obtain a sample dataset, the sample dataset including at least one sample data and sample labels corresponding to the sample data.
[0079] In essence, the sample data consists of a set of lip images of the target object within the same frame. Specifically, one set of lip images in the sample data is extracted from facial information in a recorded video file, and the other is extracted from facial information in a video converted from audio data using a pre-defined synchronization model. Sample labels are used to characterize the identification result of the target object corresponding to this set of lip images in the sample data. For example, in the insurance industry, identification is performed on videos between insurance agents and customers to determine the target object corresponding to the audio data. Alternatively, in the banking industry, lip information is compared and verified during bank account opening. Sample data can be collected from different databases or sent from the client to a database. A sample dataset is then constructed based on all sample data and all sample labels.
[0080] S602, Obtain a preset training model, input all the sample data into the preset training model, and obtain the predicted label corresponding to each sample data.
[0081] Understandably, the predicted label represents the similarity value of the predictions made by the pre-trained model to the sample data.
[0082] Specifically, after obtaining the sample dataset, a pre-set training model is acquired, and all sample data is input into the pre-set training model. The pre-set training model compares a set of lip images in the sample data, that is, it performs similarity comparison on lip images of the same sequence frames in the first and second lip image sequences. First, deep convolutional layers are applied to the first and second lip images in the sample data to obtain sample convolutional features corresponding to each frame of lip images. Then, a Long Short-Term Memory (LSTM) network is used to perform temporal processing on the sample convolutional features to obtain sample temporal features. Similarity is calculated between the temporal features of two samples from the same sequence frame to obtain sample similarity values corresponding to each sequence frame. Based on all similarity values corresponding to the same video file, the sample confidence is determined. When the sample confidence is greater than or equal to the sample confidence threshold, it is confirmed that the target objects corresponding to the lip images in the sample data are the same target object. When the sample confidence is less than the sample confidence threshold, it is confirmed that the target objects corresponding to the lip images in the sample data are not the same target object. The specific process is the same as steps S501 to S504 above, and will not be described in detail here.
[0083] S603, determine the prediction loss value corresponding to the preset training model based on the sample label and prediction label corresponding to the same sample data.
[0084] Understandably, the prediction loss is generated during the process of making predictions on the sample data.
[0085] Specifically, after obtaining the predicted labels, all predicted labels corresponding to the same sample data are arranged according to the order of the sample data in the sample dataset. Then, the sample labels are compared with the predicted labels of the same sequence. That is, according to the sorting of the sample data, the first sample label is compared with the first predicted label, and the loss value between the sample label and the predicted label is calculated by the loss function. Then, the second sample label is compared with the second predicted label, until all sample labels and all predicted labels have been compared. The loss values of all sample data are added together to obtain the prediction loss value corresponding to the preset training model.
[0086] S604, when the predicted loss value reaches the convergence condition, the preset training model after convergence is determined as the preset comparison model.
[0087] Understandably, the convergence condition can be either the predicted loss value being less than a set threshold, or the predicted loss value being very small after 500 calculations and no longer decreasing, at which point training can stop.
[0088] Specifically, after obtaining the predicted loss value, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted based on the predicted loss value. All sample data are then re-input into the preset training model with adjusted initial parameters, and iterative training is performed to obtain the predicted loss value corresponding to the preset training model with adjusted initial parameters. Then, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted again based on the predicted loss value, so that the predicted loss value of the preset training model with adjusted initial parameters reaches the preset convergence condition. In this way, the accuracy of the preset training model increases, continuously approaching the correct result, until the predicted loss value of the preset training model reaches the preset convergence condition. At this point, the converged preset training model is determined as the preset comparison model.
[0089] This invention iteratively trains a pre-set training model using a large amount of sample data and calculates the overall loss value of the pre-set training model, thus determining the predicted loss value of the pre-set training model. The initial parameters of the pre-set training model are adjusted based on the predicted loss value until the model converges, thereby determining the pre-set comparison model and ensuring that the pre-set comparison model has a high prediction accuracy.
[0090] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0091] In one embodiment, a target object recognition device is provided, which corresponds one-to-one with the target object recognition method in the above embodiments. For example... Figure 5 As shown, the target object recognition device includes a file acquisition module 11, a face detection module 12, a synchronous video module 13, an image extraction module 14, and a recognition result module 15. Detailed descriptions of each functional module are as follows:
[0092] The file acquisition module 11 is used to acquire video files and audio files corresponding to the video files;
[0093] The face detection module 12 is used to perform face detection on the target objects in the video file and obtain a target lip image sequence corresponding to each target object in the video file;
[0094] Synchronous video module 13 is used to input each of the audio files into a preset synchronization model to generate synchronized lip-sync video files corresponding to each of the audio files;
[0095] Image extraction module 14 is used to extract the synchronized lip image sequence corresponding to each of the target objects in the synchronized lip video file;
[0096] The recognition result module 15 is used to obtain a preset comparison model, and compare the target lip image sequence and the synchronous lip image sequence with the same target object through the preset comparison model to obtain the identity recognition result of the target object.
[0097] In one embodiment, the recognition result module 15 includes:
[0098] The similarity unit is used to compare the similarity of lip images in the same sequence frame of the target lip image sequence and the synchronized lip image sequence to obtain a similarity value corresponding to each sequence frame.
[0099] A confidence unit is used to determine the confidence level corresponding to the video file based on all the similarity values corresponding to the same video file.
[0100] The first identification result unit is used to confirm the identity identification result as the first identification result when the confidence level is greater than or equal to a preset confidence threshold. The first identification result indicates that the target object in the audio file and the target object in the video file are the same target object.
[0101] The second identification result unit is used to confirm the identity recognition result as the second identification result when the confidence level is less than a preset confidence threshold. The second identification result indicates that the target object in the audio file and the target object in the video file are not the same target object.
[0102] In one embodiment, the similarity unit includes:
[0103] The feature recognition subunit is used to perform feature recognition on the lip images of all sequence frames in the target lip image sequence and the synchronous lip image sequence, respectively, to obtain at least one lip feature corresponding to each frame of lip image;
[0104] The similarity value subunit is used to acquire all the lip features corresponding to the lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence, and to determine the similarity value based on the similarity between the acquired lip features of the same frame.
[0105] In one embodiment, the synchronized video module 13 includes:
[0106] The voice information unit is used to extract voice from all the audio files to obtain voice information corresponding to each of the target objects;
[0107] The lip-sync video unit is used to perform video conversion on each of the aforementioned speech information through the preset synchronization model, so as to obtain a synchronized lip-sync video file corresponding to each of the aforementioned speech information.
[0108] In one embodiment, the face detection module 12 includes:
[0109] The facial feature point unit is used to identify the face images of each target object in the video file and obtain the facial feature points corresponding to each frame of the face image;
[0110] A lip image unit is used to determine a lip image corresponding to each frame of a face image based on all the facial feature points.
[0111] An image sequence unit is used to filter all frames of the lip images to obtain the target lip image sequence.
[0112] In one embodiment, the recognition result module 15 further includes:
[0113] A sample acquisition unit is used to acquire a sample dataset, the sample dataset including at least one sample data and sample labels corresponding to the sample data;
[0114] The label prediction unit is used to obtain a preset training model, input all the sample data into the preset training model, and obtain the predicted label corresponding to each sample data.
[0115] The loss prediction unit is used to determine the predicted loss value corresponding to the preset training model based on the sample label and predicted label corresponding to the same sample data.
[0116] The model convergence unit is used to determine the converged preset training model as the preset comparison model when the predicted loss value reaches the convergence condition.
[0117] Specific limitations regarding the target object recognition device can be found in the limitations of the target object recognition method described above, and will not be repeated here. Each module in the aforementioned target object recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0118] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used in the target object identification method described in the above embodiments. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a target object identification method.
[0119] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target object recognition method described above.
[0120] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described target object identification method.
[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0123] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for identifying a target object, characterized in that, include: Obtain the video file and the corresponding audio file; Face detection is performed on each target object in the video file to obtain a sequence of target lip images corresponding to each target object in the video file; Each of the aforementioned audio files is input into a preset synchronization model to generate a synchronized lip-sync video file corresponding to each of the aforementioned audio files; Extract the synchronized lip image sequence corresponding to each of the target objects from the synchronized lip-shape video file; A preset comparison model is obtained, and the target lip image sequence and the synchronous lip image sequence are compared with the same target object through the preset comparison model to obtain the identity recognition result of the target object; The step of inputting each of the audio files into a preset synchronization model to generate a synchronized lip-sync video file corresponding to each of the audio files includes: Speech extraction is performed on all the audio files to obtain speech information corresponding to each of the target objects; The preset synchronization model is used to convert each of the speech information into video, resulting in synchronized lip-sync video files corresponding to each of the speech information. Specifically, features are extracted from the speech information and the preset face image to obtain speech features and image features. The speech features and image features of the target object are input into the generator to generate synchronized lip-sync videos corresponding to each of the speech information. A discriminator is used to synchronously judge the lip features of the synchronized videos to obtain synchronized lip-sync video files corresponding to each audio file. The step of comparing the target lip image sequence and the synchronous lip image sequence with the same target object using the preset comparison model to obtain the identity recognition result of the target object includes: The lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence are compared for similarity to obtain the similarity value corresponding to each sequence frame. The confidence level corresponding to the video file is determined based on all similarity values corresponding to the same video file; wherein, the confidence level of each target object is obtained by judging the number of lip images in each sequence frame corresponding to a similarity value greater than or equal to a preset similarity threshold; the confidence level of the video file is determined by calculating the confidence level of the video file based on the confidence levels of all target objects corresponding to the same video file. When the confidence level is greater than or equal to a preset confidence threshold, the identity recognition result is confirmed as the first recognition result, which indicates that the target object in the audio file and the target object in the video file are the same target object. When the confidence level is less than a preset confidence threshold, the identity recognition result is confirmed as the second recognition result, which indicates that the target object in the audio file and the target object in the video file are not the same target object. The step of comparing the similarity of lip images in the same sequence frames of the target lip image sequence and the synchronized lip image sequence to obtain a similarity value corresponding to each sequence frame includes: Feature recognition is performed on the lip images of all sequence frames in the target lip image sequence and the synchronous lip image sequence to obtain at least one lip feature corresponding to each frame of the lip image. Specifically, feature point recognition is performed on the lip region, and the degree of lip opening is determined based on the distance between the inner center feature point of the upper lip and the inner center feature point of the lower lip. After connecting the left corner lip feature point with the feature point closest to the left corner lip feature point on the outer contour line of the upper and lower lips to form a first vector, the angle between the first vectors is calculated to obtain the left lip drooping degree feature. After connecting the right corner lip feature point with the feature point closest to the right corner lip feature point on the outer contour line of the upper and lower lips to form a second vector, the angle between the second vectors is calculated to obtain the right lip drooping degree feature, thus obtaining at least one lip feature corresponding to each frame of the lip image. Obtain all the lip features corresponding to the lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence, and determine the similarity value based on the similarity between the obtained lip features of the same frame.
2. The target object recognition method as described in claim 1, characterized in that, The step of performing face detection on target objects in the video file to obtain a sequence of target lip images corresponding to each target object in the video file includes: The facial images of each target object in the video file are identified to obtain facial feature points corresponding to each frame of the facial image; Based on all the facial feature points, determine the lip image corresponding to each frame of the face image; The target lip image sequence is obtained by filtering all frames of the lip images.
3. The target object recognition method as described in claim 1, characterized in that, Before obtaining the preset comparison model, the following steps are included: Obtain a sample dataset, which includes at least one sample data and sample labels corresponding to the sample data; Obtain a preset training model, input all the sample data into the preset training model, and obtain the predicted label corresponding to each sample data; The prediction loss value corresponding to the preset training model is determined based on the sample label and prediction label corresponding to the same sample data. When the predicted loss value reaches the convergence condition, the preset training model after convergence is determined as the preset comparison model.
4. A target object recognition device, characterized in that, include: The file acquisition module is used to acquire video files and audio files corresponding to the video files; The face detection module is used to perform face detection on target objects in the video file and obtain a sequence of target lip images corresponding to each target object in the video file; The synchronous video module is used to input each of the audio files into a preset synchronous model and generate a synchronous lip-sync video file corresponding to each of the audio files. The image extraction module is used to extract the synchronized lip image sequence corresponding to each of the target objects in the synchronized lip video file; The identification result module is used to obtain a preset comparison model, and compare the target lip image sequence and the synchronous lip image sequence with the same target object through the preset comparison model to obtain the identification result of the target object; The synchronized video module includes: The voice information unit is used to extract voice from all the audio files to obtain voice information corresponding to each of the target objects; The lip-sync video unit is used to perform video conversion on each of the aforementioned speech information using the preset synchronization model, thereby obtaining synchronized lip-sync video files corresponding to each of the aforementioned speech information. Specifically, features are extracted from the speech information and the preset face image to obtain speech features and image features. The speech features and image features of the target object are input into the generator to generate synchronized lip-sync videos corresponding to each of the speech information. A discriminator performs synchronous judgment on the lip features of the synchronized videos to obtain synchronized lip-sync video files corresponding to each of the audio files. The recognition result module includes: The similarity unit is used to compare the similarity of lip images in the same sequence frame of the target lip image sequence and the synchronized lip image sequence to obtain a similarity value corresponding to each sequence frame. A confidence unit is used to determine the confidence level corresponding to the video file based on all the similarity values corresponding to the same video file; wherein, the confidence level of each target object is obtained by judging the number of lip images in each sequence frame corresponding to a similarity value greater than or equal to a preset similarity threshold; the confidence level corresponding to the video file is determined by calculating the confidence level of the video file based on the confidence levels of all target objects corresponding to the same video file. The first identification result unit is used to confirm the identity identification result as the first identification result when the confidence level is greater than or equal to a preset confidence threshold. The first identification result indicates that the target object in the audio file and the target object in the video file are the same target object. The second identification result unit is used to confirm the identity recognition result as the second identification result when the confidence level is less than a preset confidence threshold. The second identification result indicates that the target object in the audio file and the target object in the video file are not the same target object. The similarity unit includes: A feature recognition subunit is used to perform feature recognition on lip images of all sequence frames in the target lip image sequence and the synchronized lip image sequence, respectively, to obtain at least one lip feature corresponding to each frame of lip image; wherein, feature point recognition is performed on the lip region, and the degree of lip opening is determined based on the distance between the inner center feature point of the upper lip and the inner center feature point of the lower lip; after connecting the left corner lip feature point with the feature point closest to the left corner lip feature point on the outer contour line of the upper and lower lips to form a first vector, the angle between the first vectors is calculated to obtain the left-hand lip drooping feature; after connecting the right corner lip feature point with the feature point closest to the right corner lip feature point on the outer contour line of the upper and lower lips to form a second vector, the angle between the second vectors is calculated to obtain the right-hand lip drooping feature, thus obtaining at least one lip feature corresponding to each frame of lip image; The similarity value subunit is used to acquire all the lip features corresponding to the lip images of the same sequence frames in the target lip image sequence and the synchronized lip image sequence, and to determine the similarity value based on the similarity between the acquired lip features of the same frame.
5. The target object recognition device as described in claim 4, characterized in that, The recognition result module also includes: A sample acquisition unit is used to acquire a sample dataset, the sample dataset including at least one sample data and sample labels corresponding to the sample data; The label prediction unit is used to obtain a preset training model, input all the sample data into the preset training model, and obtain the predicted label corresponding to each sample data. The loss prediction unit is used to determine the predicted loss value corresponding to the preset training model based on the sample label and predicted label corresponding to the same sample data. The model convergence unit is used to determine the converged preset training model as the preset comparison model when the predicted loss value reaches the convergence condition.
6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the target object recognition method as described in any one of claims 1 to 3.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the target object recognition method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Identity authentication method and apparatus
CN107404381A
Virtual human image video generation method, system and device and storage medium
CN113192161A
Apparatus, method, and computer program for providing lip-sync video and apparatus, method, and computer program for displaying lip-sync video
US20230023102A1