Get-up behavior detection method and device, equipment and medium
Through cutting and speech recognition technology, the target customer's answer video data is processed, and the mouth movement and voiceprint data are combined for detection, which solves the problem of difficult answering behavior of outsiders on-line review and picture review, and improves the accuracy of the detection.
Patent Information
- Application Number
- CN202510372935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-25
AI Technical Summary
There is a risk of fraud in online review methods, especially the behavior of answering outsiders that is difficult to identify with high accuracy.
By obtaining the video data of the target customer's answers, cutting it into video clips, and using voice recognition technology to obtain the speech text content and timestamp. If the preset speech template is matched, mouth motion detection is performed. If the mouth is not moved, historical voiceprint data is obtained from the voiceprint library for detection.
It improves the accuracy of external answering behavior detection and reduces the risk of fraud in other people's answering behaviors.
Smart Images

Figure CN120148513A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fintech, and in particular, to a method, apparatus, device, and medium for detecting proxy answering behavior. Background Art
[0002] With the development of Internet technology, the face-to-face review methods in current scenarios such as loans and insurance have gradually changed from on-site review to online face-to-face review, which has greatly facilitated people's lives. However, due to the limited camera angle, the online face-to-face review method inevitably has fraud risks (such as being guided by others or answered by others on behalf).
[0003] In the prior art, the detection technology based on mouth movements can solve the detection of some violations. For example, by detecting whether the mouth moves when the applicant and the agent are speaking, some violations can be avoided. However, for violations outside the picture, it is difficult to accurately identify them only by detecting the mouth movement of the customer. Therefore, there is an urgent need for a method that can improve the accuracy of detecting proxy answering behavior by people outside the picture to reduce the fraud risk of behaviors such as being answered by others on behalf. Summary of the Invention
[0004] The embodiments of this application provide a method, apparatus, device, and medium for detecting proxy answering behavior, which can improve the accuracy of detecting proxy answering behavior by people outside the picture.
[0005] In a first aspect, the embodiments of this application provide a method for detecting proxy answering behavior, including:
[0006] Obtain the answer video data of the target customer for the target event, and cut the answer video data into multiple video segments;
[0007] Determine the target video segment to be detected currently from the multiple video segments;
[0008] Use speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content;
[0009] If the speech text content matches the preset speech template, perform mouth movement detection on the target video segment according to the timestamp and generate a mouth movement detection result;
[0010] If the mouth movement detection result is that the mouth moves, store the voiceprint data corresponding to the speech text content into the voiceprint library corresponding to the target event;
[0011] If the mouth movement detection result is that the mouth does not move, obtain the historical voiceprint data corresponding to the target event from the voiceprint library, and detect the target voiceprint data of the target video segment according to the historical voiceprint data to obtain the detection result of proxy answering behavior by people outside the picture.
[0012] In a second aspect, an embodiment of the present application further provides a substitute answering behavior detection device, which includes:
[0013] A cutting unit, configured to obtain response video data of a target customer for a target event, and cut the response video data into multiple video segments;
[0014] A determining unit, configured to determine a target video segment to be detected currently from the multiple video segments;
[0015] An obtaining unit, configured to use speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content;
[0016] A first detection unit, configured to, if the speech text content matches a preset speech template, perform mouth movement detection on the target video segment according to the timestamp and generate a mouth movement detection result;
[0017] A storage unit, configured to, if the mouth movement detection result is that there is mouth movement, store the voiceprint data corresponding to the speech text content into a voiceprint database corresponding to the target event;
[0018] A second detection unit, configured to, if the mouth movement detection result is that there is no mouth movement, obtain historical voiceprint data corresponding to the target event from the voiceprint database, and perform detection on the target voiceprint data of the target video segment according to the historical voiceprint data to obtain a detection result of a substitute answering behavior by an off-screen person.
[0019] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.
[0021] An embodiment of the present application provides a method, device, equipment and medium for detecting proxy answering behavior. The method includes: obtaining answer video data of a target customer for a target event, and cutting the answer video data into multiple video segments; determining a target video segment to be detected currently from the multiple video segments; using speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content; if the speech text content matches a preset speech template, performing mouth movement detection on the target video segment according to the timestamp and generating a mouth movement detection result; if the mouth movement detection result is that the mouth moves, storing the voiceprint data corresponding to the speech text content into the voiceprint library corresponding to the target event; if the mouth movement detection result is that the mouth does not move, obtaining historical voiceprint data corresponding to the target event from the voiceprint library, and detecting the target voiceprint data of the target video segment according to the historical voiceprint data to obtain a detection result of proxy answering behavior by an outsider. In the embodiment of the present application, on the basis of determining the target video segment to be detected currently from multiple video segments of the answer video data of the target customer for the target event, the corresponding speech text content and the timestamp of each word are obtained through speech recognition technology. If the speech text content matches the preset speech template, mouth movement detection is performed on the target video segment according to the timestamp to generate a detection result. If the detection result is that the mouth moves, the voiceprint data corresponding to the speech text content is stored into the voiceprint library corresponding to the target event. If the detection result is that the mouth does not move, historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and the target voiceprint data of the target video segment is detected accordingly to obtain a detection result of proxy answering behavior by an outsider, thereby improving the accuracy of detecting proxy answering behavior by an outsider. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 FIG. is a schematic diagram of an application scenario of the proxy answering behavior detection method provided by the embodiment of the present application;
[0024] Figure 2 FIG. is a schematic flowchart of the proxy answering behavior detection method provided by the embodiment of the present application;
[0025] Figure 3 FIG. is a schematic sub - flowchart of the proxy answering behavior detection method provided by the embodiment of the present application;
[0026] Figure 4A schematic block diagram of the proxy answer behavior detection device provided by the embodiments of the present application; and
[0027] Figure 5 A schematic block diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0029] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0030] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0031] It should be further understood that the term " / and / " as used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0032] The embodiments of the present application provide a proxy answer behavior detection method, device, equipment, and medium. The execution subject of the proxy answer behavior detection method may be the proxy answer behavior detection device provided by the embodiments of the present application, or a computer device integrated with the proxy answer behavior detection device. Among them, the proxy answer behavior detection device may be implemented in a hardware or software manner, and the computer device may be a terminal or a server. In the embodiments of the present application, the proxy answer behavior detection method provided by the present application will be described in detail taking this computer device as an example.
[0033] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the proxy answer behavior detection method provided by the embodiments of the present application. In one embodiment, the application scenario includes:
[0034] The proxy answer behavior detection method is applied to Figure 1In the computer device, in some embodiments, the computer device obtains multiple video segments of the response video data of the target customer for the target event, determines the target video segment to be currently detected from the multiple video segments, obtains the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content through speech recognition technology. If the speech text content matches the preset speech template, the computer device performs mouth movement detection on the target video segment according to the timestamp to generate a detection result. If the detection result is mouth movement, the voiceprint data corresponding to the speech text content is stored in the voiceprint library corresponding to the target event. If the detection result is no mouth movement, the historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and based on this, the voiceprint data detection of the target video segment is performed to obtain the detection result of the off-screen person answering on behalf of others. The computer device detects the off-screen person answering on behalf of others through mouth movement detection and voiceprint data detection, reducing the risk of false detection and improving the accuracy of detecting the off-screen person answering on behalf of others.
[0035] It should be noted that the application scenario of the above-mentioned answering-on-behalf detection method is only used to illustrate the technical solution of this application and is not used to limit the technical solution of this application. The above connection relationship may also have other forms.
[0036] Figure 2 It is a flowchart of the answering-on-behalf detection method provided by the embodiment of this application. As Figure 2 shown, the method includes the following steps S110 - S160.
[0037] S110. Obtain the response video data of the target customer for the target event, and cut the response video data into multiple video segments.
[0038] Among them, the target customer refers to an individual or group of customers selected as the object of analysis and research in a specific business scenario. The target event refers to an event that is associated with the target customer in a specific business scenario and has specific research significance. For example, in the video face-to-face interview scenario of a certain housing loan, the creditor will ask some questions to the debtor to confirm that he has the ability to perform the debt. The debtor participating in this video face-to-face interview of the housing loan is the target customer, and at the same time, this video face-to-face interview of the housing loan carried out for the target customer is the target event.
[0039] Specifically, it is possible to establish a connection with a video acquisition device (e.g., a camera) or a storage system to obtain the response video data of the target customer for the target event. Among them, the connection method can be selected according to the actual situation. For example, use a network interface to call the Application Programming Interface (API) to obtain from a remote server, or directly read from a local storage device. After obtaining the response video data of the target customer for the target event, verify the response video data. Among them, the content of the verification includes format verification and metadata verification. For format verification, it can include checking whether the format of the video file is correct, whether the file size meets the expectations, and whether the encoding method of the video is compatible with the subsequent processing tools. Among them, if the format of the video file is correct, the file size meets the expectations, and the encoding method of the video is compatible with the subsequent processing tools, the format verification result is qualified. For example, if the subsequent video processing software only supports the common MP4 format, and the obtained video is in AVI format, format conversion is required. For metadata verification, it can include checking whether data information such as shooting time, resolution, and frame rate is recorded in the response video data. Among them, if data information such as shooting time, resolution, and frame rate is recorded in the response video data, the metadata verification result is qualified. By verifying data information such as shooting time, resolution, and frame rate, it can be ensured that the obtained video data is complete and usable, avoiding unsuccessful or incorrect processing of the video data due to data problems in subsequent processing of the video data.
[0040] Obtain answer video data with qualified form verification results and metadata verification results, and use a specific audio and video processing tool (for example, FFmpeg) or a preset model to cut the answer video data. After a certain length of answer video data is cut by a specific tool or a preset model, several video clips can be generated. For example, take the cutting of answer video data of a video interview of an education loan using an open source cross-platform audio and video processing tool called FFmpeg as an example: first, obtain the total length of 15 minutes of interview answer video data from the video storage system, and the file name of the interview answer video data is input.mp4; secondly, according to the set parameters, the answer video data is divided according to the specified start time, end time and segment length. For example, using the FFmpeg command "ffmpeg-iinput.mp4-ss 00:00:10-t 00:00:30-c copy output1.mp4", it means that starting from the 10th second of the input.mp4 video, a 30-second segment is intercepted and saved as output1.mp4. Since the total length of the video data of the answers to the video interview of the education loan is 15 minutes, the video data of the answers to the video interview of the education loan can be cut into multiple video clips and output in the same way. Finally, after the cutting is completed, check the generated video clips. You can use a video player to play these video clips one by one to check whether the cutting time point is accurate, whether the video content is complete, whether the sound is normal, etc. If you find that there is a problem with a video clip, you can adjust the cutting time point according to the specific situation and re-execute the cutting command.
[0041] Through the above steps, the FFmpeg tool can be used to conveniently cut the answer video of the education loan video interview into multiple independent video clips, which is convenient for subsequent in-depth analysis of the answer to each question in the answer video data, and provide stronger support for the authenticity and accuracy of the creditor's answers in the subsequent loan review.
[0042] Furthermore, in some embodiments, the step of cutting the answer video data into a plurality of video segments comprises the following steps:
[0043] The answer video data is segmented and processed by using a preset voice activity detection model to obtain a plurality of video segments.
[0044] Specifically, the preset Voice Activity Detection (VAD) model is a voice processing technology mainly used to detect the presence of human voice signals. The VAD model can accurately identify the voice part, avoiding including a large number of meaningless silent or noise segments in the cut video, enabling subsequent analysts to focus on valuable voice content and improving the analysis efficiency. For example, in the answer video of a loan face-to-face interview, there may be short pauses or ambient noise when the interviewer communicates with the applicant. The VAD model can effectively filter these parts and only retain the conversation content between the two parties. The entire cutting process is automatically completed based on the preset VAD model, reducing the workload of manual marking and cutting and lowering the possibility of human errors. In addition, when processing a large number of insurance customer return visit videos, manual cutting requires a lot of time and effort, while using the VAD model for automated cutting can complete the task in a short time and improve work efficiency.
[0045] Before cutting and processing the answer video data through the preset VAD model, first use a professional audio-video processing library (such as FFmpeg or moviepy) to extract the audio track from the answer video data. Then, preprocess the extracted audio to meet the input requirements of the VAD model. The preprocessing includes but is not limited to audio format conversion (such as unifying the audio sampling rate to a specific value required by the VAD model, commonly 16kHz), normalizing the audio volume, and frame processing of the audio, etc. For example, the pydub library in Python can be used for audio format conversion and volume normalization, adjusting the audio volume to a certain standard range to improve the accuracy of VAD model detection. Input the preprocessed audio data into the loaded and initialized VAD model. The VAD model will analyze the audio data frame by frame and determine whether each frame belongs to the voice part, and output the judgment result. Among them, the output result of the VAD model can be in the form of a marking sequence. For example, "1" indicates that the frame contains voice, and "0" indicates that the frame is silent or non-voice (such as ambient noise).
[0046] According to the output result of the VAD model, the cutting points of the video segment can be determined. Map the start and end positions of the voice to the timeline of the answer video segment. For example, if it is detected that a segment of voice starts at the 5th second and ends at the 15th second of the audio, then in the answer video segment, this is a potential cutting interval. When determining the cutting points, a certain buffer time can also be considered, such as adding 0.5 seconds before the start and after the end of the voice to ensure that the voice content is completely included and avoid truncating the voice part due to detection errors.
[0047] For example, in the scenario of video review for in-person signing of a housing loan, the in-person signing officer will ask the loan applicant a series of questions, such as income situation, debt situation, purpose of purchasing a house, etc. By using a pre-set VAD model to cut the entire in-person signing video, multiple video clips of the housing loan in-person signing video can be obtained. Subsequently, by analyzing the answers of the loan applicant in the video clips, the risk of someone other than the loan applicant answering the review questions can be reduced. Specifically, load the trained VAD model, extract the audio from the answer video data and preprocess it, and input the audio into the trained VAD model for voice activity detection. When the voice part of the loan applicant's answer regarding the income situation is detected, determine its time range in the video, such as from the 30th second to the 60th second of the video, and then use a video editing tool to cut out this segment of the video. Subsequently, the reviewer can directly analyze the language expression, tone, etc. of the loan applicant's answer regarding income-related questions for this video clip to judge the authenticity and reliability of the income information, greatly improving the efficiency and accuracy of loan review.
[0048] Specifically, in some embodiments, before step S120, the method includes the following steps:
[0049] Detect whether there are video clips among the multiple video clips that have not been detected for the behavior of someone else answering on behalf of the applicant.
[0050] Determining the target video clip that needs to be detected currently from the multiple video clips includes:
[0051] If there are video clips among the multiple video clips that have not been detected for the behavior of someone else answering on behalf of the applicant, then determine a video clip from the video clips that have not been detected for the behavior of someone else answering on behalf of the applicant as the target video clip.
[0052] Among them, the target video clip refers to a certain video clip that needs to be detected currently among the multiple video clips cut from the answer video data.
[0053] Specifically, take the video face-to-face signing scenario for housing loans as an example. In the video face-to-face signing scenario for housing loans, when a bank staff member (the lender) conducts a video face-to-face signing with the borrower, the lender will ask the borrower some questions (such as asking about the lender's income, liabilities, purpose of purchasing a house, etc.) to ensure that the borrower has the ability to fulfill the repayment obligation. During the face-to-face signing process, the process of the borrower answering the questions raised by the lender will be recorded throughout to form the response video data of the target customer for the target event. The bank obtains the response video data from the credit business system and, according to the face-to-face signing process, cuts the response video data by question category, such as the response segment regarding income questions, the response segment regarding liability questions, etc. At the same time, the bank will establish a record form to record the detection status of each video segment. Among them, all the multiple video segments formed by the first cut of the response video data are "undetected", and their status will be marked as "detected" after being detected as the target video segments. Each time during the detection, the first "undetected" video segment is selected as the target video segment in chronological order among the multiple video segments, and its status is updated to "detected" after the detection is completed until all the video segments are detected.
[0054] On the one hand, by only focusing on the undetected video segments, it is possible to avoid repeated processing of the detected segments, thus saving time and computing resources. Especially when processing a large number of insurance customer return visit videos, the processing speed can be significantly improved. On the other hand, selecting the target segment from the undetected segments according to certain rules can ensure that all video segments are finally detected, without missing any segment that may have the behavior of someone else answering on behalf of the person in the video, thereby improving the risk prevention and control ability.
[0055] S120. Determine the target video segment that needs to be detected currently from multiple video segments.
[0056] Specifically, in the response video data of the target customer for the target time, after cutting the response video data through a preset voice activity detection model, multiple video segments can be obtained. According to the preset selection rule, the video segment that needs to be detected currently among the obtained multiple video segments is used as the target video segment. For example, in the response video data of target customer A during the face-to-face review for a second-hand housing loan at a certain bank, after cutting the response video data of target customer A through a preset voice activity detection model, multiple video segments are obtained. Among these multiple video segments, there are multiple non-voice video segments and multiple voice-containing video segments. At this time, the video segments corresponding to the multiple voice-containing video segments are arranged in chronological order from front to back, and the video segment ranked first is used as the target video segment that needs to be detected currently.
[0057] S130: Acquire the speech text content corresponding to the target video clip and the timestamp of each word in the speech text content by using speech recognition technology.
[0058] Among them, Automatic Speech Recognition (ASR) technology refers to the technology of converting the vocabulary content in human speech into computer-readable text. It is widely used in fields such as intelligent voice assistants and voice input.
[0059] Specifically, the audio data in the target video clip is extracted through video editing software or a special audio extraction tool, and the extracted audio data in the target video clip is preprocessed (for example, format conversion or noise reduction processing). The extracted audio data is uploaded to a speech recognition engine (for example, Google Cloud Speech to Text or iFlytek Engine). By calling the speech recognition engine, it will automatically convert the speech content in the audio data into speech text content and output it. In addition, while outputting the speech text content, the start and end timestamps corresponding to each word in the speech text content will also be provided.
[0060] For example, during a video interview for a housing loan, bank staff (creditors) will ask the debtor multiple questions, such as "What is your current monthly income?" "Do you have any other debts?" etc. The target video clip of the debtor answering the creditor's questions is processed according to the above method. First, the audio data is extracted from the target video clip, and after pre-processing such as format conversion and noise reduction, the iFlytek engine interface is called. The voice text content recognized by the iFlytek engine may be "My current monthly income is about 15,000 yuan, and I have no other debts." At the same time, each word in the voice text content output by the iFlytek engine is timestamped, such as "I" at 0.5-0.6 seconds, "eye" at 0.6-0.7 seconds, etc. Bank staff (creditors) can use these timestamps and combine them with video data such as the debtor's mouth movements when speaking to more accurately judge the authenticity and credibility of the debtor's answers.
[0061] Specifically, in some embodiments, before step S140, the method includes the following steps:
[0062] Based on a preset word vector model, convert the speech text content into a first word vector, and convert the preset speech template into a second word vector;
[0063] Calculate the cosine angle between the first word vector and the second word vector to obtain a cosine angle value;
[0064] If the cosine angle value is greater than or equal to a preset angle threshold, it is determined that the speech text content matches the preset speech template;
[0065] If the cosine angle value is less than the preset angle threshold, it is determined that the speech text content does not match the preset speech template.
[0066] Among them, the preset word vector model refers to a model constructed based on machine learning or deep learning technology, which can map words or phrases in natural language into word vectors. The word vector not only contains the semantic information of the word, but also reflects the semantic relationship between words, such as semantic similarity, semantic relevance, etc. The preset speech template refers to a series of texts pre-set for verifying speech text content, which are used to match and compare with the input speech text content, and usually includes a series of key sentences, vocabulary combinations, and the possible order or logical relationship of their appearance, etc.
[0067] Specifically, preprocess the obtained speech text content and the preset speech template, including removing punctuation marks, converting to lowercase, removing stop words, etc., to improve the accuracy and efficiency of subsequent word vector calculations. For example, use the functions in the Natural Language Toolkit (NLTK) library to remove stop words. For the preprocessed speech text content, it can be first split into individual words or phrases through a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model, and then based on the locally pre-stored and trained word vector model (for example, Word2Vec model or GloVe model), convert each word or phrase into the corresponding word vector. Finally, the foregoing word vectors can be combined into a first word vector representing the entire speech text content by means of weighted averaging. In the same way, after preprocessing the preset speech template, perform word segmentation and word vector conversion to obtain a second word vector.
[0068] After obtaining the first word vector and the second word vector, substitute the first word vector and the second word vector into the cosine similarity formula to calculate the cosine angle between the two, and the cosine angle value can be obtained. The cosine similarity formula is as follows:
[0069]
[0070] Among them, represents the first word vector; represents the second word vector; · represents the dot product of vectors; represents the modulus of the first vector; Represents the modulus of the second vector. The cosine angle value can be obtained by substituting the first vector and the second vector into the cosine similarity formula. The range of this cosine angle value is between [-1, 1]. Among them, the closer the cosine angle value is to 1, the more similar the directions of the two vectors are, that is, the higher the matching degree between the speech text content and the semantics of the preset speech template.
[0071] Compare the cosine angle value calculated by the cosine similarity formula with the preset angle threshold. If the cosine angle value is greater than or equal to the preset angle threshold, it means that the similarity between the speech text content and the preset speech template is high, that is, it means that the speech text content matches the preset speech template; if the cosine angle value is less than the preset angle threshold, it means that the similarity between the speech text content and the preset speech template is low, that is, it means that the speech text content does not match the preset speech template.
[0072] For example, there is currently a customer A who conducts a video face-to-face review for an auto consumption loan at Bank B. Assume that the preset speech template is "The loan amount you applied for is 200,000 yuan, the loan term is 3 years, the annual interest rate is 5.8%, and the equal principal and interest repayment method is adopted. Do you confirm the above information?". The speech text content corresponding to this question in the video clip of customer A's answer is "The loan I applied for is 200,000 yuan, to be repaid in 3 years, with an annual interest rate of 5.8%, and equal principal and interest repayment. Confirm no problem.". Then, on the basis of preprocessing steps such as removing punctuation marks, converting to lowercase, and removing stop words for this speech text content of customer A, use the pre-trained BERT model to split it into word groups ["loan", "200,000 yuan", "in 3 years", "annual interest rate 5.8%", "equal principal and interest", "repayment", "confirm", "no problem"], and then based on the word vector model (such as Word2Vec model or GloVe model) that is pre-stored locally and has completed training, convert each word or word group into the corresponding word vector. For example: loan [0.12, -0.45,..., 0.78]; 200,000 yuan [0.34, 0.21,..., -0.56]; in 3 years [-0.09, 0.67,..., 0.12]... Assume that all word weights are the same, and calculate the first word vector by the method of weighted average. Similarly, calculate the second word vector according to the same steps. Finally, substitute the first word vector and the second word vector into the aforementioned cosine similarity formula to calculate the cosine angle between the two, and the cosine angle value can be obtained. Assume that the finally calculated cosine angle value is 0.92, and the preset angle threshold is 0.8. Since 0.92 is greater than 0.8, the similarity between the speech text content of customer A's answer and the preset speech template is high, that is, the speech text content of customer A matches the preset speech template.
[0073] S140. If the speech text content matches the preset speech template, perform mouth movement detection on the target video segment according to the time stamp and generate a mouth movement detection result.
[0074] Specifically, after obtaining the speech text content corresponding to the target video segment and the time stamp of each word in the speech text content through speech recognition technology, match the speech text content with the preset speech template. If the match is successful, perform mouth movement detection on the target video segment according to the time stamp and generate a mouth movement detection result; if the match is unsuccessful, obtain the video segment corresponding to each word in the speech text content according to the time stamp, and delete or discard the video segment.
[0075] Specifically, in some embodiments, as Figure 3 shown, S140 includes the following steps:
[0076] S141. Obtain multiple target image frames from the target video segment according to the time stamp.
[0077] S142. Determine the mouth movement probability corresponding to each of the target image frames based on a preset mouth movement model.
[0078] S143. Determine the mouth movement confidence of the target video segment according to each of the mouth movement probabilities.
[0079] S144. When the mouth movement confidence is greater than a preset confidence threshold, regard the mouth movement as the mouth movement detection result.
[0080] S145. When the mouth movement confidence is less than or equal to the confidence threshold, regard the mouth not moving as the mouth movement detection result.
[0081] Specifically, open the target video segment through a video processing library (such as OpenCV) function to establish a connection with the video file. The time stamps are stored in a list. Obtain the frame rate of the target video segment through the video processing library, and convert the time stamps into corresponding video frame numbers. Extract the corresponding target image frames frame by frame from the target video segment according to the video frame numbers. By traversing the video frame numbers corresponding to the time stamps and using the functions of the video processing library to read and save these frames, multiple target image frames can be obtained.
[0082] Then, a mouth movement model is trained by a deep neural network (e.g., a model trained based on a lightweight network structure such as MobileNetV3). Among them, a lightweight network model (such as MobileNetV3) is selected for predicting the mouth movement probability. On the premise of ensuring a certain accuracy, the amount of calculation and memory occupation are reduced, enabling this method to run relatively efficiently on different hardware devices. The mouth region image in each target image frame is preprocessed (e.g., normalized to scale the pixel values to the range of 0-1), and then the preprocessed target image frame is input into the locally stored and pre-trained mouth movement model for prediction. Using the pre-trained mouth movement model can effectively capture the mouth movement features and improve the accuracy of predicting the mouth movement probability.
[0083] The mouth movement model will output a probability value for each target image frame respectively (i.e., the mouth movement probability corresponding to each target image frame). The probability value represents the possibility that the mouth in the target image frame is in a moving state, and its value range is between 0 and 1. By weighted averaging of the probability values, the mouth movement confidence of the target video segment can be obtained. The method of calculating the confidence based on probability comprehensively considers the information of multiple target image frames, further improving the reliability of the detection result. By comparing the obtained mouth movement confidence with the preset confidence threshold, the mouth movement detection result can be obtained.
[0084] For example, in a loan application process, in order to confirm the authenticity of the identity of the loan applicant and the compliance of the operation, the bank will require the applicant to record a video, and the video content is the applicant reading the loan contract terms.
[0085] The bank system extracts the image frames corresponding to the reading of each clause from the video of the applicant reading the loan contract terms according to the timestamp information in the video content. For example, when the applicant reads the sentence "I promise to repay the principal and interest of the loan on time and in full", the system accurately extracts multiple image frames during the reading of this sentence based on the timestamp. For each extracted image frame, after preprocessing, it is input into the pre-trained mouth movement model to obtain the mouth movement probability in each image frame. By weighted averaging the mouth movement probabilities in each image frame, the mouth movement confidence of the video corresponding to the applicant reading the sentence "I promise to repay the principal and interest of the loan on time and in full" can be obtained. Suppose it is 0.1. Suppose the confidence threshold preset by the bank is 0.4. Since the obtained mouth movement confidence is 0.1, which is less than the confidence threshold preset by the bank, the bank can consider that the applicant's mouth is not moving, or there may be a behavior of someone else answering on behalf of the applicant. The bank can then take measures to further verify the situation to ensure the authenticity and compliance of the loan application.
[0086] S150. If the mouth movement detection result indicates mouth movement, store the voiceprint data corresponding to the voice text content into the voiceprint database corresponding to the target event.
[0087] Specifically, for example, in the face-to-face review video of a housing loan application, the staff of the bank will ask multiple questions regarding the applicant's situation and simultaneously record the video data of their answers. The bank system will perform mouth movement detection on the applicant's answer video data. When the detection result indicates mouth movement, it shows that the applicant himself / herself is answering the questions of the bank staff, that is, there is no behavior of an outsider answering on behalf.
[0088] The system intercepts the audio clip of the applicant's answer to the questions from the audio corresponding to the face-to-face review video according to the timestamp of each word in the voice text content, performs preprocessing such as noise reduction and normalization on it, and then extracts the voiceprint data using voiceprint recognition technology. Since this business scenario belongs to the target event of housing loan application, the system stores the voiceprint data into the voiceprint database corresponding to housing loan application.
[0089] The voiceprint data extracted in the subsequent question-answering session of the face-to-face review can be compared and verified with the voiceprint data stored in the current voiceprint database to confirm the identity of the applicant in real time and effectively prevent the risk of identity fraud in loan business.
[0090] S160. If the mouth movement detection result indicates no mouth movement, obtain the historical voiceprint data corresponding to the target event from the voiceprint database, and detect the target voiceprint data of the target video segment based on the historical voiceprint data to obtain the detection result of outsider answering on behalf.
[0091] Specifically, multiple answer data of the target customer are recorded in the answer video data of the target customer for the target event. The server real-time detects the mouth movement behavior of the target customer when answering questions. At the same time, the voiceprint data corresponding to the voice text content with the detection result of mouth movement is stored into the voiceprint database corresponding to the target event. When the detection result for the target customer's answer to a certain question indicates no mouth movement, at this time, obtain the historical voiceprint data corresponding to the target event in the voiceprint database, and detect the target voiceprint data of the target video segment based on the historical voiceprint data, so as to obtain the detection result of outsider answering on behalf. If there is voice output when there is no mouth movement, there is likely an abnormal situation of an outsider answering on behalf. Through subsequent detection based on historical voiceprint data, such potential risks can be captured in a timely manner. For example, in scenarios such as loans and important contract signings, it can ensure the authenticity and security of the business and prevent fraud behavior from occurring.
[0092] For example, in the face-to-face signing video of applying for a car loan from the bank, the staff on the bank side will ask a series of questions according to the process to confirm her identity. Suppose among the 5 questions asked by the bank staff, the mouth movement detection results for Xiaohong's answers to the first 4 questions are all mouth movements, and the mouth movement detection result for Xiaohong's answer to the 5th question is no mouth movement. Then, the voiceprint data corresponding to the voice text content of Xiaohong's answers to the first 4 questions is stored in the voiceprint database corresponding to the car loan. At the same time, the voiceprint data of the first 4 questions corresponding to the car loan is retrieved from the voiceprint database, and the voiceprint data of the 5th question is detected based on the voiceprint data of the first 4 questions, so as to obtain the detection result of whether the answer to Xiaohong's 5th question is a behavior of being answered by someone off-camera.
[0093] Specifically, in some embodiments, the detecting the target voiceprint data of the target video segment according to the historical voiceprint data to obtain the detection result of the behavior of being answered by someone off-camera includes the following steps:
[0094] Obtain the number of voiceprints of the historical voiceprint data;
[0095] If the number of voiceprints is greater than a preset voiceprint number threshold, detect the target voiceprint data according to the historical voiceprint data to obtain the detection result of the behavior of being answered by someone off-camera.
[0096] Specifically, obtain the historical voiceprint data stored in the voiceprint database, compare the number of voiceprints of the historical voiceprint data with the preset voiceprint number threshold. If the number of voiceprints is greater than the preset voiceprint number threshold, detect the target number of voiceprints according to the historical voiceprint data to obtain the detection result of the behavior of being answered by someone off-camera. Rich data samples can more comprehensively reflect the variation range of the voiceprint characteristics of the target object, effectively reduce misjudgments caused by individual differences in voiceprint characteristics or environmental factors, and make the detection result more accurate and reliable. If the number of voiceprints is less than the preset voiceprint number threshold, do not detect the target voiceprint data. When the historical voiceprint data is less, no detection is performed, avoiding waste of invalid computing resources due to insufficient data.
[0097] For example, during the application process of a large commercial loan, the bank requires the applicant to answer a series of key questions about the loan purpose, repayment plan, etc. during the video interview. After the bank's system detects that the applicant's mouth is not moving but there is a voice answering the question (the voiceprint data of the video segment where the applicant's mouth is not moving but there is a voice answering the question is the target voiceprint data), it retrieves the voiceprint database of the large commercial loan and compares the number of voiceprints recorded in the voiceprint database with the preset voiceprint number threshold. Suppose the preset voiceprint number threshold is 3, and the number of voiceprints recorded in the voiceprint database is 4. Since 4 is greater than 3, the target voiceprint data of this record can be detected according to the historical voiceprint data, so as to obtain the detection result of the behavior of being answered by someone off-camera.
[0098] Further, in some embodiments, detecting the target voiceprint data according to the historical voiceprint data to obtain the detection result of the off-screen person's proxy answer behavior includes the following steps:
[0099] Obtain the voiceprint similarity values between each of the historical voiceprint data and the target voiceprint data;
[0100] If there is historical voiceprint data with the voiceprint similarity value less than a preset voiceprint similarity threshold, then taking the existence of the off-screen person's proxy answer behavior as the detection result of the off-screen person's proxy answer behavior.
[0101] Specifically, for the audio of the target video segment, use a professional voiceprint recognition tool (for example, the Kaldi voiceprint recognition toolkit) to extract the feature vectors of the target voiceprint data. At the same time, also extract the feature vectors for each piece of historical voiceprint data obtained. Adopt a voiceprint similarity calculation algorithm (for example, adopt the cosine similarity algorithm), and calculate the feature vectors of the target voiceprint data and the feature vectors of each piece of historical voiceprint data one by one, so as to obtain the voiceprint similarity values between each historical voiceprint data and the target voiceprint data.
[0102] Pre-set a voiceprint similarity threshold in the system, for example, 0.7. Compare the calculated voiceprint similarity values with 0.7 one by one. As long as there is a voiceprint similarity value between a certain piece of historical voiceprint data and the target voiceprint data less than this threshold, it is determined that there is an off-screen person's proxy answer behavior, and this is used as the detection result of the off-screen person's proxy answer behavior; if all voiceprint similarity values are greater than or equal to the threshold, it is determined that it is the person himself / herself answering, that is, taking the non-existence of the off-screen person's proxy answer behavior as the detection result of the off-screen person's proxy answer behavior. In scenarios involving important services, such as loan services, this strict detection method can effectively prevent fraud behaviors, ensure that the person himself / herself participates in the service process, strongly guarantee the authenticity and security of the service, and safeguard the legitimate rights and interests of financial institutions and users.
[0103] In summary, in the embodiments of the present application, on the basis of determining the target video segment to be detected currently from multiple video segments of the target customer's answer video data for the target event, the corresponding speech text content and the time stamp of each word are obtained through speech recognition technology. If the speech text content matches the preset speech template, the mouth movement of the target video segment is detected according to the time stamp to generate a detection result. If the detection result is mouth movement, the voiceprint data corresponding to the speech text content is stored in the voiceprint library corresponding to the target event. If the detection result is that the mouth does not move, the historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and accordingly, the target voiceprint data of the target video segment is detected to obtain the detection result of the off-screen person's proxy answer behavior, thereby improving the accuracy of the detection of the off-screen person's proxy answer behavior.
[0104] Figure 4 is a schematic block diagram of the answering-on-behalf behavior detection device provided by an embodiment of the present application. As Figure 4 shown, corresponding to the above answering-on-behalf behavior detection method, the present invention also provides an answering-on-behalf behavior detection device 1000. The answering-on-behalf behavior detection device 1000 includes units for executing the above answering-on-behalf behavior detection method. Specifically, please refer to Figure 4 The answering-on-behalf behavior detection device 1000 includes a cutting unit 1001, a determining unit 1002, an obtaining unit 1003, a first detection unit 1004, a storage unit 1005, and a second detection unit 1006, where:
[0105] The cutting unit 1001 is configured to obtain the answering video data of the target customer for the target event, and cut the answering video data into multiple video segments;
[0106] The determining unit 1002 is configured to determine the target video segment to be detected currently from the multiple video segments;
[0107] The obtaining unit 1003 is configured to use speech recognition technology to obtain the speech text content corresponding to the target video segment and the time stamp of each word in the speech text content;
[0108] The first detection unit 1004 is configured to, if the speech text content matches a preset speech template, perform mouth movement detection on the target video segment according to the time stamp and generate a mouth movement detection result;
[0109] The storage unit 1005 is configured to, if the mouth movement detection result is mouth movement, store the voiceprint data corresponding to the speech text content into the voiceprint library corresponding to the target event;
[0110] The second detection unit 1006 is configured to, if the mouth movement detection result is no mouth movement, obtain the historical voiceprint data corresponding to the target event from the voiceprint library, and detect the target voiceprint data of the target video segment according to the historical voiceprint data to obtain an off-screen answering-on-behalf behavior detection result.
[0111] In some embodiments, the cutting unit 1001 is specifically configured to:
[0112] Perform cutting processing on the answering video data through a preset voice activity detection model to obtain multiple video segments.
[0113] In some embodiments, the first detection unit 1004 is specifically configured to:
[0114] Obtain multiple target image frames from the target video segment according to the time stamp;
[0115] Determine the mouth movement probability corresponding to each of the target image frames based on a preset mouth movement model;
[0116] Determine the mouth movement confidence of the target video segment according to each of the mouth movement probabilities;
[0117] When the mouth movement confidence is greater than a preset confidence threshold, take the mouth movement as the mouth movement detection result;
[0118] When the mouth movement confidence is less than or equal to the confidence threshold, take the mouth not moving as the mouth movement detection result.
[0119] In some embodiments, the second detection unit 1006 is specifically configured to:
[0120] Obtain the number of voiceprints of the historical voiceprint data;
[0121] If the number of voiceprints is greater than a preset voiceprint number threshold, detect the target voiceprint data according to the historical voiceprint data to obtain the off-screen voice answering behavior detection result.
[0122] Furthermore, the second detection unit 1006 is also specifically configured to:
[0123] Obtain the voiceprint similarity values between each of the historical voiceprint data and the target voiceprint data;
[0124] If there is historical voiceprint data with a voiceprint similarity value less than a preset voiceprint similarity threshold, take the existence of an off-screen voice answering behavior as the off-screen voice answering behavior detection result.
[0125] In some embodiments, before the determination unit 1002, the voice answering behavior detection device 1000 further includes:
[0126] A third detection unit, configured to detect whether there are video segments in multiple video segments that have not been subjected to off-screen voice answering behavior detection;
[0127] At this time, the determination unit 1002 is specifically configured to:
[0128] If there are video segments in multiple video segments that have not been subjected to off-screen voice answering behavior detection, determine a video segment from the video segments that have not been subjected to off-screen voice answering behavior detection as the target video segment.
[0129] In some embodiments, before the first detection unit 1004, the voice answering behavior detection device 1000 further includes:
[0130] A conversion unit, configured to convert the speech text content into a first word vector and convert a preset speech template into a second word vector based on a preset word vector model;
[0131] A calculation unit, configured to calculate a cosine angle between the first word vector and the second word vector to obtain a cosine angle value;
[0132] A first determination subunit, configured to determine that the speech text content matches the preset speech template if the cosine angle value is less than or equal to a preset angle threshold;
[0133] A second determination subunit, configured to determine that the speech text content does not match the preset speech template if the cosine angle value is greater than the preset angle threshold.
[0134] In summary, based on determining a target video segment to be detected currently from multiple video segments of video data answered by a target customer for a target event, the answering-on-behalf behavior detection device 1000 in the embodiment of the present application obtains the corresponding speech text content and the time stamp of each character thereof through speech recognition technology. If the speech text content matches a preset speech template, a mouth movement detection is performed on the target video segment according to the time stamp to generate a detection result. If the detection result is mouth movement, the voiceprint data corresponding to the speech text content is stored in a voiceprint library corresponding to the target event. If the detection result is no mouth movement, historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and based on this, a detection is performed on the target voiceprint data of the target video segment to obtain an answering-on-behalf behavior detection result of an off-screen person, thereby improving the accuracy of the answering-on-behalf behavior detection of the off-screen person.
[0135] The above answering-on-behalf behavior detection device can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 5 shown.
[0136] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 800 is a terminal or a server.
[0137] Referring to Figure 5 , the computer device 800 includes a processor 802, a memory, and a network interface 805 connected through a system bus 801. Among them, the memory may include a non-volatile storage medium 803 and an internal memory 804.
[0138] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions, and when the program instructions are executed, the processor 802 can be made to execute an answering-on-behalf behavior detection method.
[0139] The processor 802 is used to provide computing and control capabilities to support the operation of the entire computer device 800.
[0140] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can be made to execute a method for detecting a proxy answering behavior.
[0141] The network interface 805 is used for network communication with other devices. Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 800 to which the solution of this application is applied. The specific computer device 800 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0142] Among them, the processor 802 is used to run the computer program 8032 stored in the memory to implement the following steps:
[0143] Obtain the answer video data of the target customer for the target event, and cut the answer video data into multiple video segments;
[0144] Determine the target video segment to be detected currently from the multiple video segments;
[0145] Use speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content;
[0146] If the speech text content matches the preset speech template, perform mouth movement detection on the target video segment according to the timestamp and generate a mouth movement detection result;
[0147] If the mouth movement detection result is mouth movement, store the voiceprint data corresponding to the speech text content into the voiceprint library corresponding to the target event;
[0148] If the mouth movement detection result is no mouth movement, obtain the historical voiceprint data corresponding to the target event from the voiceprint library, and detect the target voiceprint data of the target video segment according to the historical voiceprint data to obtain a detection result of a proxy answering behavior by an off-screen person.
[0149] It should be understood that in the embodiments of the present application, the processor 802 may be a central processing unit (CPU), and the processor 802 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0151] Therefore, the present application also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:
[0152] Obtain the response video data of the target customer for the target event, and cut the response video data into multiple video segments;
[0153] Determine the target video segment to be detected currently from the multiple video segments;
[0154] Use speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content;
[0155] If the speech text content matches the preset speech template, perform mouth movement detection on the target video segment according to the timestamp and generate a mouth movement detection result;
[0156] If the mouth movement detection result is mouth movement, store the voiceprint data corresponding to the speech text content into the voiceprint library corresponding to the target event;
[0157] If the mouth movement detection result indicates that the mouth is not moving, obtain the historical voiceprint data corresponding to the target event from the voiceprint database, and detect the target voiceprint data of the target video segment based on the historical voiceprint data to obtain the detection result of the off-screen voice substitution behavior.
[0158] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which can store program codes.
[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0160] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0161] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0162] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the essence of the technical solution of this application, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.
[0163] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for detecting answering behavior, characterized in that: include: Acquire the target customer's answer video data for the target event, and cut the answer video data into multiple video clips; Determine a target video segment that currently needs to be detected from multiple video segments; Using speech recognition technology to obtain the speech text content corresponding to the target video clip and the timestamp of each word in the speech text content; If the speech text content matches the preset speech template, performing mouth movement detection on the target video clip according to the timestamp and generating a mouth movement detection result; If the mouth movement detection result is mouth movement, storing the voiceprint data corresponding to the speech text content in the voiceprint library corresponding to the target event; If the mouth movement detection result is that the mouth does not move, historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and the target voiceprint data of the target video clip is detected based on the historical voiceprint data to obtain the off-screen person answering the phone behavior detection result.
2. The method for detecting answering behavior according to claim 1, characterized in that: The performing mouth movement detection on the target video segment according to the timestamp and generating a mouth movement detection result includes: Acquire a plurality of target image frames from the target video segment according to the timestamp; Determining the mouth movement probability corresponding to each of the target image frames based on a preset mouth movement model; Determining a mouth movement confidence of the target video segment according to each of the mouth movement probabilities; When the mouth movement confidence is greater than a preset confidence threshold, taking the mouth movement as the mouth movement detection result; When the mouth movement confidence is less than or equal to the confidence threshold, the absence of mouth movement is taken as the mouth movement detection result.
3. The method for detecting answering behavior according to claim 1, characterized in that: The detecting the target voiceprint data of the target video clip according to the historical voiceprint data to obtain the off-screen person answering behavior detection result includes: Obtain the number of voiceprints of the historical voiceprint data; If the number of voiceprints is greater than a preset voiceprint number threshold, the target voiceprint data is detected according to the historical voiceprint data to obtain the detection result of the off-screen person answering the call.
4. The method for detecting answering behavior according to claim 3, characterized in that: The detecting the target voiceprint data according to the historical voiceprint data to obtain the detection result of the off-screen person answering the phone includes: Obtaining voiceprint similarity values between each of the historical voiceprint data and the target voiceprint data; If there is historical voiceprint data whose voiceprint similarity value is less than a preset voiceprint similarity threshold, the existence of a person outside the picture answering the question on behalf of another person will be taken as the person outside the picture answering question on behalf of another person detection result.
5. The method for detecting answering behavior according to claim 1, characterized in that: The step of cutting the answer video data into a plurality of video segments comprises: The answer video data is segmented and processed by using a preset voice activity detection model to obtain a plurality of video segments.
6. The method for detecting answering behavior according to claim 1, characterized in that: Before determining the target video segment currently to be detected from the multiple video segments, the method includes: Detecting whether there is a video clip among the multiple video clips that has not been detected for the off-screen person answering behavior; The step of determining a target video segment that needs to be detected currently from a plurality of video segments includes: If there is a video segment in the plurality of video segments for which the off-screen person answering the question behavior detection has not been performed, a video segment is determined from the video segments for which the off-screen person answering the question behavior detection has not been performed as the target video segment.
7. The method for detecting answering behavior according to claim 1, characterized in that: If the speech text content matches the preset speech template, before performing mouth movement detection on the target video segment according to the timestamp and generating a mouth movement detection result, the method includes: Based on a preset word vector model, convert the speech text content into a first word vector, and convert the preset speech template into a second word vector; Calculate the cosine angle between the first word vector and the second word vector to obtain a cosine angle value; If the cosine angle value is greater than or equal to a preset angle threshold, it is determined that the speech text content matches the preset speech template; If the cosine angle value is less than the preset angle threshold, it is determined that the speech text content does not match the preset speech template.
8. A device for detecting answering calls on behalf of others, characterized in that: include: A cutting unit, used for obtaining the target customer's answer video data for the target event, and cutting the answer video data into multiple video clips; A determination unit, used to determine a target video segment that currently needs to be detected from multiple video segments; An acquisition unit, configured to acquire the speech text content corresponding to the target video clip and the timestamp of each word in the speech text content by using speech recognition technology; A first detection unit, configured to perform mouth movement detection on the target video clip according to the timestamp and generate a mouth movement detection result if the speech text content matches a preset speech template; a storage unit, configured to store the voiceprint data corresponding to the speech text content into a voiceprint library corresponding to the target event if the mouth movement detection result is mouth movement; The second detection unit is used to obtain historical voiceprint data corresponding to the target event from the voiceprint library if the mouth movement detection result is that the mouth does not move, and detect the target voiceprint data of the target video clip based on the historical voiceprint data to obtain the detection result of the off-screen person answering the question behavior.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for detecting answering behavior as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting answering behavior on behalf of another as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice detection method based on multiple sound areas, related device and storage medium
CN111833899A
Cosmetic live broadcast efficacy verbal skill monitoring method and system
CN114168711A
Double-recording method based on speaker detection and related equipment
CN117649691A
Sound and lip synchronous identification method, device and equipment and storage medium
CN118334554A
Over-person answering-on-behalf detection method based on face-to-face video and related equipment
CN119206566A