Answering behavior detection method, device, equipment and medium
By segmenting online interview video data and detecting speech and mouth movements, combined with voiceprint data analysis, the problem of detecting unauthorized responses during online interviews has been solved, improving the accuracy and security of the detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-03-25
- Publication Date
- 2026-04-10
AI Technical Summary
Existing online verification methods pose a risk of fraud, especially since it is difficult to accurately identify instances of unauthorized individuals answering on behalf of others, a problem that current mouth movement detection technologies struggle to effectively address.
By acquiring video data of the target customer's answers, cutting it into multiple video segments, using speech recognition technology to obtain the speech text content and timestamps, and combining mouth movement detection and voiceprint data, the system generates a result for detecting off-screen person answering on behalf of the user.
It improves the accuracy of detecting unauthorized responses, reduces the risk of false detections, and ensures the authenticity and security of loan, insurance, and other business transactions.
Smart Images

Figure CN120148513B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of financial technology, and in particular to a method and device for detecting an answering behavior, and a related equipment and medium. BACKGROUND
[0002] With the development of Internet technology, the face-to-face audit mode in current loan and insurance scenarios is gradually changing from on-site audit to online audit, which greatly facilitates people's life. However, due to the limited camera angle, the online audit mode inevitably has fraud risks (such as being guided by others or being answered by others).
[0003] The existing technology based on mouth movement detection can solve part of the irregular behavior detection, such as detecting whether the mouth moves when the insured and the agent speak, which can avoid part of the irregular behavior. However, for irregular behavior outside the picture, it is difficult to accurately identify the irregular behavior outside the picture only by detecting the mouth movement of the customer. Therefore, there is an urgent need for a method that can improve the accuracy of the out-of-picture answering behavior detection to reduce the fraud risk of being answered by others. SUMMARY
[0004] The embodiments of the present application provide a method and device for detecting an answering behavior, which can improve the accuracy of the out-of-picture answering behavior detection.
[0005] In a first aspect, the embodiments of the present application provide a method for detecting an answering behavior, comprising:
[0006] obtaining answer video data of a target customer for a target event, and cutting the answer video data into a plurality of video clips;
[0007] determining a target video clip that needs to be detected from the plurality of video clips;
[0008] obtaining the voice text content corresponding to the target video clip and the time stamp of each character in the voice text content by using a voice recognition technology;
[0009] if the voice text content matches a preset script template, performing mouth movement detection on the target video clip according to the time stamp and generating a mouth movement detection result;
[0010] if the mouth movement detection result is mouth movement, storing the voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event;
[0011] if the mouth movement detection result is no mouth movement, obtaining historical voiceprint data corresponding to the target event from the voiceprint library, and detecting target voiceprint data of the target video clip according to the historical voiceprint data to obtain an out-of-picture answering behavior detection result.
[0012] In a second aspect, the embodiments of the present application further provide a device for detecting a stand-in behavior, which comprises:
[0013] a cutting unit, configured to acquire answer video data of a target client for a target event, and cut the answer video data into a plurality of video clips;
[0014] a determining unit, configured to determine a target video clip which needs to be detected currently from the plurality of video clips;
[0015] an acquiring unit, configured to acquire voice text content corresponding to the target video clip and a time stamp of each word in the voice text content by using a voice recognition technology;
[0016] a first detecting unit, configured to, if the voice text content matches a preset script template, perform a mouth movement detection on the target video clip according to the time stamp and generate a mouth movement detection result;
[0017] a storing unit, configured to, if the mouth movement detection result is mouth movement, store voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event;
[0018] a second detecting unit, configured to, if the mouth movement detection result is no mouth movement, acquire historical voiceprint data corresponding to the target event from the voiceprint library, and perform a detection on target voiceprint data of the target video clip according to the historical voiceprint data, to obtain a stand-in behavior detection result.
[0019] In a third aspect, the embodiments of the present application further provide a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above method when executing the computer program.
[0020] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the above method.
[0021] The embodiment of the application provides a kind of answering behavior detection method, device, equipment and medium.The method comprises: obtaining the answering video data of target customer to target event, and the answering video data is cut into multiple video clips;From multiple video clips, determine the target video clip currently needing detection;The voice text content corresponding to the target video clip is obtained using speech recognition technology and the time stamp of each word in the voice text content;If the voice text content matches the preset dialogue template, then according to the time stamp, mouth movement detection is carried out on the target video clip and mouth movement detection result is generated;If the mouth movement detection result is mouth movement, then the voiceprint data corresponding to the voice text content is stored in the voiceprint library corresponding to the target event;If the mouth movement detection result is mouth not movement, then the historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and the target voiceprint data of the target video clip is detected according to the historical voiceprint data, to obtain the off-screen person answering behavior detection result.In the embodiment of the application, the target video clip currently needing detection is determined from multiple video clips of the answering video data of target customer to target event, and the voice text content corresponding thereto and the time stamp of each word are obtained by speech recognition technology, if the voice text content matches the preset dialogue template, then according to the time stamp, mouth movement detection is carried out on the target video clip and mouth movement detection result is generated, if the detection result is mouth movement, then the voiceprint data corresponding to the voice text content is stored in the voiceprint library corresponding to the target event, if the detection result is mouth not movement, then the historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and the target voiceprint data of the target video clip is detected according to the historical voiceprint data, to obtain the off-screen person answering behavior detection result, so as to improve the accuracy of off-screen person answering behavior detection. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without any creative effort.
[0023] Figure 1 An application scenario diagram of the answering behavior detection method provided by the embodiment of the application is provided.
[0024] Figure 2 A flowchart of the answering behavior detection method provided by the embodiment of the application is provided.
[0025] Figure 3 A sub-flowchart of the answering behavior detection method provided by the embodiment of the application is provided.
[0026] Figure 4a schematic block diagram of a computer device provided by an embodiment of the present application.
[0027] Figure 5 a schematic block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of the present application.
[0029] It should be understood that, when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0030] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0031] It should be further understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0032] The embodiments of the present application provide a method, device, equipment and medium for detecting a proxy answering behavior. The execution subject of the method for detecting a proxy answering behavior can be a device for detecting a proxy answering behavior provided by the embodiments of the present application, or a computer device integrated with the device for detecting a proxy answering behavior. The device for detecting a proxy answering behavior can be realized in the form of hardware or software, and the computer device can be a terminal or a server. The embodiments of the present application take the computer device as an example to describe the method for detecting a proxy answering behavior provided by the present application in detail.
[0033] Please refer to Figure 1 , Figure 1 The application scenario of the method for detecting a proxy answering behavior provided by the embodiments of the present application is shown in the figure. In an embodiment, the application scenario includes:
[0034] The method for detecting a proxy answering behavior is applied to Figure 1In some embodiments, the computer device obtains a plurality of video clips of answer video data of a target customer for a target event, and determines a target video clip to be currently required to be detected from the plurality of video clips. The computer device obtains voice text content corresponding to the target video clip and a time stamp of each word in the voice text content through voice recognition technology. If the voice text content matches a preset script template, the computer device generates a detection result by performing mouth movement detection on the target video clip according to the time stamp. If the detection result is mouth movement, the computer device stores voiceprint data corresponding to the voice text content to a voiceprint library corresponding to the target event. If the detection result is no mouth movement, the computer device obtains historical voiceprint data corresponding to the target event from the voiceprint library, and detects target voiceprint data of the target video clip to obtain an off-screen answer behavior detection result. The computer device detects the off-screen answer behavior through mouth movement detection and voiceprint data detection, reduces the risk of false detection, and improves the accuracy of off-screen answer behavior detection.
[0035] It should be noted that the application scenarios of the above-mentioned answer behavior detection method are only used to illustrate the technical solutions of the present application, and are not used to limit the technical solutions of the present application. The above-mentioned connection relationship can also have other forms.
[0036] Figure 2 A flowchart of an answer behavior detection method provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the method includes the following steps S110-S160. Figure 2
[0037] S110, obtaining answer video data of a target customer for a target event, and cutting the answer video data into a plurality of video clips.
[0038] The target customer refers to a customer individual or group selected as an analysis and research object in a specific business scenario. The target event refers to an event associated with the target customer and having specific research significance in a specific business scenario. For example, in a video interview scenario of a certain housing loan, the creditor will ask the debtor some questions to confirm the debtor's ability to perform the debt, and the debtor participating in the video interview is the target customer, and the video interview for the target customer about the certain housing loan is the target event.
[0039] Specifically, the answer video data of the target customer for the target event can be obtained by establishing a connection with a video acquisition device (e.g., a camera) or a storage system. The connection method can be selected according to the actual situation, such as using a network interface to call an application programming interface (API) to obtain from a remote server or directly reading from a local storage device. After obtaining the answer video data of the target customer for the target event, the answer video data is checked. The content of the check includes form checking and metadata checking. The form checking can include checking whether the format of the video file is correct, whether the file size meets the expectation, and whether the encoding method of the video is compatible with the subsequent processing tool. If the format of the video file is correct, the file size meets the expectation, and the encoding method of the video is compatible with the subsequent processing tool, the form checking result is qualified. For example, if the subsequent video processing software only supports the common MP4 format, and the obtained video is in AVI format, format conversion is needed. The metadata checking can include checking whether the shooting time, resolution, frame rate, and other data information are recorded in the answer video data. If the shooting time, resolution, frame rate, and other data information are recorded in the answer video data, the metadata checking result is qualified. By checking the shooting time, resolution, frame rate, and other data information, it can be ensured that the obtained video data is complete and usable, and the problem of unsuccessful or incorrect processing of the video data in subsequent processing due to data problems is avoided.
[0040] The answer video data with both the form verification result and the metadata verification result being qualified is cut using a specific audio and video processing tool (for example, FFmpeg) or a preset model. A certain length of answer video data is cut by the specific tool or the preset model to generate a plurality of video clips. For example, an open-source cross-platform audio and video processing tool named FFmpeg is used to cut the answer video data of an education student loan video interview. First, the total length of the answer video data of the interview is 15 minutes, and the file name of the answer video data of the interview is input.mp4. Second, according to the set parameters, the answer video data is cut according to the specified start time, end time and clip length. For example, the command "ffmpeg -i input.mp4 -ss 00:00:10 -t 00:00:30 -c copy output1.mp4" of FFmpeg is used, which means that the 30-second clip is cut from the 10th second of the input.mp4 video and saved as output1.mp4. Since the total length of the answer video data of the education student loan video interview is 15 minutes, the answer video data of the education student loan video interview can be cut into a plurality of video clips in the same way and output. Finally, after cutting, the generated video clips are checked. The video player can be used to play these video clips one by one to check whether the cutting time point is accurate, whether the video content is complete, whether the sound is normal, etc. If a problem is found in a certain video clip, the cutting time point can be adjusted according to the specific circumstances, and the cutting command can be executed again.
[0041] Through the above steps, the FFmpeg tool can conveniently cut the answer video of the education student loan video interview into a plurality of independent video clips, which facilitates the in-depth analysis of the answers to each question in the answer video data and provides stronger support for the authenticity and accuracy of the answers of the subsequent loan auditors.
[0042] Further, in some embodiments, the cutting of the answer video data into a plurality of video clips includes the following steps:
[0043] The answer video data is cut by a preset voice activity detection model to obtain a plurality of video clips.
[0044] Specifically, the preset voice activity detection model (VAD) is a voice processing technology mainly used to detect the presence of human voice signals. The VAD model can accurately identify the voice part, avoiding including a large amount of meaningless silence or noise segments in the cut video, so that the subsequent analysts can focus on valuable voice content and improve analysis efficiency. For example, in the answer video of loan face interview, there may be short pauses or ambient noise when the interviewer and the applicant communicate, and the VAD model can effectively filter these parts and only keep the conversation content of the two parties. The entire cutting process is automatically completed based on the preset VAD model, reducing the workload of manual marking and cutting and reducing the possibility of human error. In addition, when processing a large number of insurance customer follow-up videos, manual cutting requires a lot of time and effort, while using the VAD model for automatic cutting can complete the task in a short time and improve work efficiency.
[0045] Before cutting the answer video data through the preset VAD model, the audio track is extracted from the answer video data using a professional audio and video processing library (such as FFmpeg or moviepy). Then, the extracted audio is preprocessed to meet the input requirements of the VAD model. Preprocessing includes but is not limited to audio format conversion (such as converting the audio sampling rate to a specific value required by the VAD model, commonly 16kHz), normalizing the audio volume, and framing the audio. For example, the pydub library in Python can be used for audio format conversion and volume normalization, adjusting the volume of the audio to a certain standard range to improve the accuracy of the VAD model detection. The preprocessed audio data is input into the loaded and initialized VAD model. The VAD model analyzes and judges each frame of audio data to determine whether it belongs to the voice part, and outputs the judgment result. The output result of the VAD model can be in the form of a label sequence, for example, "1" represents that the frame contains voice, and "0" represents that the frame is silent or non-voice (such as ambient noise)
[0046] According to the output result of the VAD model, the cutting points of the video segment can be determined. According to the start and end positions of the voice, it is mapped to the time axis of the answer video segment. For example, if a voice is detected starting from the 5th second of the audio and ending at the 15th second, then in the answer video segment, this is a potential cutting interval. When determining the cutting point, a certain buffer time can also be added, such as increasing 0.5 seconds before and after the start and end of the voice, to ensure that the voice content is completely included and avoid the voice part being truncated due to detection errors.
[0047] For example, in the housing loan face-to-face signing video review scene, the face-to-face signing officer will ask the loaner a series of questions, such as income, debt, and housing purchase purpose. By presetting the VAD model to cut the entire face-to-face signing video, a plurality of video clips about the housing loan face-to-face signing video can be obtained. Subsequently, the answers of the loaner in the video clips are analyzed to reduce the risk of a third party replacing the loaner to answer the review questions. Specifically, the trained VAD model is loaded, and the audio in the answer video data is extracted and preprocessed. The audio is input into the trained VAD model for voice activity detection. When the voice part of the loaner's answer to the income question is detected, the time range thereof in the video is determined, such as from the 30th second to the 60th second of the video. Then, a video editing tool is used to cut out this video segment. Subsequently, the review personnel can directly analyze the language expression and tone of the loaner when answering the income-related question in this video segment, judge the authenticity and reliability of the income information, and greatly improve the efficiency and accuracy of the loan review.
[0048] Specifically, in some embodiments, before step S120, the method comprises the following steps:
[0049] detecting whether there is a video segment in the plurality of video segments that has not been subjected to the third-party answer behavior detection;
[0050] The determining of the target video segment from the plurality of video segments comprises:
[0051] If there is a video segment in the plurality of video segments that has not been subjected to the third-party answer behavior detection, a video segment that has not been subjected to the third-party answer behavior detection is determined as the target video segment.
[0052] The target video segment refers to a certain video segment that needs to be detected from the plurality of video segments cut from the answer video data.
[0053] Specifically, taking the housing loan video face-to-face signing scene as an example. In the housing loan video face-to-face signing scene, when the bank staff (the loan party) and the borrower conduct video face-to-face signing, the loan party will ask the borrower some questions (for example, inquire about the borrower's income, liabilities, housing purchase purpose, etc.) to ensure that the borrower has the ability to perform the repayment obligation. The process of the borrower answering the questions raised by the loan party during the face-to-face signing process will be recorded throughout to form the answer video data of the target customer to the target event. The bank obtains the answer video data from the credit business system and cuts the answer video data according to the question categories, such as the answer segment of the income question, the answer segment of the liability question, etc. At the same time, the bank will establish a record table to record the detection status of each video segment. Among them, the multiple video segments formed by cutting the answer video data for the first time are all "undetected", and after being detected as target video segments, their status will be marked as "detected". Each time of detection selects the first "undetected" video segment in the multiple video segments as the target video segment in time sequence, updates its status to "detected" after detection, and all video segments are detected.
[0054] On the one hand, by only focusing on the undetected video segments, repeated processing of the detected segments can be avoided, thereby saving time and computing resources. Especially when processing a large number of insurance customer follow-up videos, the processing speed can be greatly improved. On the other hand, selecting the target segment from the undetected segments according to certain rules can ensure that all video segments are eventually detected and no segments that may exist in the behavior of answering outside the camera are missed, thereby improving the risk prevention and control capability.
[0055] S120, determining a target video segment currently needing to be detected from the multiple video segments.
[0056] Specifically, in the answer video data of the target customer to the target event, multiple video segments can be obtained by cutting the answer video data through a preset voice activity detection model. According to a preset selection rule, a video segment currently needing to be detected is taken as a target video segment from the obtained multiple video segments. For example, in the answer video data of target customer A in a second-hand housing loan face-to-face audit in a certain bank, multiple video segments are obtained by cutting the answer video data of target customer A through a preset voice activity detection model. The multiple video segments contain multiple non-voice video segments and multiple voice-containing video segments. At this time, the video segments corresponding to the multiple voice-containing video segments are arranged in time sequence from front to back in order of time, and the video segment arranged in the first position is taken as the target video segment currently needing to be detected.
[0057] S130. Use speech recognition technology to obtain the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content.
[0058] Automatic Speech Recognition (ASR) technology refers to the technology of converting the lexical content of human speech into computer-readable text. It is widely used in fields such as intelligent voice assistants and voice input.
[0059] Specifically, audio data from the target video segment is extracted using video editing software or specialized audio extraction tools, and then preprocessed (e.g., format conversion or noise reduction). The extracted audio data is then uploaded to a speech recognition engine (e.g., Google Cloud Speech-to-Text or iFlytek engine). By calling the speech recognition engine, it automatically converts the speech content in the audio data into speech-to-text content and outputs it. Furthermore, along with the output speech-to-text content, it also provides the start and end timestamps for each word in the speech-to-text content.
[0060] For example, during a video interview for a mortgage loan, bank staff (the creditor) will ask the debtor several questions, such as "What is your current monthly income?" and "Do you have any other liabilities?" The target video segment where the debtor answers these questions is processed using the method described above. First, audio data is extracted from the target video segment. After preprocessing such as format conversion and noise reduction, the iFlytek engine interface is called. The speech-text content recognized by the iFlytek engine might be, "My current monthly income is approximately 15,000 yuan, and I have no other liabilities." Simultaneously, each word in the speech-text content output by the iFlytek engine has a timestamp, such as "I" at 0.5-0.6 seconds, "eye" at 0.6-0.7 seconds, etc. Bank staff (the creditor) can use these timestamps, combined with subsequent video data such as the debtor's mouth movements, to more accurately determine the authenticity and credibility of the debtor's answers.
[0061] Specifically, in some embodiments, prior to step S140, the method includes the following steps:
[0062] Based on a preset word vector model, the speech text content is converted into a first word vector, and the preset speech template is converted into a second word vector;
[0063] Calculate the cosine angle between the first word vector and the second word vector to obtain the cosine angle value;
[0064] If the cosine angle value is greater than or equal to a preset angle threshold, then the voice text content is determined to match the preset speech template.
[0065] If the cosine angle value is less than the preset angle threshold, it is determined that the voice text content does not match the preset dialogue template.
[0066] wherein the preset word vector model refers to a model constructed based on machine learning or deep learning technology, which can map words or phrases in natural language into word vectors. The word vector not only contains semantic information of the word, but also reflects semantic relationship between words, such as semantic similarity, semantic correlation, etc. The preset dialogue template refers to a series of texts set in advance for verifying voice text content, which is used for matching comparison with the input voice text content, and usually contains a series of key sentences, word combinations, and their possible order or logical relationship, etc.
[0067] Specifically, the obtained voice text content and the preset dialogue template are preprocessed, including removing punctuation, converting to lowercase, removing stop words, etc., to improve the accuracy and efficiency of subsequent word vector calculation. For example, the stop word removal is performed using the function in the Natural Language Toolkit (NLTK) library. For the preprocessed voice text content, it can be first split into single words or phrases by a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model, and then each word or phrase is converted into a corresponding word vector based on a locally pre-stored and trained word vector model (e.g., Word2Vec model or GloVe model). Finally, the aforementioned word vectors can be combined into a first word vector representing the entire voice text content by weighted averaging. In the same way, the preset dialogue template is preprocessed, tokenized and word vector converted to obtain a second word vector.
[0068] After obtaining the first word vector and the second word vector, the first word vector and the second word vector are substituted into the cosine similarity formula to calculate the cosine angle between them, i.e., the cosine angle value. The cosine similarity formula is as follows:
[0069]
[0070] wherein, represents the first word vector; represents the second word vector; represents the vector dot product; represents the modulus of the first vector; represents a module of the second vector. The cosine angle value is calculated by substituting the first vector and the second vector into the cosine similarity formula, and the cosine angle value ranges between [-1, 1]. The closer the cosine angle value is to 1, the more similar the directions of the two vectors are, that is, the higher the matching degree of the semantic content of the voice text and the preset dialogue template is.
[0071] The cosine angle value calculated by the cosine similarity formula is compared with the preset angle threshold value. If the cosine angle value is greater than or equal to the preset angle threshold value, it indicates that the similarity between the voice text content and the preset dialogue template is high, that is, the voice text content matches the preset dialogue template. If the cosine angle value is less than the preset angle threshold value, it indicates that the similarity between the voice text content and the preset dialogue template is low, that is, the voice text content does not match the preset dialogue template.
[0072] For example, there is an A customer who conducts a video interview for a car consumption loan in B bank. Assuming that the preset dialogue template is "The loan amount you apply for is 200,000 yuan, the loan period is 3 years, the annual interest rate is 5.8%, and the equal principal and interest repayment method is adopted. Do you confirm the above information?". The voice text content corresponding to the question in the video segment answered by the A customer is "The loan I apply for is 200,000 yuan, 3 years, 5.8% annual interest rate, equal principal and interest repayment. Confirm no problem." Then, based on the preprocessing steps of removing punctuation, converting to lowercase, and removing stop words, the pre-trained BERT model is used to split it into word groups ["loan", "200,000", "3 years", "5.8% annual interest rate", "equal principal and interest", "repayment", "confirm", "no problem"], and then based on the local pre-stored and trained word vector model (for example, Word2Vec model or GloVe model), each word or word group is converted into a corresponding word vector. For example: loan [0.12, -0.45,..., 0.78]; 200,000 [0.34, 0.21,..., -0.56]; 3 years [-0.09, 0.67,..., 0.12]... Assuming that all word weights are the same, the first word vector is calculated by weighted average. Similarly, the second word vector is calculated by the same steps. Finally, the first word vector and the second word vector are substituted into the aforementioned cosine similarity formula to calculate the cosine angle between them, and the cosine angle value is obtained. Assuming that the finally calculated cosine angle value is 0.92, and the preset angle threshold value is 0.8, since 0.92 is greater than 0.8, the similarity between the voice text content answered by the A customer and the preset dialogue template is high, that is, the voice text content of the A customer matches the preset dialogue template.
[0073] S140, if the speech text content matches the preset dialogue template, performing mouth movement detection on the target video segment according to the timestamp and generating a mouth movement detection result.
[0074] Specifically, after obtaining the speech text content corresponding to the target video segment and the timestamp of each word in the speech text content through speech recognition technology, the speech text content is matched with the preset dialogue template. If the matching is successful, mouth movement detection is performed on the target video segment according to the timestamp and a mouth movement detection result is generated. If the matching is unsuccessful, the video segment corresponding to each word in the speech text content is obtained according to the timestamp, and the video segment is deleted or discarded.
[0075] Specifically, in some embodiments, as shown in FIG. 1, S140 includes the following steps: Figure 3
[0076] S141, obtaining a plurality of target image frames from the target video segment according to the timestamp.
[0077] S142, determining a mouth movement probability corresponding to each target image frame based on a preset mouth movement model.
[0078] S143, determining a mouth movement confidence of the target video segment according to each mouth movement probability.
[0079] S144, when the mouth movement confidence is greater than a preset confidence threshold, mouth movement is taken as the mouth movement detection result.
[0080] S145, when the mouth movement confidence is less than or equal to the confidence threshold, no mouth movement is taken as the mouth movement detection result.
[0081] Specifically, the target video segment is opened by a video processing library (for example, OpenCV) function to establish a connection with the video file. The timestamp is stored in a list, the frame rate of the target video segment is obtained by a video processing library, and the timestamp is converted into a corresponding video frame number. The corresponding target image frame is extracted from the target video segment according to the video frame number. By traversing the video frame number corresponding to the timestamp, the function of the video processing library is used to read and save these frames, and a plurality of target image frames are obtained.
[0082] A mouth movement model (e.g., a model trained based on a lightweight network structure such as MobileNetV3) is trained by the deep neural network. The lightweight network model (e.g., MobileNetV3) is selected to predict the mouth movement probability, which reduces the computational load and memory occupation while ensuring a certain accuracy, so that the method can be relatively efficiently run on different hardware devices. The mouth region image in each target image frame is preprocessed (e.g., normalized to scale the pixel value to the range of 0-1), and then the preprocessed target image frame is input into the locally stored and pre-trained mouth movement model for prediction. Using the pre-trained mouth movement model can effectively capture the mouth movement features and improve the accuracy of predicting the mouth movement probability.
[0083] The mouth movement model outputs a probability value for each target image frame (i.e., the mouth movement probability corresponding to each target image frame), which represents the possibility of the mouth being in a movement state in the target image frame and has a value range of 0-1. The mouth movement confidence of the target video segment can be obtained by weighting and averaging the probability values. The confidence is calculated based on the probability, which comprehensively considers the information of multiple target image frames, further improving the reliability of the detection result. The obtained mouth movement confidence is compared with the preset confidence threshold, and the mouth movement detection result is obtained by comparison.
[0084] For example, in a loan application process, the bank will ask the applicant to record a video to confirm the authenticity of the identity and the compliance of the operation, in which the applicant reads the loan contract terms.
[0085] The bank system extracts the image frames corresponding to the reading of each term from the video of the applicant reading the loan contract terms according to the timestamp information in the video content. For example, when the applicant reads the sentence "I promise to repay the loan principal and interest on time and in full", the system accurately extracts multiple image frames in the reading process of this sentence according to the timestamp. For each extracted image frame, after preprocessing, it is input into the pre-trained mouth movement model to obtain the mouth movement probability in each image frame. The mouth movement confidence of the video corresponding to the applicant reading the sentence "I promise to repay the loan principal and interest on time and in full" can be obtained by weighting and averaging the mouth movement probability in each image frame, which is assumed to be 0.1. Assuming that the bank's preset confidence threshold is 0.4, since the obtained mouth movement confidence is 0.1, which is less than the bank's preset confidence threshold, the bank can consider that the applicant's mouth is not moving, or there may be a third party answering behavior. The bank can take measures to further verify the situation to ensure the authenticity and compliance of the loan application.
[0086] S150, if the mouth movement detection result is mouth movement, the voiceprint data corresponding to the speech text content is stored in the voiceprint library corresponding to the target event.
[0087] Specifically, for example, in a face-to-face video of a housing loan application, the staff of the bank will ask the applicant a plurality of questions according to the situation of the applicant, and record the video data of the answers of the applicant at the same time. The bank system will detect the mouth movement of the applicant according to the video data of the answers of the applicant, and when the detection result is mouth movement, it indicates that the applicant answers the questions of the staff of the bank himself / herself, that is, there is no third person answering the questions.
[0088] According to the time stamp of each word in the speech text content, the system extracts the audio segment of the applicant answering the questions from the audio corresponding to the face-to-face video of the applicant, and after preprocessing such as noise reduction and normalization, the voiceprint data is extracted by using the voiceprint recognition technology. Since the business scenario belongs to the target event of housing loan application, the system stores the voiceprint data in the voiceprint library corresponding to the housing loan application.
[0089] The voiceprint data extracted in the subsequent question link of the face-to-face interview can be compared and verified with the voiceprint data stored in the current voiceprint library to confirm the identity of the applicant in real time, effectively preventing the risk of identity fraud in the loan business.
[0090] S160, if the mouth movement detection result is mouth movement, the voiceprint data corresponding to the speech text content is stored in the voiceprint library corresponding to the target event, and the target voiceprint data of the target video segment is detected according to the historical voiceprint data, and the third person answering behavior detection result is obtained.
[0091] Specifically, in the answer video data of the target customer to the target event, a plurality of answer data of the target customer are recorded, and the server detects the mouth movement behavior of the target customer when answering the questions in real time. At the same time, the voiceprint data corresponding to the speech text content with the detection result of mouth movement is stored in the voiceprint library corresponding to the target event. When the detection result of the target customer answering a question is mouth movement, the historical voiceprint data corresponding to the target event in the voiceprint library is obtained, and the target voiceprint data of the target video segment is detected according to the historical voiceprint data, so as to obtain the third person answering behavior detection result. If the mouth does not move but there is voice output, there may be an abnormal situation of third person answering. Through subsequent detection based on historical voiceprint data, potential risks can be captured in time, such as in the scenarios of loan, important contract signing, etc., to ensure the authenticity and security of the business and prevent fraud.
[0092] For example, in the face-to-face video interview with the bank for applying for a car loan, the staff of the bank will ask Xiaohong a series of questions to confirm her identity according to the process. Assuming that among the 5 questions asked by the staff of the bank, the mouth movement detection result of Xiaohong's answers to the first 4 questions is mouth movement, and the mouth movement detection result of Xiaohong's answer to the 5th question is no mouth movement. Then the voiceprint data corresponding to the voice text content of Xiaohong's answers to the first 4 questions is stored in the voiceprint library corresponding to the car loan, and the voiceprint data of the first 4 questions corresponding to the car loan is obtained from the voiceprint library, and the voiceprint data of the 5th question is detected according to the voiceprint data of the first 4 questions, so as to obtain the detection result of whether Xiaohong's answer to the 5th question is a behavior of answering by an outsider.
[0093] Specifically, in some embodiments, the step of detecting the target voiceprint data of the target video segment according to the historical voiceprint data to obtain a detection result of the behavior of answering by an outsider includes the following steps:
[0094] Obtaining the number of voiceprints of the historical voiceprint data;
[0095] If the number of voiceprints is greater than a preset number of voiceprint threshold, detecting the target voiceprint data according to the historical voiceprint data to obtain the detection result of the behavior of answering by an outsider.
[0096] Specifically, the historical voiceprint data stored in the voiceprint library is obtained, and the number of voiceprints of the historical voiceprint data is compared with a preset number of voiceprint threshold. If the number of voiceprints is greater than the preset number of voiceprint threshold, the target voiceprint data is detected according to the historical voiceprint data to obtain the detection result of the behavior of answering by an outsider. Rich data samples can more comprehensively reflect the voiceprint feature variation range of the target object, effectively reduce the misjudgment caused by individual differences or environmental factors of voiceprint features, and make the detection result more accurate and reliable. If the number of voiceprints is less than the preset number of voiceprint threshold, the target voiceprint data is not detected. When the historical voiceprint data is less, no detection is performed, so as to avoid the waste of invalid computing resources caused by insufficient data.
[0097] For example, in the process of applying for a large commercial loan, the bank requires the applicant to answer a series of key questions about the purpose of the loan, the repayment plan, etc. in the video interview. After detecting that the applicant's mouth does not move but has a voice answer (the voiceprint data of the video segment of the applicant's mouth not moving but having a voice answer is the target voiceprint data), the bank's system obtains the voiceprint library of the large commercial loan, and compares the number of voiceprints recorded in the voiceprint library with a preset number of voiceprint threshold. Assuming that the preset number of voiceprint threshold is 3 and the number of voiceprints recorded in the voiceprint library is 4, since 4 is greater than 3, the target voiceprint data can be detected according to the historical voiceprint data, so as to obtain the detection result of the behavior of answering by an outsider.
[0098] Further, in some embodiments, the detecting the target voiceprint data according to the historical voiceprint data to obtain the off-screen answering behavior detection result comprises the following steps:
[0099] obtaining a voiceprint similarity value between each historical voiceprint data and the target voiceprint data;
[0100] if there is historical voiceprint data with a voiceprint similarity value less than a preset voiceprint similarity threshold, then the off-screen answering behavior exists as the off-screen answering behavior detection result.
[0101] Specifically, for the audio of the target video segment, a professional voiceprint recognition tool (for example, Kaldi voiceprint recognition tool package) is used to extract the feature vector of the target voiceprint data. Meanwhile, the feature vector extraction is also performed on each piece of historical voiceprint data obtained. A voiceprint similarity calculation algorithm (for example, a cosine similarity algorithm) is used to calculate the feature vector of the target voiceprint data and the feature vector of each piece of historical voiceprint data one by one, so as to obtain the voiceprint similarity value between each historical voiceprint data and the target voiceprint data.
[0102] A voiceprint similarity threshold is preset in the system, for example, 0.7. The calculated voiceprint similarity value is compared with 0.7 one by one, as long as there is historical voiceprint data with a voiceprint similarity value less than the threshold, it is determined that the off-screen answering behavior exists, and this is taken as the off-screen answering behavior detection result; if all voiceprint similarity values are greater than or equal to the threshold, it is determined that the person answers, that is, the off-screen answering behavior does not exist as the off-screen answering behavior detection result. In the scene involving important business, such as loan business, this strict detection method can effectively prevent fraud, ensure that the person participates in the business process, effectively guarantee the authenticity and security of the business, and maintain the legitimate rights and interests of financial institutions and users.
[0103] In summary, in the embodiment of the present application, the target video segment currently to be detected is determined from the multiple video segments of the target customer answering the video data of the target event, the corresponding voice text content and the time stamp of each character are obtained through voice recognition technology, if the voice text content matches the preset script template, the mouth movement detection result of the target video segment is generated according to the time stamp, if the detection result is mouth movement, the voiceprint data corresponding to the voice text content is stored in the voiceprint library corresponding to the target event, if the detection result is no mouth movement, the historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and the target voiceprint data of the target video segment is detected according to the historical voiceprint data to obtain the off-screen answering behavior detection result, thereby improving the accuracy of off-screen answering behavior detection.
[0104] Figure 4 is a schematic block diagram of the answering behavior detection device provided by the embodiment of the present application. As shown in Figure 4 corresponding to the above answering behavior detection method, the present application also provides an answering behavior detection device 1000. The answering behavior detection device 1000 comprises units for executing the above answering behavior detection method. Specifically, please refer to Figure 4 , the answering behavior detection device 1000 comprises a cutting unit 1001, a determining unit 1002, an obtaining unit 1003, a first detection unit 1004, a storage unit 1005 and a second detection unit 1006, wherein:
[0105] The cutting unit 1001 is configured to obtain answer video data of a target customer for a target event, and cut the answer video data into a plurality of video clips;
[0106] The determining unit 1002 is configured to determine a target video clip currently in need of detection from the plurality of video clips;
[0107] The obtaining unit 1003 is configured to obtain voice text content corresponding to the target video clip and a time stamp of each word in the voice text content by using a voice recognition technology;
[0108] The first detection unit 1004 is configured to, if the voice text content matches a preset dialogue template, perform mouth movement detection on the target video clip according to the time stamp and generate a mouth movement detection result;
[0109] The storage unit 1005 is configured to, if the mouth movement detection result is mouth movement, store voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event;
[0110] The second detection unit 1006 is configured to, if the mouth movement detection result is no mouth movement, obtain historical voiceprint data corresponding to the target event from the voiceprint library, and detect target voiceprint data of the target video clip according to the historical voiceprint data to obtain an off-screen person answering behavior detection result.
[0111] In some embodiments, the cutting unit 1001 is specifically configured to:
[0112] cut the answer video data by using a preset voice activity detection model to obtain the plurality of video clips.
[0113] In some embodiments, the first detection unit 1004 is specifically configured to:
[0114] obtain a plurality of target image frames from the target video clip according to the time stamp;
[0115] determine a mouth movement probability corresponding to each of the target image frames based on the preset mouth movement model;
[0116] determine a mouth movement confidence of the target video segment according to the mouth movement probabilities;
[0117] when the mouth movement confidence is greater than a preset confidence threshold, determine mouth movement as the mouth movement detection result;
[0118] when the mouth movement confidence is less than or equal to the confidence threshold, determine no mouth movement as the mouth movement detection result.
[0119] In some embodiments, the second detection unit 1006 is specifically configured to:
[0120] obtain a number of voiceprints of the historical voiceprint data;
[0121] if the number of voiceprints is greater than a preset voiceprint number threshold, detect the target voiceprint data according to the historical voiceprint data to obtain the off-screen answer behavior detection result.
[0122] Further, the second detection unit 1006 is further specifically configured to:
[0123] obtain a voiceprint similarity value between each of the historical voiceprint data and the target voiceprint data;
[0124] if there is historical voiceprint data with a voiceprint similarity value less than a preset voiceprint similarity threshold, determine that there is off-screen answer behavior as the off-screen answer behavior detection result.
[0125] In some embodiments, before the determination unit 1002, the answer behavior detection apparatus 1000 further comprises:
[0126] a third detection unit configured to detect whether there is a video segment that has not been detected for off-screen answer behavior in the plurality of video segments;
[0127] At this time, the determination unit 1002 is specifically configured to:
[0128] if there is a video segment that has not been detected for off-screen answer behavior in the plurality of video segments, determine a video segment that has not been detected for off-screen answer behavior as the target video segment.
[0129] In some embodiments, before the first detection unit 1004, the answer behavior detection apparatus 1000 further comprises:
[0130] The conversion unit is configured to convert the voice text content into a first word vector and convert the preset script template into a second word vector based on a preset word vector model;
[0131] The calculation unit is configured to calculate a cosine included angle between the first word vector and the second word vector to obtain a cosine included angle value.
[0132] The first determination sub-unit is configured to determine that the voice text content matches the preset script template if the cosine included angle value is less than or equal to a preset included angle threshold.
[0133] The second determination sub-unit is configured to determine that the voice text content does not match the preset script template if the cosine included angle value is greater than the preset included angle threshold.
[0134] In summary, the answer behavior detection device 1000 in the embodiment of the present application determines the target video segment to be detected from the multiple video segments of the video data of the target customer answering the target event, acquires the corresponding voice text content and the time stamp of each word through voice recognition technology, determines that the voice text content matches the preset script template, generates a detection result through mouth movement detection of the target video segment according to the time stamp, stores the voiceprint data corresponding to the voice text content to the voiceprint library corresponding to the target event if the detection result is mouth movement, acquires the historical voiceprint data corresponding to the target event from the voiceprint library if the detection result is no mouth movement, and detects the target voiceprint data of the target video segment to obtain the off-screen person answer behavior detection result, thereby improving the accuracy of the off-screen person answer behavior detection.
[0135] The above-mentioned answer behavior detection device can be realized in the form of a computer program, which can run on a computer device as shown in Figure 5 .
[0136] Please refer to Figure 5 , Figure 5 is a schematic block diagram of a computer device provided by the embodiment of the present application. The computer device 800 is a terminal or a server.
[0137] Please refer to Figure 5 , the computer device 800 includes a processor 802, a memory and a network interface 805 connected through a system bus 801, wherein the memory can include a non-volatile storage medium 803 and an internal memory 804.
[0138] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions, which, when executed, can cause the processor 802 to perform an answer behavior detection method.
[0139] The processor 802 is configured to provide computing and control capabilities to support the operation of the entire computer device 800.
[0140] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803, which, when executed by the processor 802, causes the processor 802 to perform an answer behavior detection method.
[0141] The network interface 805 is configured to communicate with other devices via a network. Those skilled in the art can understand that the network interface 805 can be implemented in various forms, such as a network adapter, a modem, a Bluetooth module, a wireless transceiver, or the like. Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 800 to which the scheme of the present application is applied. The specific computer device 800 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0142] The processor 802 is configured to run the computer program 8032 stored in the memory to implement the following steps:
[0143] Obtain answer video data of a target customer for a target event, and cut the answer video data into a plurality of video clips;
[0144] Determine a target video clip currently needed to be detected from the plurality of video clips;
[0145] Obtain voice text content corresponding to the target video clip and a time stamp of each word in the voice text content by using a voice recognition technology;
[0146] If the voice text content matches a preset script template, perform mouth movement detection on the target video clip according to the time stamp and generate a mouth movement detection result;
[0147] If the mouth movement detection result is mouth movement, store voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event;
[0148] If the mouth movement detection result is no mouth movement, obtain historical voiceprint data corresponding to the target event from the voiceprint library, and detect target voiceprint data of the target video clip according to the historical voiceprint data to obtain an off-screen answer behavior detection result.
[0149] It should be understood that, in the embodiments of the present application, the processor 802 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0150] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by instructing relevant hardware by a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.
[0151] Therefore, the present application also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by a processor to perform the following steps:
[0152] Obtaining answer video data of a target customer for a target event, and cutting the answer video data into a plurality of video clips;
[0153] Determining a target video clip currently needed to be detected from the plurality of video clips;
[0154] Obtaining voice text content corresponding to the target video clip and a time stamp of each word in the voice text content by using a voice recognition technology;
[0155] If the voice text content matches a preset dialogue template, performing mouth movement detection on the target video clip according to the time stamp and generating a mouth movement detection result;
[0156] If the mouth movement detection result is mouth movement, storing voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event;
[0157] If the mouth movement detection result is that the mouth is not moving, historical voiceprint data corresponding to the target event is obtained from the voiceprint library, and target voiceprint data of the target video segment is detected according to the historical voiceprint data, to obtain an off-screen answering behavior detection result.
[0158] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, and various computer-readable storage media that can store program codes.
[0159] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0160] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed.
[0161] The steps in the method embodiments of the present application can be adjusted, combined, and reduced in sequence according to actual needs. The units in the device embodiments of the present application can be combined, divided, and reduced according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0162] The integrated unit, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the parts that make contributions to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application.
[0163] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for detecting an answering behavior, characterized by, The method comprises the following steps: acquiring answer video data of a target customer for a target event, and cutting the answer video data into a plurality of video clips; determining a target video clip currently in need of detection from the plurality of video clips; acquiring voice text content corresponding to the target video clip and a time stamp of each word in the voice text content by using a voice recognition technology; if the voice text content matches a preset script template, performing mouth movement detection on the target video clip according to the time stamp and generating a mouth movement detection result; if the mouth movement detection result is mouth movement, storing voiceprint data corresponding to the voice text content into a voiceprint library corresponding to the target event; if the mouth movement detection result is no mouth movement, acquiring historical voiceprint data corresponding to the target event from the voiceprint library, and detecting target voiceprint data of the target video clip according to the historical voiceprint data to obtain an off-screen person answering behavior detection result; the step of detecting the target voiceprint data according to the historical voiceprint data to obtain the off-screen person answering behavior detection result comprises the following steps: acquiring a voiceprint quantity of the historical voiceprint data; if the voiceprint quantity is greater than a preset voiceprint quantity threshold, detecting the target voiceprint data according to the historical voiceprint data to obtain the off-screen person answering behavior detection result; if the voiceprint quantity is less than the preset voiceprint quantity threshold, not detecting the target voiceprint data.
2. The method of claim 1, wherein the method further comprises: the step of performing mouth movement detection on the target video clip according to the time stamp and generating a mouth movement detection result comprises the following steps: acquiring a plurality of target image frames from the target video clip according to the time stamp; determining a mouth movement probability corresponding to each target image frame based on a preset mouth movement model; determining a mouth movement confidence of the target video clip according to each mouth movement probability; when the mouth movement confidence is greater than a preset confidence threshold, taking mouth movement as the mouth movement detection result; when the mouth movement confidence is less than or equal to the confidence threshold, taking no mouth movement as the mouth movement detection result.
3. The method of claim 1, wherein the method further comprises: the step of detecting the target voiceprint data according to the historical voiceprint data to obtain the off-screen person answering behavior detection result comprises the following steps: acquiring a voiceprint similarity value between each historical voiceprint data and the target voiceprint data; if there is historical voiceprint data with a voiceprint similarity value less than a preset voiceprint similarity threshold, taking the existence of off-screen person answering behavior as the off-screen person answering behavior detection result.
4. The method of claim 1, wherein the method further comprises: the step of cutting the answer video data into a plurality of video clips comprises the following step: cutting the answer video data by using a preset voice activity detection model to obtain a plurality of video clips.
5. The method of claim 1, wherein the method further comprises: before the step of determining a target video clip currently in need of detection from the plurality of video clips, the method comprises the following step: detecting whether there is a video clip that has not been subjected to off-screen person answering behavior detection in the plurality of video clips; the step of determining a target video clip currently in need of detection from the plurality of video clips comprises the following steps: If there is a video clip that has not been detected for the off-screen answering behavior among the plurality of video clips, a video clip is determined as the target video clip from the video clip that has not been detected for the off-screen answering behavior.
6. The method of claim 1, wherein the method further comprises: Before the mouth movement detection and the generation of the mouth movement detection result of the target video clip according to the time stamp, if the voice text content matches the preset dialogue template, the method comprises the following steps of: Converting the voice text content into a first word vector and the preset dialogue template into a second word vector based on a preset word vector model; Calculating a cosine angle between the first word vector and the second word vector to obtain a cosine angle value; If the cosine angle value is greater than or equal to a preset angle threshold, it is determined that the voice text content matches the preset dialogue template; If the cosine angle value is less than the preset angle threshold, it is determined that the voice text content does not match the preset dialogue template.
7. An answering behavior detection apparatus characterized by comprising: Comprise: The cutting unit is used for obtaining the answering video data of the target customer for the target event, and cutting the answering video data into a plurality of video clips; The determination unit is used for determining the target video clip that needs to be detected from the plurality of video clips; The acquisition unit is used for acquiring the voice text content corresponding to the target video clip and the time stamp of each word in the voice text content by using the voice recognition technology; The first detection unit is used for performing the mouth movement detection on the target video clip according to the time stamp and generating the mouth movement detection result if the voice text content matches the preset dialogue template; The storage unit is used for storing the voiceprint data corresponding to the voice text content into the voiceprint library corresponding to the target event if the mouth movement detection result is the mouth movement; The second detection unit is used for acquiring the historical voiceprint data corresponding to the target event from the voiceprint library if the mouth movement detection result is the non-movement of the mouth, and detecting the target voiceprint data of the target video clip according to the historical voiceprint data to obtain the off-screen answering behavior detection result; The second detection unit is specifically used for acquiring the number of voiceprints of the historical voiceprint data; if the number of voiceprints is greater than a preset number of voiceprint threshold, the target voiceprint data is detected according to the historical voiceprint data to obtain the off-screen answering behavior detection result; if the number of voiceprints is less than the preset number of voiceprint threshold, the target voiceprint data is not detected.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the answering behavior detection method in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to realize the answering behavior detection method in any one of claims 1 to 6.
Citation Information
Patent Citations
Cosmetic live broadcast efficacy verbal skill monitoring method and system
CN114168711A
Double-recording method based on speaker detection and related equipment
CN117649691A
Sound and lip synchronous identification method, device and equipment and storage medium
CN118334554A
Over-person answering-on-behalf detection method based on face-to-face video and related equipment
CN119206566A