Illegal behavior identification method and device, equipment and storage medium
By dividing face-to-face review videos and building preset reply templates and voiceprint analysis in the face-to-face review, the loopholes in identifying violations during online reviews were solved, and the accurate identification and risk control of violations were achieved.
Patent Information
- Application Number
- CN202510696851.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-08
AI Technical Summary
There are loopholes in the identification method of identifying violations in the mid-line review of existing technologies, especially in the field of medical and health insurance, it is difficult to accurately distinguish between customers' answers and others' answers, resulting in an increase in incorrect underwriting decisions and financial risks of insurance companies.
By obtaining face-to-face review videos, extracting user response content, using voice activity detection and voice recognition technology to divide face-to-face review videos, building preset reply templates, matching user response content, determining standard audio intervals, and analyzing voiceprint characteristics to determine whether the voiceprint belongs to the same character.
Effectively identifying whether there are any violations in the face-to-face review video improves the accuracy and efficiency of the face-to-face review and reduces the financial risks of the insurance company.
Smart Images

Figure CN120452457A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology and can be applied to the medical health field and the financial field. In particular, it relates to a method, device, equipment and storage medium for identifying illegal behaviors. Background Art
[0002] In various fields, such as healthcare and insurance, the development of the internet has led to the adoption of online audits in many face-to-face interview scenarios. This shift has significantly improved customer convenience. For example, when applying for medical expense reimbursement insurance, traditional in-person audits required customers to visit the insurance company's office, a time-consuming and laborious process. However, with online audits, customers can complete the process without leaving their homes, significantly saving time and effort while also significantly improving audit efficiency.
[0003] However, online interviews, lacking the oversight of in-person interviews, present significant security risks and offer some clients the opportunity to cheat. In healthcare insurance scenarios, common interview violations include guidance and proxy answers. For example, when a client applies for critical illness insurance, if they are supervised by someone else during the interview, they may be led to conceal their medical history or exaggerate the severity of their current condition. Once a claim is settled, these unscrupulous clients may resort to using guidance or proxy answers as an excuse to cheat, severely infringing on the insurance company's interests.
[0004] Currently, a common method for detecting violations is to monitor whether a customer's mouth moves when speaking. If the mouth moves, the customer is deemed to be answering on their own; if the mouth remains motionless, it is suspected that someone else is answering on their behalf. However, this method has vulnerabilities. For example, when applying for long-term care insurance, the customer may lip-sync while someone else is actually answering. In this case, the system may mistakenly interpret this as a normal response, resulting in the violation not being identified. This can lead to incorrect underwriting decisions and increase the financial risk for the insurance company. Summary of the Invention
[0005] The present invention provides a method, device, equipment and storage medium for identifying violations to solve the technical problem that the identification method of violations in face-to-face audits in related technologies has identification loopholes.
[0006] In a first aspect, the present invention provides a method for identifying illegal behavior, comprising:
[0007] Obtain the target user's face-to-face interview video;
[0008] Obtaining the target user's responses to the interview questions in the interview video;
[0009] Determining, from the user reply content, standard user reply content that matches a preset reply template;
[0010] Obtaining a standard audio interval corresponding to the standard user response content in the face-to-face review video;
[0011] Analyzing the voiceprint of each of the standard audio intervals;
[0012] Determine whether the voiceprints of each standard audio interval belong to the voiceprint of the same person. If so, it is determined that there is no violation in the face-to-face review video of the target user. Otherwise, it is determined that there is a violation in the face-to-face review video of the target user.
[0013] In a second aspect, the present invention provides a device for identifying illegal behavior, comprising:
[0014] Video acquisition module, used to obtain the target user's face-to-face interview video;
[0015] A response content acquisition module is used to obtain the target user's response content to the face-to-face review questions in the face-to-face review video;
[0016] a determination module, configured to determine, from the user reply content, standard user reply content that matches a preset reply template;
[0017] An audio interval acquisition module, configured to acquire a standard audio interval in the face-to-face review video corresponding to the standard user response content;
[0018] An analysis module, configured to analyze the voiceprint of each of the standard audio intervals;
[0019] The determination module is used to determine whether the voiceprints of each standard audio interval belong to the voiceprint of the same person. If so, it is determined that there is no violation in the face-to-face review video of the target user; otherwise, it is determined that there is a violation in the face-to-face review video of the target user.
[0020] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for identifying illegal behaviors when executing the computer program.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for identifying illegal behaviors are implemented.
[0022] In the solution implemented by the above-mentioned method, device, equipment and storage medium for identifying violations, the target user's answer audio to each interview question in the interview video (the standard audio interval corresponding to the standard user's answer content) is separated, and the voiceprint of the target user's answer audio to each interview question is further analyzed to see whether it belongs to the same person. If it belongs to the same person, it can be understood that the interview questions in the interview video are all answered by the same person. This can avoid other people answering the interview questions in the interview video on behalf of the target user, thereby effectively identifying the target user's interview violation. In summary, the above-mentioned solution can solve the technical problem that the existing technology has identification loopholes in the identification method of interview violations. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] Figure 1 1 is a flow chart of a method for identifying illegal behavior in one embodiment of the present invention;
[0025] Figure 2 yes Figure 1 A flow chart of a specific implementation of step S120;
[0026] Figure 3 is another flow chart of a method for identifying illegal behavior in one embodiment of the present invention;
[0027] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step S130;
[0028] Figure 5 yes Figure 4 A flowchart of a specific implementation of step S1322;
[0029] Figure 6 yes Figure 1 Another specific implementation flow diagram of step S130;
[0030] Figure 7 This is another flowchart of a method for identifying illegal behavior in one embodiment of the present invention;
[0031] Figure 8 1 is a schematic structural diagram of a device for identifying illegal behavior in one embodiment of the present invention;
[0032] Figure 9is a structural diagram of a computer device in one embodiment of the present invention;
[0033] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0035] Figure 1 A flowchart of a method for identifying illegal behavior provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method for identifying illegal behavior provided by the embodiment of the present invention includes the following steps.
[0036] Step S110: Obtain the face-to-face interview video of the target user.
[0037] Specifically, the face-to-face interview video of the target user obtained in step S110 can be a face-to-face interview video of any field, and the acquisition method can be any feasible method, which is not limited in detail here.
[0038] As a specific example, in the healthcare sector, when a customer (i.e., target user) applies for healthcare insurance-related services, such as critical illness insurance or medical insurance claims, the insurance company needs to conduct an in-person interview with the customer. To this end, the insurance company can use an online platform with video recording capabilities to conduct the interview. After the interview is completed, the system automatically saves the interview video and can upload it to the insurance company's secure server, allowing the insurance company to obtain the target user's interview video.
[0039] Step S120: Obtain the target user's response content to the interview questions in the interview video.
[0040] Specifically, obtaining the target user's responses to the interview questions in the interview video can be done in any feasible manner. For example, speech recognition software can be used to process the received interview video, extract the audio portion, and convert it into text. The generated text can then be segmented. This can be done in any feasible manner. For example, since the interview process includes questions from the interviewer and the target user's responses, an algorithm can be used to distinguish between the two. For example, this can be determined based on tone, pauses, specific question guides (such as "Excuse me" and "Can you explain?"), and the logical connection between the answers. For example, if the interviewer asks, "What was the condition that led to your most recent hospitalization?" the next continuous text is likely the user's response. In this way, the entire interview conversation can be segmented into question-answer pairs.
[0041] In some embodiments of the present invention, Figure 2 As shown, step S120 includes the following steps.
[0042] Step S121: input the face-to-face interview video into a voice activity detection model to obtain the target user's speaking time interval in the face-to-face interview video;
[0043] Step S122: dividing the face-to-face review video into sub-face-to-face review videos according to each of the speaking time intervals, wherein the sub-face-to-face review videos correspond one-to-one to the speaking time intervals;
[0044] Step S123: Perform voice recognition on the sub-interview video to obtain text content in the sub-interview video, and determine the text content in the sub-interview video as the user reply content.
[0045] Specifically, in step S121, the core function of the Voice Activity Detection (VAD) model is to accurately identify speaking time periods within an audio stream. When the interview video is fed into the VAD model, the model pre-processes the audio portion of the video. This includes removing background noise and adjusting the audio gain to improve audio quality and facilitate more accurate voice activity detection.
[0046] Furthermore, the model can determine whether a speech signal exists at each time point by analyzing the characteristics of the audio signal, such as short-time energy and zero-crossing rate. When the characteristics of the audio signal meet the preset conditions for the presence of speech, the model can mark this time period as speaking time. By analyzing the audio frame by frame, the complete speaking time interval of the target user in the face-to-face interview video is ultimately determined. For example, in a medical health insurance face-to-face interview video, even if there is some hospital background noise in the surrounding environment, such as broadcast sounds and footsteps, the voice activity detection model can eliminate these noise interferences through accurate analysis of the audio features and accurately identify the time period when the customer is speaking.
[0047] Regarding step S122, the face-to-face interview video segmentation process can be based on video editing technology. According to each speaking time interval, corresponding video segments are extracted from the face-to-face interview video. These video segments are called sub-face-to-face interview videos. Each sub-face-to-face interview video corresponds one-to-one with a specific speaking time interval. For example, if the target user has three speaking time intervals during the face-to-face interview process, the original face-to-face interview video will be divided into three sub-face-to-face interview videos, each corresponding to a speaking time interval.
[0048] As a specific example, suppose a customer applies for a critical illness insurance claim. During the interview, the interviewer sequentially asks the customer about their medical condition, treatment process, and expenses. In step S121, the voice activity detection model analyzes the entire interview video and identifies the time intervals during which the customer speaks when answering the interviewer's questions about their medical condition, such as from the 2nd to 4th minute of the video, and from the 6th to 8th minute when answering questions about their treatment process. Proceeding to step S122, the original interview video is divided into multiple sub-interview videos based on these speaking time intervals. For example, the first sub-interview video contains content from the 2nd to 4th minute, showing only the customer's responses to their medical condition during this period; the second sub-interview video contains content from the 6th to 8th minute, describing the customer's treatment process. In step S123, speech recognition is performed on each of these sub-interview videos. In the first sub-interview video, the client described how she was diagnosed with stomach cancer at the end of last year, discovered through a gastroscopy at XX Hospital. The speech recognition system converted this information into text. In the second sub-interview video, the client described her treatment process, including surgery and chemotherapy, for a total of six courses. This information was also accurately converted into text. The text content recognized from each sub-interview video constituted the client's complete responses to the interview questions.
[0049] In this way, through the solution of steps S121-S123, the voice activity detection model first determines the speaking time interval, and then divides the interview video into sub-interview videos for speech recognition. This avoids performing speech recognition on the entire interview video, significantly reducing computational complexity. For example, in a long interview video, if the interviewer spends a long time asking questions and the client speaks relatively little, this solution only requires speech recognition for the speaking portion, significantly improving processing speed and saving time and cost. After dividing the interview video by speaking time interval, the content of each sub-interview video is relatively simple, focusing primarily on the client's answers to one or several related questions. This allows the speech recognition system to reduce contextual interference during processing and focus on specific speech content, thereby improving speech recognition accuracy. For example, when a client answers complex questions about medical expenses, the sub-interview video only contains this portion of content, allowing the speech recognition system to more accurately identify key information such as the amount and item. Each sub-interview video corresponds to a speaking time interval, which in turn corresponds to the client's response to a specific question or category of questions. This one-to-one correspondence makes subsequent analysis of the user's responses more convenient and organized. For example, when reviewing a claim application, an insurance company can review the customer's medical condition, treatment process, cost rationality, etc. based on the content of different sub-interview videos, thereby improving the accuracy and efficiency of the review.
[0050] Step S130: determining standard user reply content that matches a preset reply template in the user reply content.
[0051] Specifically, pre-set response templates correspond to pre-set interview questions and can be customized based on the specific interview questions set during the application. For example, a series of pre-set response templates can be constructed based on the characteristics of various health insurance products, common questions, and relevant regulations and policies. These templates can cover various types of questions that may arise during the interview process. For example, for medical insurance claims, the templates can include standard answer formats for information such as the reason for medical treatment, the time and location of treatment, the treatment method, and detailed expense information. For critical illness insurance claims, the templates can include standard representations for information such as the time of diagnosis, the hospital where the illness was diagnosed, descriptions of symptoms, and medical history. When constructing templates, it is also important to fully consider the differences in representation across regions and healthcare systems, ensuring that the templates are as universal and comprehensive as possible. Templates can take the form of structured text frames, which can include both mandatory and optional fields. For example, a template for detailed medical expense information might specify mandatory fields such as expense item name, amount, and payment method, while optional fields might include whether medical insurance reimbursement is available and the reimbursement percentage. Templates can also include keywords and semantic rules for subsequent matching.
[0052] Specifically, when determining the standard user reply content that matches the preset reply template, any feasible method can be used, for example, a text matching algorithm in natural language processing can be used to match the user reply content obtained in step S120 with the preset reply template. The text matching algorithm may include a matching algorithm based on word vectors, such as a cosine similarity algorithm. Specifically, the user reply content and the text in the preset reply template can be converted into word vector representations, and the degree of their matching is measured by calculating the cosine similarity between the word vectors of the two. The higher the similarity, the closer the user reply content is to the preset reply template. Semantic analysis algorithms, such as dependency syntactic analysis and semantic role labeling, can also be combined. By analyzing the sentence structure and semantic relationship between words in the user reply content and the preset reply template, the matching situation can be judged more accurately. For example, if the preset reply template requires the description of disease symptoms, the chronological order of the symptoms must be clarified, then the semantic analysis algorithm can determine whether the user reply content meets this requirement.
[0053] According to the results of the matching algorithm, the user response content that best matches the preset response model is determined and determined as the standard user response content.
[0054] Furthermore, after determining the standard user reply content, the user reply content can be verified and supplemented. If the user reply content contains missing information, the user can be prompted to supplement the relevant information according to the template requirements of the standard user reply content.
[0055] In some embodiments of the present invention, Figure 3 As shown, the steps of setting the preset reply template include the following steps.
[0056] Step S1311, determining the target field to which the face-to-face review video belongs;
[0057] Step S1312: Obtain regulatory documents for the target field, historical response templates that meet the requirements of the regulatory documents, and a professional dictionary for the target field;
[0058] Step S1313: Input the regulatory document, the historical reply template, and the professional dictionary into a preset model to obtain a template syntax tree corresponding to the preset reply template with semantic constraints.
[0059] The nodes of the template syntax tree represent the syntax elements of the target domain, and the semantic constraints are used to represent the context of the syntax elements represented by the nodes in the target domain.
[0060] Specifically, for step S1311, when constructing a preset reply template, the target field to which the interview video belongs must be clearly defined. This is because the interview content and requirements in different fields vary greatly. For example, the interview focus and compliance requirements in the financial field, education field, and medical and health insurance field are completely different. Determining the target field will help to accurately collect relevant information and build a template that conforms to the characteristics of the field. In actual operation, the field to which it belongs can be determined by the subject of the interview video, the type of business involved, or related identification information. For example, if the interview video revolves around the customer's application for medical expense reimbursement, consultation on major disease insurance terms, etc., it can be clearly stated that its target field is the medical and health insurance field.
[0061] Regarding step S1312, since each target area is subject to corresponding laws, regulations, and regulatory policies, each target area has corresponding regulatory documents. For the healthcare insurance sector, regulatory documents may include insurance industry standards, medical insurance claims regulations, and customer rights protection laws. These documents can define the legal and compliance boundaries of insurance business operations and serve as an important basis for constructing pre-set response templates. For example, information such as medical insurance reimbursement coverage and time limits for claim applications can be reflected in the templates to ensure that the insurance company's face-to-face review and subsequent business processing meet regulatory requirements. These regulatory documents can be obtained through channels such as the official websites of relevant departments, documents published by industry associations, and relevant legal and regulatory databases. Furthermore, historical response templates that meet the requirements of regulatory documents are proven in past face-to-face reviews and comply with regulatory requirements. They contain standard responses and formats for various common questions. For example, in healthcare insurance, historical response templates for describing the disease diagnosis process and treatment options can provide a reference framework for constructing new templates. These historical response templates can be collected by reviewing internal company business documents, case libraries, and past successful claims cases. Specifically, the professional dictionary in the target field can include terms, abbreviations and professional expressions that are unique to the target field. In the field of medical and health insurance, professional dictionaries can cover medical terms (such as the names of various diseases and surgical procedures) as well as insurance industry terms (such as deductibles, insured amounts, and claim ratios). These professional terms are crucial for accurately understanding and answering interview questions. For example, when describing disease symptoms and treatment processes, using professional medical terms can convey information more accurately. Professional dictionaries can be compiled through medical professional books, insurance industry standard terminology manuals, and authoritative online medical and insurance knowledge bases.
[0062] Regarding step S1313, the preset model can be any model that, after inputting regulatory documents, historical response templates, and a professional dictionary, can output a template syntax tree corresponding to the preset response template with semantic constraints. For example, the preset model can be a grammatical analysis model based on natural language processing technology, such as a dependency parsing model or a semantic role labeling model. The collected regulatory documents, historical response templates, and professional dictionary are input into this preset model. The model will conduct an in-depth analysis of this data to extract the grammatical structure and semantic information.
[0063] Furthermore, the preset model can generate a template syntax tree corresponding to a preset response template with semantic constraints by analyzing the input data. The template syntax tree is a tree structure in which each node represents a grammatical element in the target domain. For example, in the field of medical health insurance, a node may represent grammatical elements such as "disease name", "medical treatment time", and "medical expense amount". The connection relationship between nodes reflects the logical relationship between these grammatical elements in the sentence structure. Semantic constraints are used to clarify the specific context of the grammatical element represented by each node in the target domain. For example, the semantic constraint of the "disease name" node may stipulate that it must be a disease name that exists in a medical professional dictionary and is within the scope of medical insurance reimbursement. In this way, through the template syntax tree, not only the grammatical structure of the preset response template is specified, but also the semantic scope of each grammatical element is constrained, making the constructed template more accurate and in line with actual business needs.
[0064] As can be understood, by incorporating regulatory documents into the template-building process based on steps S1311-S1313, the pre-set response templates are ensured to fully comply with industry regulations and regulatory requirements. This helps insurance companies avoid legal risks and regulatory penalties arising from non-compliant operations, maintaining their image as legal and compliant. Furthermore, by utilizing historical response templates, successful past business practices can be passed down. Simultaneously, historical templates can be optimized through pre-set models in response to new regulatory requirements and business development, enabling them to continuously adapt to changing market environments and business needs. Furthermore, the use of professional dictionaries ensures that the terminology used in the pre-set response templates is accurate and standardized, meeting the professional requirements of the healthcare insurance sector. This helps improve the accuracy of face-to-face communication, avoids misunderstandings caused by inappropriate terminology, and enhances the insurance company's professional image and service quality. Furthermore, the semantic constraints of the template syntax tree make the pre-set response templates more semantically sound and accurate. Clarifying the semantic scope of each grammatical element allows for more targeted and rational communication between the interviewer and the client, improving the efficiency and quality of the interview and providing a reliable information foundation for subsequent business processes such as claims review.
[0065] In some embodiments of the present invention, Figure 4As shown, step S130 includes the following steps.
[0066] Step S1321, calculate the semantic similarity between the preset response template and the alternative user response content in the user response content;
[0067] Step S1322, calculate the structural similarity between the preset response template and the alternative user response content;
[0068] Step S1324, weight the semantic similarity and the structural similarity to obtain the similarity between each alternative user response content and the preset response template;
[0069] Step S1325, determine the alternative user response content with the maximum similarity value as the standard user response content.
[0070] Specifically, for step S1321, the alternative user response content is the content in the user response content. Each alternative user response content can be a complete semantic representation. For example, it can be a sentence or a phrase. Additionally, in step S1321, when calculating the semantic similarity between the preset response template and the alternative user response content, a word vector-based method can be used. For example, the word vector-based method can be the cosine similarity word vector algorithm. First, the preset response template and the alternative user response content can be preprocessed, including removing stop words (such as meaningless function words like "de", "le", and "zai"), converting the text to lowercase, and performing stemming (such as converting "running" to "run") and other operations to simplify the text and improve the accuracy of subsequent calculations. Further, using a pre-trained word vector model, the preprocessed text is converted into a word vector representation. These models can map each word to a low-dimensional vector space through learning a large amount of text data, making the distance between semantically similar words in the vector space relatively close as well. For example, in the field of medical health insurance, for the two words "heart disease" and "cardiovascular disease", after being processed by the word vector model, the angle between their word vectors in the space will be relatively small, indicating that they have a high semantic similarity.
[0071] For step S1322, the method of calculating the structural similarity between the preset reply template and the alternative user reply content can be any feasible method. For example, when calculating the structural similarity, the preset reply template and the alternative user reply content can be subjected to text structure analysis. For example, a dependency syntax analysis algorithm can be used to perform text structure analysis. The algorithm can analyze the grammatical dependency relationship between each word in a sentence and determine the subject, predicate, object, attributive, adverbial, complement and other structural components of the sentence. For example, for the sentence "I was hospitalized for three days in XX hospital because of a cold", the dependency syntax analysis can identify that "I" is the subject, "hospitalized" is the predicate, "in XX hospital" is the location adverbial, "because of a cold" is the cause adverbial, and "three days" is the complement.
[0072] Furthermore, key structural features can be extracted from the analyzed text structure. These features may include sentence length, the number and distribution of words of different parts of speech in a sentence, the depth of the sentence hierarchy, etc. For example, sentence length can reflect the level of detail of the text; the distribution of words of different parts of speech can reflect the focus of the text's expression. For example, a text with more nouns may focus on describing things, while a text with more verbs may focus on describing actions. In the field of medical and health insurance, for texts describing the claims process, structural features may include the number of steps mentioned, the level of detail described in each step, etc. Furthermore, by comparing the structural features of the preset reply template and the alternative user reply content, the structural similarity between the preset reply template and the alternative user reply content can be obtained.
[0073] Regarding step S1324, the weights of semantic similarity and structural similarity can be set based on the specific needs of the application. For example, in the healthcare insurance field, if the accuracy and professionalism of the user's response are more important, a higher weight, such as 0.7, may be assigned to semantic similarity; while if the logical order and completeness of the response are more important, a relatively lower but still important weight, such as 0.3, may be assigned to structural similarity.
[0074] Thus, based on the determined weights, the semantic similarity calculated in step S1321 and the structural similarity calculated in step S1322 can be weighted and summed. Assuming the semantic similarity is S1 and the structural similarity is S2, with weights w1 and w2, respectively, the final similarity S = w1*S1+w2*S2. For example, if the semantic similarity S1 is 0.8, the weight w1 is 0.7, the structural similarity S2 is 0.7, and the weight w2 is 0.3, the final similarity S = 0.7*0.8+0.3*0.7 = 0.77.
[0075] Regarding step S1325, after calculating the similarity between each candidate user reply and the preset reply template, a series of similarity values are obtained, and these similarity values are compared to find the largest similarity value.
[0076] It is understood that, based on steps S1321-S1325, both semantic similarity and structural similarity can be considered simultaneously, enabling evaluation of candidate user responses from multiple perspectives. Semantic similarity ensures the consistency of the user response's meaning with the preset template, while structural similarity ensures the logical and expressive rationality of the response. This comprehensive evaluation approach is more comprehensive and accurate than relying solely on semantics or structure, enabling a more precise determination of whether a user response meets standards. Furthermore, in the healthcare insurance sector, users express themselves in a variety of ways, and focusing solely on semantics can overlook important structural information, such as information integrity and logical order. For example, when describing a claims process, even if semantics are similar, incorrect step sequence can lead to misunderstandings. By combining structural similarity, it is possible to better adapt to such complex and changing business scenarios and improve the accuracy of user response evaluations. Furthermore, accurately determining standard user response content helps insurance companies improve efficiency and accuracy in business processes such as reviewing claims applications and answering customer inquiries. Reviewers can quickly reference standard user responses to determine whether the information provided is complete and accurate, thereby reducing review time and error rates, improving customer service quality and company operational efficiency. Furthermore, by analyzing the similarity between large numbers of user responses and pre-set response templates, insurance companies can identify common user expression patterns and questions, thereby optimizing and improving pre-set response templates, further improving the rationality and efficiency of business processes.
[0077] In some embodiments of the present invention, Figure 5 As shown, step S1322 includes the following steps.
[0078] Step S13221: constructing a first abstract dependency tree corresponding to the candidate user reply content and a second abstract dependency tree corresponding to the preset reply template;
[0079] Step S13222: Calculate the minimum number of operations required for mutual conversion between the first abstract dependency tree and the second abstract dependency tree;
[0080] Step S13223: Determine the structural similarity between the first abstract dependency tree and the second abstract dependency tree based on the minimum number of operations, wherein the minimum number of operations is inversely proportional to the value of the structural similarity.
[0081] In some embodiments of the present invention, Figure 6 As shown, after step S1322 and before step S1324, the following steps are also included.
[0082] Step S13231, evaluating the semantics and structural importance of each candidate user's reply content in the target domain;
[0083] Step S13232 : determining the weights of the semantic similarity and the structural similarity according to the importance of the semantics and the structure in the target domain, respectively.
[0084] The weight value of the semantic similarity is proportional to the importance of the semantics in the target domain, and the weight value of the structural similarity is proportional to the importance of the structure in the target domain.
[0085] Specifically, for step S13221, the method of constructing the first abstract dependency tree and the second abstract dependency tree can be a dependency syntactic analysis method. Dependency syntactic analysis can reveal the dependency relationship between each word in the sentence, determine the core dominant word of each word and the dependency type between them (such as subject-predicate relationship, verb-object relationship, etc.). For example, for the sentence "The patient received treatment in the hospital", after dependency syntactic analysis, "accept" is the core verb, "patient" is its subject, "treatment" is its object, and "in the hospital" is an adverbial indicating location, and has a specific dependency relationship with "accept". After completing the dependency syntactic analysis, the analysis results are presented in a tree structure to form an abstract dependency tree. Among them, the nodes of the tree represent the words in the sentence, and the edges represent the dependency relationships between words. In order to focus more on structural features, the dependency tree can be abstracted, and some specific part-of-speech details may be ignored, retaining only key dependency relationships and core word information. For the abstract dependency tree constructed for the above sentence, the root node may be "accept", under which there is a "patient" node representing the subject, a "treatment" node representing the object, and a "hospital" node representing the location adverbial, which are connected by edges to show the dependency relationship.
[0086] In this way, the above dependency parsing and abstract dependency tree construction are performed on the candidate user reply content and the preset reply template, respectively, to obtain a first abstract dependency tree and a second abstract dependency tree. These two trees represent the sentence structure of the candidate user reply content and the preset reply template, respectively.
[0087] Regarding step S13222, when calculating the minimum number of operations required to convert between two abstract dependency trees, these operations may include node insertion, deletion, and replacement. For example, if a node in the first tree is missing but exists in the second tree and has a significant impact on the structure, an insertion operation may be required. If a node has different dependencies in the two trees, a node replacement operation may be required.
[0088] The method for calculating the minimum number of operations required to convert the first abstract dependency tree and the second abstract dependency tree can be any feasible method, for example, some algorithms based on graph theory or dynamic programming can be used to calculate the minimum number of operations. For example, through a dynamic programming algorithm, starting from the root node of the tree, the corresponding nodes and their subtrees of the two trees are gradually compared, and the minimum number of operations required to convert one tree into the other is recorded. Assuming that the structure of a subtree in the first abstract dependency tree is significantly different from the corresponding subtree structure in the second abstract dependency tree, multiple node insertion, deletion, and replacement operations are required to make them consistent. The algorithm will accurately calculate the minimum number of these operations.
[0089] Regarding step S13223, since the minimum number of operations reflects the degree of difference between the two tree structures, the fewer the number of operations, the more similar the structures of the two trees are. Therefore, the minimum number of operations is inversely proportional to the structural similarity.
[0090] As a specific example, suppose the pre-set response template is "The insured must provide a medical diagnosis certificate issued by a reputable hospital to apply for a medical insurance claim." After dependency parsing and abstraction processing, the second abstract dependency tree is constructed with "provide" as the root node. Below it is the "insured" node representing the subject, the "disease diagnosis certificate" node, the "diagnosis certificate" node modifying "disease diagnosis certificate" and the "issued by a reputable hospital" node, and the "to apply for a medical insurance claim" node representing the purpose. Suppose the device selects the user response content as "I want to provide a hospital diagnosis certificate to apply for a medical insurance claim." Thus, the first abstract dependency tree is constructed with "provide" as the root node, with the "I" node representing the subject, the "diagnosis certificate" node, the "hospital" node modifying "diagnosis certificate," and the "to apply for a medical insurance claim" node representing the purpose. Further comparison of the two abstract dependency trees reveals that the "medical certificate" in the first abstract dependency tree lacks the modifier "formal," requiring a node insertion. Furthermore, while "I" and "insured" share similar semantics, they can be considered different representations in the template's canonical expression, requiring a node replacement. Therefore, the minimum number of operations required to transform the first abstract dependency tree into the second is calculated to be two.
[0091] It can be understood that by constructing an abstract dependency tree and calculating the minimum number of operations, the structural differences between candidate user responses and the preset response template can be accurately quantified. This quantification method makes the assessment of structural similarity more objective and accurate, avoiding the ambiguity of subjective judgment. In the healthcare insurance business, a large number of different user responses can be compared for structural similarity using a unified quantitative standard, helping to select responses that best match the preset template structure. Furthermore, the dependency tree construction process emphasizes the dependencies between words in a sentence, highlighting key features of sentence structure. In healthcare insurance scenarios, accurate structural information is crucial for understanding user needs and ensuring the standardization of responses. For example, in the description of claim application requirements, the logical relationships and structural order of each information point are key. This method can accurately capture these key features and determine whether the user response fully and accurately conveys the relevant information. Furthermore, users may use a variety of vocabulary and sentence structures in their expressions, but the core structure of a sentence tends to be relatively stable. Structural similarity calculation based on the abstract dependency tree can better adapt to these complex expression variations and accurately determine the structural consistency of user responses with the preset template. Even if the words are different, as long as the structure is similar, the degree of matching with the template can be identified, which improves the inclusiveness and adaptability to various forms of expression and helps to more comprehensively process the user's response content.
[0092] Step S140: Obtain a standard audio interval in the face-to-face review video that corresponds to the standard user response content.
[0093] Specifically, within the interview video, based on the previously determined standard user responses, time stamping or other video analysis techniques can be used to locate the audio sections corresponding to these standard user responses, thereby obtaining the standard audio interval. For example, if the standard user response is from 3 minutes 20 seconds to 4 minutes 10 seconds in the video, the audio of this time period corresponds to the standard audio interval. This step determines the specific audio range for subsequent voiceprint analysis.
[0094] As a specific example, consider a standard user response during a health insurance interview, where a customer describes their symptoms and treatment process when applying for a critical illness insurance claim. The interviewer asks, "Please describe your symptoms and treatment after becoming ill in detail." The standard user response is, "After becoming ill, I often felt fatigued and had a cough. I later underwent surgery in the hospital, followed by several chemotherapy sessions." Video analysis determined that the audio segment corresponding to this response is from the 5th minute 30 seconds to the 6th minute 40 seconds of the video, which is the standard audio segment.
[0095] Step S150: Analyze the voiceprint of each standard audio interval.
[0096] Specifically, a voiceprint is a biometric feature derived by extracting features from a speaker's speech signal. Each person's voiceprint is unique, just like a fingerprint. When analyzing voiceprints within standard audio intervals, the audio within these intervals can be preprocessed, including noise removal and volume adjustment, to improve audio quality. Signal processing and machine learning algorithms can then be used to extract characteristic parameters from the audio, such as Mel-frequency cepstral coefficients. These characteristic parameters reflect the speaker's voice characteristics and are used for subsequent voiceprint recognition and comparison.
[0097] As a specific example, the audio within the aforementioned standard audio interval is processed. First, interference such as hospital background noise is removed. Then, voiceprint analysis software is used to extract Mel-frequency cepstral coefficient features. The audio signal is then framed, and a set of Mel-frequency cepstral coefficient feature parameters is extracted for each frame. Ultimately, a series of data representing the speaker's voiceprint characteristics within the standard audio interval is obtained.
[0098] Step S160, determining whether the voiceprints of each of the standard audio intervals belong to the voiceprint of the same person, if so, determining that there is no violation in the face-to-face review video of the target user, otherwise determining that there is a violation in the face-to-face review video of the target user.
[0099] Specifically, the voiceprint features extracted from each standard audio interval are compared with each other. If the voiceprint feature matching degree of all standard audio intervals reaches a certain threshold, it indicates that these voices come from the same person, that is, it is determined that there is no violation in the face-to-face review video of the target user. On the contrary, if the voiceprint feature matching degree is lower than the threshold, it means that there may be voices of different people, that is, it is determined that there is a violation in the face-to-face review video of the target user. For example, the voiceprint matching degree threshold is set to 80%. When the voiceprint matching degree of each standard audio interval is higher than 80%, it is determined that there is no violation; if the voiceprint matching degree of some intervals is lower than 80%, it is determined that there is a violation.
[0100] As a specific example, assume that there are multiple standard audio segments during an interview, each corresponding to a different response to a question. The voiceprint features extracted from these standard audio segments are compared. If the calculated similarity of these voiceprint features in the voiceprint database is below a set threshold of 75% (hypothetically), the interview video of the target user is judged to have violated the rules. For example, in one case, the voiceprint of an audio segment describing the medical condition and another segment explaining the cost of treatment had a similarity of only 60%, which was below the threshold, thus determining a violation.
[0101] Thus, in an embodiment of the present invention, the target user's answer audio to each interview question in the interview video (the standard audio interval corresponding to the standard user's answer content) is separated, and the voiceprint of the target user's answer audio to each interview question is further analyzed to see whether it belongs to the same person. If it belongs to the same person, it can be understood that the interview questions in the interview video are all answered by the same person. This can avoid other people answering the interview questions in the interview video when the target user answers them, thereby effectively identifying the target user's interview violation behavior. In summary, the above scheme can solve the technical problem that the identification method of interview violation behavior in the existing technology has identification loopholes.
[0102] In some embodiments of the present invention, Figure 7 As shown, after step S160, the following steps are also included.
[0103] Step S170, after determining that the target user's face-to-face review video has violated the rules, determine whether the target user's lips in the face-to-face review video are moving. If so, it is determined that the violation has been misjudged; otherwise, it is determined that the violation has not been misjudged.
[0104] Specifically, after determining that a violation has occurred in the target user's interview video, further observation can be made to see if the target user's lips move during the interview. If the lips move while speaking, the voiceprint analysis may have misjudged the user, as the lip movement and voice normally appear to be from the same person. If the lips do not move, but the voiceprint analysis indicates a different voice, then it is more likely that a violation has occurred, indicating that the violation has not been misjudged.
[0105] As a specific example, after determining a violation, the face-to-face review video is reviewed again. If the target user's lips are noticeably moving while speaking (corresponding to the audio segment where the violation occurred), such as during the audio segment describing a surgical procedure, the violation may have been misjudged. Conversely, if the target user's lips remain motionless during the audio playback, such as during the audio segment mentioning a treatment fee, the violation is not misjudged.
[0106] Understandably, while voiceprint analysis is a relatively reliable authentication technology, it can be affected by various factors, such as ambient noise, differences in recording equipment, and changes in the speaker's state, leading to misjudgments. By using a supplementary method to observe the target user's lip movements, the risk of misjudgment can be reduced to a certain extent. In the healthcare insurance sector, misjudgment can lead to unfair treatment for customers, affecting both the customer experience and the company's reputation. Adding this step helps improve the accuracy of the judgment. Furthermore, the synchronization of lip movement and voice provides an intuitive basis for determining whether the speaker is the same person. Combining voiceprint analysis and lip synchronization observation allows for two distinct perspectives in determining violations, making the results more reliable. When handling complex healthcare insurance audit scenarios, this multi-dimensional judgment approach can better address various potential scenarios, providing stronger support for companies to accurately identify violations and safeguarding the legitimate rights and interests of both the company and its customers.
[0107] It should be understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. The software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0108] In one embodiment, a device for identifying illegal behaviors is provided, which corresponds to the method for identifying illegal behaviors in the above embodiment. Figure 8 As shown, the generating device includes a video acquisition module 810, a reply content acquisition module 820, a determination module 830, an audio interval acquisition module 840, an analysis module 850, and a judgment module 860. The functional modules are described in detail as follows:
[0109] Video acquisition module 810, used to obtain the target user's face-to-face interview video;
[0110] The answer content acquisition module 820 is used to obtain the user answer content of the target user to the interview questions in the interview video;
[0111] A determination module 830 is configured to determine standard user reply content that matches a preset reply template;
[0112] An audio interval acquisition module 840 is configured to acquire, from the user response content, a standard audio interval corresponding to the standard user response content in the face-to-face review video;
[0113] An analysis module 850 is configured to analyze the voiceprint of each standard audio interval;
[0114] The determination module 860 is used to determine whether the voiceprints of each of the standard audio intervals belong to the voiceprints of the same person. If so, it is determined that there is no violation in the face-to-face review video of the target user; otherwise, it is determined that there is a violation in the face-to-face review video of the target user.
[0115] In one embodiment, the reply content acquisition module 820 is specifically configured to:
[0116] Inputting the face-to-face review video into a voice activity detection model to obtain a speaking time interval of the target user in the face-to-face review video;
[0117] Dividing the face-to-face review video into sub-face-to-face review videos according to each of the speaking time intervals, wherein the sub-face-to-face review videos correspond one-to-one to the speaking time intervals;
[0118] Perform voice recognition on the sub-interview video to obtain text content in the sub-interview video, and determine the text content in the sub-interview video as the user reply content.
[0119] In one embodiment, the determination module 860 is further configured to:
[0120] After determining that the target user's face-to-face review video has violated the rules, determine whether the target user's lips in the face-to-face review video are moving. If so, it is determined that the violation has been misjudged; otherwise, it is determined that the violation has not been misjudged.
[0121] In one embodiment, the determination module 830 is further configured to:
[0122] Determine the target area to which the interview video belongs;
[0123] Obtain regulatory documents in the target field, historical response templates that meet the requirements of the regulatory documents, and professional dictionaries in the target field;
[0124] The regulatory document, the historical reply template and the professional dictionary are input into a preset model to obtain a template syntax tree corresponding to the preset reply template with semantic constraints, wherein the nodes of the template syntax tree represent the various grammatical elements of the target domain, and the semantic constraints are used to represent the context of the grammatical elements represented by the nodes in the target domain.
[0125] In one embodiment, the determination module 830 is specifically configured to:
[0126] Calculating semantic similarity between the preset reply template and the candidate user reply content in the user reply content;
[0127] Calculating the structural similarity between the preset reply template and the candidate user reply content;
[0128] Weighting the semantic similarity and the structural similarity to obtain the similarity between the answer content of each candidate user and the preset answer template;
[0129] The candidate user reply content with the largest similarity value is determined as the standard user reply content.
[0130] In one embodiment, the determination module 830 is further configured to:
[0131] Evaluate the semantics and structural importance of each candidate user response in the target domain;
[0132] The weights of the semantic similarity and the structural similarity are determined according to the importance of the semantics and the structure in the target domain, respectively, wherein the weight value of the semantic similarity is proportional to the importance of the semantics in the target domain, and the weight value of the structural similarity is proportional to the importance of the structure in the target domain.
[0133] In one embodiment, the determination module 830 is further configured to:
[0134] Constructing a first abstract dependency tree corresponding to the candidate user reply content and a second abstract dependency tree corresponding to the preset reply template;
[0135] Calculating a minimum number of operations required for mutual conversion between the first abstract dependency tree and the second abstract dependency tree;
[0136] The structural similarity between the first abstract dependency tree and the second abstract dependency tree is determined according to the minimum number of operations, wherein the minimum number of operations is inversely proportional to the value of the structural similarity.
[0137] The present invention provides a device for identifying violations, which separates the target user's answer audio to each interview question in the interview video (the standard audio interval corresponding to the standard user's answer content), and further analyzes whether the voiceprint of the target user's answer audio to each interview question belongs to the same person. If they belong to the same person, it can be understood that the interview questions in the interview video are all answered by the same person. This can avoid other people answering the interview questions on the target user's behalf when the target user answers the interview questions in the interview video, thereby effectively identifying the target user's interview violations. In summary, the above scheme can solve the technical problem that the identification method of interview violations in the existing technology has identification loopholes.
[0138] The specific definition of the device for identifying illegal behavior can be found in the definition of the method for identifying illegal behavior above, and will not be repeated here. The various modules in the above-mentioned device for identifying illegal behavior can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0139] Based on the above-mentioned identification methods of violations, such as Figure 9 As shown, an embodiment of the present invention further provides a schematic structural diagram of a device for identifying illegal behaviors, which includes a processor 91 and a memory 92 coupled to the processor 91. The memory 92 stores a computer program, which, when executed by the processor 91, causes the processor 91 to perform the steps of the method for identifying illegal behaviors in the above embodiment.
[0140] For other details about how the processor 91 in the above-mentioned violation identification device implements the above-mentioned technical solution, please refer to the description of the violation identification method provided in the above-mentioned invention embodiment, which will not be repeated here.
[0141] Among them, the processor 91 can also be called a CPU (Central Processing Unit), and the processor 91 may be an integrated circuit chip with signal processing capabilities; the processor 91 can also be a general-purpose processor, DSP (Digital Signal Process), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, among which the general-purpose processor can be a microprocessor or the processor 91 can also be any conventional processor, etc.
[0142] like Figure 10As shown, an embodiment of the present invention further provides a schematic diagram of the structure of a computer-readable storage medium, on which a readable computer program 101 is stored; wherein, the computer program 101 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), and other media that can store program code, or a terminal device such as a computer, server, mobile phone, or tablet.
[0143] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0144] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.
[0145] In addition, the functional modules in the various embodiments of the present invention may be integrated into a single processing module, each module may exist physically separately, or two or more modules may be integrated into a single module. The integrated modules may be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may be stored in a computer-readable storage medium.
[0146] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0147] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium), or a semiconductor medium (e.g., an SSD (solid state disk)).
[0148] The technical solution provided by the present invention is introduced in detail above. Specific examples are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
[0149] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) containing computer-usable program code.
[0150] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0151] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0153] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for identifying illegal behavior, characterized in that: include: Obtain the target user's face-to-face interview video; Obtaining the target user's responses to the interview questions in the interview video; Determining, from the user reply content, standard user reply content that matches a preset reply template; Obtaining a standard audio interval corresponding to the standard user response content in the face-to-face review video; Analyzing the voiceprint of each of the standard audio intervals; Determine whether the voiceprints of each standard audio interval belong to the voiceprint of the same person. If so, determine that there is no violation in the face-to-face review video of the target user. Otherwise, determine that there is a violation in the face-to-face review video of the target user.
2. The method for identifying illegal behavior according to claim 1, characterized in that: The obtaining of the target user's response to the interview questions in the interview video includes: Inputting the face-to-face review video into a voice activity detection model to obtain a speaking time interval of the target user in the face-to-face review video; Dividing the face-to-face review video into sub-face-to-face review videos according to each of the speaking time intervals, wherein the sub-face-to-face review videos correspond one-to-one to the speaking time intervals; Perform voice recognition on the sub-interview video to obtain text content in the sub-interview video, and determine the text content in the sub-interview video as the user reply content.
3. The method for identifying illegal behavior according to claim 1, characterized in that: After determining whether the voiceprints of the standard audio intervals belong to the same person, if so, determining that the face-to-face review video of the target user does not violate any rules; otherwise, determining that the face-to-face review video of the target user does violate any rules, the method further includes: After determining that the target user's face-to-face review video has violated the rules, determine whether the target user's lips in the face-to-face review video are moving. If so, it is determined that the violation has been misjudged; otherwise, it is determined that the violation has not been misjudged.
4. The method for identifying illegal behavior according to claim 1, characterized in that: The steps for setting the preset reply template include: Determine the target area to which the interview video belongs; Obtain regulatory documents in the target field, historical response templates that meet the requirements of the regulatory documents, and professional dictionaries in the target field; The regulatory document, the historical reply template and the professional dictionary are input into a preset model to obtain a template syntax tree corresponding to the preset reply template with semantic constraints, wherein the nodes of the template syntax tree represent the various grammatical elements of the target domain, and the semantic constraints are used to represent the context of the grammatical elements represented by the nodes in the target domain.
5. The method for identifying illegal behavior according to claim 4, characterized in that: Determining, from the user reply content, standard user reply content that matches a preset reply template includes: Calculating semantic similarity between the preset reply template and the candidate user reply content in the user reply content; Calculating the structural similarity between the preset reply template and the candidate user reply content; Weighting the semantic similarity and the structural similarity to obtain the similarity between the answer content of each candidate user and the preset answer template; The candidate user reply content with the largest similarity value is determined as the standard user reply content.
6. The method for identifying illegal behavior according to claim 5, characterized in that: After calculating the structural similarity between the preset answer template and the candidate user answer content, and before weighting the semantic similarity and the structural similarity to obtain the similarity between each candidate user answer content and the preset answer template, the method further includes: Evaluate the semantics and structural importance of each candidate user response in the target domain; The weights of the semantic similarity and the structural similarity are determined according to the importance of the semantics and the structure in the target domain, respectively, wherein the weight value of the semantic similarity is proportional to the importance of the semantics in the target domain, and the weight value of the structural similarity is proportional to the importance of the structure in the target domain.
7. The method for identifying illegal behavior according to claim 5, characterized in that: The calculating of the structural similarity between the preset reply template and the candidate user reply content includes: Constructing a first abstract dependency tree corresponding to the candidate user reply content and a second abstract dependency tree corresponding to the preset reply template; Calculating a minimum number of operations required for mutual conversion between the first abstract dependency tree and the second abstract dependency tree; The structural similarity between the first abstract dependency tree and the second abstract dependency tree is determined according to the minimum number of operations, wherein the minimum number of operations is inversely proportional to the value of the structural similarity.
8. A device for identifying illegal behavior, characterized in that: include: Video acquisition module, used to obtain the target user's face-to-face interview video; A response content acquisition module is used to obtain the target user's response content to the face-to-face review questions in the face-to-face review video; a determination module, configured to determine, from the user reply content, standard user reply content that matches a preset reply template; An audio interval acquisition module, configured to acquire a standard audio interval in the face-to-face review video corresponding to the standard user response content; An analysis module, configured to analyze the voiceprint of each of the standard audio intervals; The determination module is used to determine whether the voiceprints of each standard audio interval belong to the voiceprint of the same person. If so, it is determined that there is no violation in the face-to-face review video of the target user; otherwise, it is determined that there is a violation in the face-to-face review video of the target user.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for identifying illegal behaviors according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for identifying illegal behaviors according to any one of claims 1 to 7 are implemented.