Character detection method and apparatus thereof

By acquiring biometric and textual information from multimedia data during dual-recording quality inspection, and utilizing voiceprint and lip movement recognition technologies to detect the role of the interactive object, the problem of inaccurate speaker identification in existing technologies is solved, achieving efficient role detection and preventing unauthorized operations.

CN116127432BActive Publication Date: 2026-04-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-02-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

During the dual-recording quality inspection process, existing technologies cannot effectively determine whether the speaker is the intended role, and there are cheating and violation operations such as substitute questioning and answering and inducement to answer. In addition, manual quality inspection is costly and has the risk of missed detection.

Method used

By acquiring biometric features and text information from multimedia data, feature vectors are extracted using voiceprint recognition and lip movement recognition technologies. Combined with role tag information, the system can detect whether the interactive object is the expected role.

Benefits of technology

It achieves accurate matching between the speaker and the question and answer content, clearly determines whether the role of the interaction object is the expected role, prevents violations such as substitute question and answer and inducement to answer, and reduces the cost of manual quality inspection and the risk of missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127432B_ABST
    Figure CN116127432B_ABST
Patent Text Reader

Abstract

This application discloses a role detection method and apparatus, relating to the field of artificial intelligence, particularly natural language processing, speech technology, and deep learning. The specific implementation involves: acquiring first multimedia data of the interactive object within the question-and-answer time interval of the current question-and-answer session; extracting biometric features from the first multimedia data and obtaining a first biometric vector of the interactive object based on the biometric features; acquiring text information from the first multimedia data and obtaining first role label information from the text information; and detecting whether the interactive object is the expected role within the question-and-answer time interval based on the first role label information and the first biometric vector. This method can achieve speaker role separation during dual-recording quality inspection, clearly and accurately detecting whether the role of the currently speaking interactive object is the expected role, thus addressing cheating and violation operations such as substitute question-and-answer and induced answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, particularly to technologies such as natural language processing, speech technology and deep learning, and specifically to a role detection method and apparatus that can be used in scenarios such as smart finance. Background Technology

[0002] Dual recording refers to the simultaneous audio and video recording during the sale of financial products (such as wealth management products and insurance products). The execution of dual recording plays a crucial role in comprehensively and objectively recording key sales stages, such as marketing presentations, relevant risk and key information prompts, and consumer confirmation and feedback, during the process of financial services or product purchases by both financial institutions and consumers.

[0003] In practice, the dual recording process has several noteworthy issues. For example, some sales personnel, finding the process cumbersome, or seeking to circumvent regulations and avoid risks, or aiming to exploit loopholes to close sales, fail to strictly adhere to the relevant regulations, creating potential risks for financial consumers. Furthermore, in stages requiring consumers to confirm their awareness of product risks and express their purchase intentions, sales personnel, to avoid duplicate recordings, may induce consumers to confirm or answer on their behalf. Controlling and identifying such violations through post-event manual review via full audio recording would be extremely costly and carries the risk of missed recordings.

[0004] In related technologies, artificial intelligence (AI) is used to perform quality control on the recording process. AI methods include face detection, face recognition, document recognition, speech synthesis, and speech recognition. However, the dual-recording quality control process in these technologies has loopholes, which can lead to cheating and other violations such as substitute answers and misleading responses. Summary of the Invention

[0005] This application provides a role detection method, apparatus, electronic device, and storage medium.

[0006] According to a first aspect of this application, a role detection method is provided, comprising:

[0007] Retrieve the first multimedia data of the currently interacting object;

[0008] Extract biometric features from the first multimedia data, and obtain the first biometric feature vector of the interactive object based on the biometric features in the first multimedia data;

[0009] Obtain the text information of the first multimedia data, and obtain the first role tag information of the first multimedia data based on the text information;

[0010] Based on the first role tag information and the first biometric vector, detect whether the interactive object is the expected role.

[0011] According to a second aspect of this application, a role detection device is provided, comprising:

[0012] The first acquisition module is used to acquire the first multimedia data of the current interactive object;

[0013] The feature extraction module is used to extract biological features from the first multimedia data;

[0014] The second acquisition module is used to acquire the first biometric vector of the interactive object based on the biometric features in the first multimedia data;

[0015] The third acquisition module is used to acquire the text information of the first multimedia data and acquire the first role tag information of the first multimedia data based on the text information;

[0016] The detection module is used to detect whether the interactive object is the expected role based on the first role tag information and the first biometric vector.

[0017] According to a third aspect of this application, an electronic device is provided, comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0021] According to a fourth aspect of this application, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0022] According to a fifth aspect of this application, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in the first aspect above.

[0023] The technology of this application solves the problem of speaker role separation in the dual recording quality inspection process, determines the combination of the speaker's biometrics and text, and forms a match between the speaker and the question and answer content. It can clearly and accurately detect whether the role of the current interactive object is the expected role, so as to deal with cheating and violation operations such as substitute question and answer and inducement to answer.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0025] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of this application. Wherein:

[0026] Figure 1 A flowchart illustrating a role detection method provided in this application embodiment;

[0027] Figure 2 This is a schematic diagram of the structure of the ECAPA-TDNN voiceprint recognition model provided in the embodiments of this application;

[0028] Figure 3 A flowchart illustrating another role detection method provided in this application embodiment;

[0029] Figure 4 A flowchart illustrating the method for obtaining a reference biometric vector and the mapping relationship between the reference biometric vector and reference role tag information provided in this application embodiment;

[0030] Figure 5 A block diagram of a role detection device provided in an embodiment of this application;

[0031] Figure 6 A block diagram of another role detection device provided in the embodiments of this application;

[0032] Figure 7 A block diagram of yet another role detection device provided in the embodiments of this application;

[0033] Figure 8 A block diagram of another role detection device provided in the embodiments of this application;

[0034] Figure 9 This is a block diagram of an electronic device used to implement the role detection method in the embodiments of this application. Detailed Implementation

[0035] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0036] In the description of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone.

[0037] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms "a" and "the" as used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0038] It should be understood that although the terms first, second, third, etc., may be used to describe various information in the embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" and "suppose" as used herein can be interpreted as "when," "when," or "in response to a determination."

[0039] First, it should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0040] It should also be noted that the technology in this application can be used in smart finance scenarios. Common steps in smart finance dual-recording quality inspection scenarios may include: opening remarks, customer presentation of valid identification, product instruction manual reading, product risk disclosure statement reading, text display and signature recognition, and closing remarks. Artificial intelligence capabilities commonly used in each step of smart finance dual-recording quality inspection scenarios may include text-to-speech (TTS) technology, automatic speech recognition (ASR) technology, facial recognition technology, ID card recognition technology, and handwriting recognition technology.

[0041] In dual recording, multiple stages require confirmation of the consumer's intent through voice responses (e.g., after product descriptions or risks are read, the consumer answers whether they understand). While speech recognition technology can be used to transcribe these responses, it cannot determine the speaker's identity or whether someone has answered on behalf of the consumer. Furthermore, during the recording process, sales personnel may act against the consumer's true intentions, inducing them to answer in a manner indicating they already understand. Speech recognition alone cannot identify the speaker of such responses; if it is a salesperson, it constitutes a violation. Therefore, it is necessary to perform question-and-answer role detection in the question-and-answer phase of the dual recording quality inspection scenario (including phases using speech recognition technology) to determine whether the interactive object (i.e., the speaker) within the question-and-answer period matches the expected role for that period. In other words, it is crucial to determine whether the roles taking turns answering questions (e.g., sales personnel and customers) during the dual recording quality inspection question-and-answer process conform to the expected correspondence.

[0042] The role detection method and apparatus of this application are described below with reference to the accompanying drawings. It should be noted that the role detection scheme in the embodiments of this application can be used for question-and-answer role detection in dual-recording quality inspection scenarios.

[0043] Figure 1 This is a flowchart illustrating a role detection method provided in an embodiment of this application. Figure 1 As shown, the method may include, but is not limited to, the following steps.

[0044] In step 101, the first multimedia data of the current interactive object is obtained.

[0045] Optionally, in one implementation, the role detection method in the embodiments of this application can be used in a dual-recording quality inspection scenario. The current interaction object can be the object that needs to speak within the question-and-answer time interval of the current question-and-answer session. When performing role detection on the object, the first multimedia data of the interaction object within the question-and-answer time interval of the current question-and-answer session can be obtained.

[0046] In the embodiments of this application, the interactive object within the question-and-answer time interval can refer to the speaker who speaks within the question-and-answer time interval.

[0047] In embodiments of this application, the first multimedia data may include voice data, and / or, the first multimedia data may include video data. In one possible implementation, the first multimedia data may include voice data. For example, in a question-and-answer session of a dual-recording quality inspection scenario, continuous voice segments can be extracted within the question-and-answer time interval as the voice data of the interactive object within that time interval. In another possible implementation, the first multimedia data may include video data; for example, in a question-and-answer session of a dual-recording quality inspection scenario, continuous video segments can be extracted within the question-and-answer time interval as the video data of the interactive object within that time interval.

[0048] For example, in multiple stages of a dual-recording quality inspection scenario, it is necessary to confirm the consumer's intentions through voice responses. This includes tasks such as reading product instructions, and having the consumer answer whether they understand the product risks. These all fall under the category of question-and-answer sessions. In these sessions, speech recognition technology is typically used to extract continuous audio segments within the question-and-answer timeframe as the primary multimedia data for the interaction.

[0049] In step 102, biometric features are extracted from the first multimedia data, and the first biometric vector of the interactive object is obtained based on the biometric features in the first multimedia data.

[0050] In one embodiment of this application, the first multimedia data includes voice data, and the biometric features of the first multimedia data may include acoustic features, and the first biometric feature vector may include a voiceprint feature vector.

[0051] In embodiments of this application, the first multimedia data includes video data, and the biometrics of the first multimedia data may include lip movement features, wherein the first biometric vector includes a lip movement feature vector.

[0052] In the embodiments of this application, the first multimedia data includes voice data and video data, and the first biometric vector may include a voiceprint feature vector and a lip movement feature vector.

[0053] In one possible implementation, taking the first multimedia data including speech data and the first biometric vector as a voiceprint feature vector as an example, acoustic features can be extracted from the speech data using voiceprint feature extraction technology. For example, these acoustic features can be filter FBank features and / or MFCC (Mel-Frequency Cepstral Coefficients) features. In one possible implementation, these acoustic features can be FBank features. Compared to MFCC features, the FBank features used in this application do not require a Discrete Cosine Transform (DCT) process, have lower computational cost, and better fit the characteristics of human ear reception. After extracting the FBank features from the speech data, the FBank features can be input into a preset voiceprint recognition model to obtain a voiceprint feature vector that can represent the speaker's role in the speech data.

[0054] It should be noted that the aforementioned voiceprint recognition model can be pre-trained. In one implementation, the ECAPA-TDNN voiceprint recognition model can be used. The input to this ECAPA-TDNN voiceprint recognition model can be an 80-dimensional FBank feature of length T. For example, as shown... Figure 2 The diagram shows the structure of the ECAPA-TDNN voiceprint recognition model. The first part is a Conv1D+ReLU+BN module, which, due to the presence of one-dimensional convolution, is equivalent to a TDNN (Time Delay Neural Network) module. The second part is an N-layer SE-Res2Net module (3 layers). The third part is also a TDNN module, which fuses the features output from the aforementioned multi-layer SE-Res2Net modules. The fourth part is an Attention Statistical Pooling layer, used to statistically pool the output. The fifth part is an FC+BN module, used to perform a linear transformation on the final features; the output dimension is a pre-defined voiceprint vector dimension, here set to 192. The final part is an AAM-Softmax layer, used to classify the output; the number of class labels corresponds to the number of speakers. This indicates that the voiceprint recognition model is obtained through supervised training. It is important to note that this application uses the AAM loss function, which aims to maximize the inter-class distance and minimize the intra-class distance. Compared to the Softmax layer, adding additive angular margins can widen the inter-class distance.

[0055] In the embodiments of this application, the ECAPA-TDNN voiceprint recognition model is trained using zhvoice Chinese speech data. The zhvoice dataset consists of eight open-source datasets, containing information on text, speech, and speaker. After noise reduction and silence removal, it contains approximately 3200 speakers, nearly 900 hours of audio, and approximately 1.13 million text entries, totaling 13 million characters. Furthermore, this application employs audio data enhancement methods such as random cropping, adding background noise, adjusting speech rate, and adjusting volume to enhance the robustness of the model. The Nesterov momentum optimization algorithm is used for model training. In one implementation, the output of the fifth part of the model, i.e., a 192-dimensional tensor, is taken as the voiceprint feature vector of the interactive object.

[0056] In another possible implementation, taking the first multimedia data as video data and the first biometric vector as a lip movement feature vector as an example, a pre-set lip movement recognition model can be used to process the video data to obtain the lip movement feature vector of the interactive object. This lip movement recognition model can be pre-trained, and existing technologies can be used to train the lip movement recognition model. This application does not specifically limit this, nor will it elaborate further.

[0057] In step 103, the text information of the first multimedia data is obtained, and the first role tag information of the first multimedia data is obtained based on the text information.

[0058] In one possible implementation, taking the first multimedia data including voice data as an example, speech recognition technology can be used to perform speech recognition on the voice data to obtain the text information of the voice data.

[0059] In another possible implementation, taking the first multimedia data including video data as an example, lip movement recognition technology can be used to perform lip movement recognition on the video data to obtain the lip movement information of the video data. By utilizing the mapping relationship between lip movement and text, the text information corresponding to the lip movement information can be obtained, and the text information is the text information of the video data.

[0060] In embodiments of this application, the first role tag information of the first multimedia data can be obtained by using keyword tagging. In one possible implementation, a first keyword can be extracted from the text information. Based on a pre-established mapping relationship between keywords and role tag information, the role tag information corresponding to the first keyword can be obtained, and this role tag information can be determined as the first role tag information of the first multimedia data. For example, a business person says, "Hello, I am the business manager of financial products, my name is Zhang San," during a Q&A session. The keyword "business manager" in this text information has corresponding role tag information, which can be determined as the first role tag information of the first multimedia data.

[0061] It is worth noting that other technical means can also be used to obtain the first role tag information of the first multimedia data. In one possible implementation, semantic role annotation technology is used to perform semantic role annotation on the text information to obtain the semantic role annotation results, and the first role tag information of the first multimedia data is obtained based on the semantic role annotation results.

[0062] In step 104, the interaction object is detected as the expected role based on the first role tag information and the first biometric vector.

[0063] Optionally, in one implementation, the interaction object can be detected as the expected question-and-answer role within the question-and-answer time interval based on the first role label information and the first biometric vector.

[0064] In the embodiments of this application, the first role tag information and the first biometric vector of the first multimedia data can be used to detect whether the interaction object is the expected role within the question-and-answer time interval.

[0065] In one possible implementation, reference role label information corresponding to the first biometric vector is determined, and the reference role label information is compared with the first role label information to determine whether the interaction object is the expected role within the question-and-answer time period.

[0066] In another possible implementation, a reference biometric vector corresponding to the first role label information is determined. This reference biometric vector is then matched with the first biometric vector to determine whether the interaction object is the expected role within the question-and-answer time period. For example, the similarity between the reference biometric vector and the first biometric vector is calculated. If the similarity is greater than or equal to a certain threshold, the interaction object can be determined to be the expected role within the question-and-answer time period. If the similarity is less than the threshold, the interaction object can be determined to be an unexpected role within the question-and-answer time period, i.e., not the expected role within the question-and-answer time period.

[0067] By implementing the embodiments of this application, multimedia data of the interactive object can be obtained. By using the text information and biometric vectors corresponding to the multimedia data, it is possible to detect whether the interactive object is the expected role. This solves the problem of separating the speaker role in the dual recording quality inspection process, and determines the combination of the speaker's biometrics and text to form a match between the speaker and the question and answer content. It can clearly and accurately detect whether the role of the currently speaking interactive object is the expected role, so as to deal with cheating and violation operations such as substitute question and answer and inducement to answer.

[0068] Figure 3 A flowchart illustrating another role detection method provided in an embodiment of this application. For example... Figure 3 As shown, the method may include, but is not limited to, the following steps.

[0069] In step 301, the first multimedia data of the interactive object within the question-and-answer time interval of the current question-and-answer session is obtained.

[0070] In the embodiments of this application, step 301 can be implemented in any of the ways described in the various embodiments of this application. The embodiments of this application do not limit this, nor will they be described in detail.

[0071] In step 302, biometric features are extracted from the first multimedia data, and the first biometric vector of the interactive object is obtained based on the biometric features in the first multimedia data.

[0072] In the embodiments of this application, step 302 can be implemented in any of the ways described in the various embodiments of this application. The embodiments of this application do not limit this, nor will they be described in detail.

[0073] In step 303, the text information of the first multimedia data is obtained, and the first role tag information of the first multimedia data is obtained based on the text information.

[0074] In the embodiments of this application, step 303 can be implemented in any of the embodiments of this application. The embodiments of this application do not limit this and will not be described in detail.

[0075] In step 304, a reference biometric vector and the mapping relationship between the reference biometric vector and the reference role tag information are determined.

[0076] In one possible implementation, the reference biometric vector can be pre-registered by the interactive object. For example, a biometric vector registration function can be provided to obtain the biometric vector registered by the interactive object, and this registered biometric vector can be stored in a database as a reference biometric vector. Furthermore, when registering the biometric vector of an interactive object, the reference biometric vector can also be bound to the interactive object's reference role tag information to obtain a mapping relationship between the reference biometric vector and the reference role tag information. This mapping relationship can then be stored for use in subsequent question-and-answer role detection.

[0077] In another possible implementation, the reference biometric vector can be a biometric vector obtained from the multimedia data of the interacting object during the initial question-and-answer session in a dual-recording quality inspection scenario. This biometric vector is then cached as a reference biometric vector for use in subsequent question-and-answer role detection. The mapping relationship between this reference biometric vector and reference role label information can be as follows: during the initial question-and-answer session in a dual-recording quality inspection scenario, corresponding role label information is obtained from the text information obtained from the multimedia data of the interacting object. This role label information is then used as reference role label information. A mapping relationship between the reference biometric vector and the reference role label information is established and cached for use in subsequent question-and-answer role detection.

[0078] It is worth noting that, in the embodiments of this application, when exiting the dual-recording quality inspection scenario, the cached reference biometric vector and the mapping relationship between the reference biometric vector and the reference role tag information can be deleted.

[0079] In step 305, the similarity between the first biometric vector and the reference biometric vector is calculated to obtain the similarity between the first biometric vector and the reference biometric vector.

[0080] In one possible implementation, the diagonal cosine value (cosine of the included angle) can be used to calculate the similarity between the first biometric vector and the reference biometric vector. Alternatively, other similarity calculation methods can be used to calculate the similarity between the first biometric vector and the reference biometric vector, such as calculating similarity through distance (e.g., Euclidean distance, or Manhattan distance). This application does not specifically limit or elaborate on this method.

[0081] In step 306, in response to a similarity greater than or equal to a preset threshold, the first reference role label information corresponding to the reference biometric vector is determined.

[0082] In one possible implementation, when the similarity is greater than or equal to a pre-set threshold, the reference biometric vector and the first biometric vector can be considered to correspond to the same person; when the similarity is less than or equal to the threshold, the reference biometric vector and the first biometric vector can be considered to be biometric vectors of different people.

[0083] In embodiments of this application, when the similarity between the reference biometric vector and the first biometric vector is greater than or equal to a preset threshold, the mapping relationship between the reference biometric vector and the reference role label information can be used to determine the first reference role label information corresponding to the reference biometric vector. As an example, the preset threshold can be 0.6.

[0084] In step 307, in response to the fact that the first role label information is the same as the first reference role label information, the interaction object is determined to be the expected role within the question-and-answer time interval.

[0085] In the embodiments of this application, the first reference role label information corresponding to the reference biometric vector can be compared with the first role label information. If the first role label information and the first reference role label information are the same, such as both being business personnel or both being customer A, then the interaction object can be determined as the expected role within the question-and-answer time period.

[0086] Optionally, in some embodiments of this application, the method may further include: in response to the similarity being less than a preset threshold, and / or the first role tag information being different from the first reference role tag information, determining that the interaction object is an unexpected role within the question-and-answer time interval.

[0087] In this embodiment, if the similarity between the reference biometric vector and the first biometric vector is less than a preset threshold, the interaction object can be considered as an unexpected role within the question-and-answer time interval, that is, the interaction object is not the expected role within the question-and-answer time interval.

[0088] In this embodiment, if the first role tag information is determined to be different from the first reference role tag information, the interaction object can be considered to be an unexpected role within the question-and-answer time period.

[0089] In this embodiment, if the similarity between the reference biometric vector and the first biometric vector is less than a preset threshold, and the first role label information is different from the first reference role label information, the interaction object can be considered as an unexpected role within the question-and-answer time interval.

[0090] For example, taking the first biometric vector as the voiceprint feature vector and the reference biometric vector as the reference voiceprint feature vector, the mapping relationship between the reference voiceprint feature vector and the reference role label information can be determined. The similarity between the voiceprint feature vector of the interaction object and the reference voiceprint feature vector is calculated. If the similarity is greater than or equal to a preset threshold, the first reference role label information corresponding to that reference voiceprint feature vector is determined. The first reference role label information is compared with the first role label information. If the first role label information and the first reference role label information are the same, such as both being business personnel or both being customer A, then the interaction object can be determined as the expected role within the question-and-answer time period.

[0091] If the similarity is less than a preset threshold, and / or the first role label information is different from the first reference role label information, then the interaction object can be considered as an unexpected role within the question-and-answer time period, that is, the interaction object is not the expected role within the question-and-answer time period.

[0092] By implementing the embodiments of this application, it is possible to further determine whether the roles (such as business personnel and customers) taking turns answering questions during the dual-recording quality inspection question-and-answer process conform to the expected correspondence by using biometric comparison technology such as similarity calculation. It can clearly detect whether the role of the currently speaking interaction object is the expected role, so as to deal with cheating and violation operations such as substitute question-and-answer and inducement to answer.

[0093] It should be noted that the reference biometric vector in this application can be a biometric vector obtained through the multimedia data of the interacting object during the initial question-and-answer session in a dual-recording quality inspection scenario. This biometric vector is then cached as a reference biometric vector for use in subsequent question-and-answer role detection. The mapping relationship between this reference biometric vector and reference role label information can be as follows: during the initial question-and-answer session in a dual-recording quality inspection scenario, text information obtained through the multimedia data of the interacting object is used to obtain corresponding role label information. This role label information is then used as reference role label information. A mapping relationship between the reference biometric vector and the reference role label information is established and cached for use in subsequent question-and-answer role detection. Optionally, such as... Figure 4 As shown, the reference biometric vector, and the mapping relationship between the reference biometric vector and the reference role label information, can be obtained in advance through the following steps:

[0094] Step 401: During the first question-and-answer session in the dual-recording quality inspection scenario, acquire the second multimedia data of the interactive object within the question-and-answer time interval of the first question-and-answer session.

[0095] In the embodiments of this application, the initial question-and-answer session can be understood as the first question-and-answer session that begins in the current dual-recording quality inspection scenario.

[0096] In the embodiments of this application, the implementation method of obtaining the second multimedia data in step 401 is similar to the implementation method of obtaining the first multimedia data. It can be implemented by the method of obtaining the first multimedia data in this application, and will not be described again here.

[0097] Step 402: Extract biometric features from the second multimedia data, obtain the second biometric vector of the interactive object based on the biometric features in the second multimedia data, and determine the second biometric vector as the reference biometric vector.

[0098] In the embodiments of this application, the implementation of biometric features in the second multimedia data and the second biometric vector of the interactive object is similar to the implementation of biometric features in the first multimedia data and the first biometric vector of the interactive object. It can be implemented using the implementation method of biometric features in the first multimedia data and the first biometric vector of the interactive object in this application, and will not be described again here.

[0099] Step 403: Obtain the text information of the second multimedia data, obtain the second role tag information of the second multimedia data based on the text information of the second multimedia data, and determine the second role tag information as the reference role tag information.

[0100] In the embodiments of this application, the implementation method for obtaining the text information of the second multimedia data and the second role tag information of the second multimedia data is similar to the implementation method for obtaining the text information of the first multimedia data and the first role tag information of the first multimedia data. It can be implemented using the implementation method for obtaining the text information of the first multimedia data and the first role tag information of the first multimedia data in this application, and will not be described again here.

[0101] Step 404: Bind the reference biometric vector to the reference role tag information to establish a mapping relationship between the reference biometric vector and the reference role tag information.

[0102] In this application, the biometric vectors of the interactive objects are cached in the first question-and-answer session of the dual-recording quality inspection scenario for use in subsequent question-and-answer role detection.

[0103] To ensure that the interaction object's role during the initial question-and-answer session in a dual-recording quality inspection scenario is the expected role, and thus to guarantee the correct mapping between the cached reference biometric vector and reference role label information, the interaction object can be authenticated before binding the reference biometric vector to the reference role label information. Optionally, in some embodiments of this application, valid identity information of the interaction object can be obtained, and role-based authentication can be performed on the interaction object based on this valid identity information. If the interaction object passes the role-based authentication, the reference biometric vector is bound to the reference role label information. This authentication includes, but is not limited to, ID card recognition and / or facial recognition.

[0104] Optionally, in some embodiments of this application, when an unexpected role is detected as the interaction target during the question-and-answer period, a warning can be issued using a preset method. For example, a warning can be issued via a pop-up window. Alternatively, a warning can be issued via voice announcement. Other methods can also be used for warnings, which are not specifically limited in this application and will not be elaborated upon further.

[0105] Figure 5 This is a block diagram of a role detection device provided in an embodiment of this application. Figure 5 As shown, the character detection device may include: a first acquisition module 501, a feature extraction module 502, a second acquisition module 503, a third acquisition module 504, and a detection module 505.

[0106] The first acquisition module 501 is used to acquire the first multimedia data of the current interactive object.

[0107] The feature extraction module 502 is used to extract biological features from the first multimedia data.

[0108] The second acquisition module 503 is used to acquire a first biometric vector of the interactive object based on the biometric features in the first multimedia data. In one possible implementation, the first multimedia data includes voice data, and the first biometric vector includes a voiceprint feature vector; and / or, the first multimedia data includes video data, and the first biometric vector includes a lip movement feature vector.

[0109] The third acquisition module 504 is used to acquire the text information of the first multimedia data and acquire the first role tag information of the first multimedia data based on the text information.

[0110] In one possible implementation, the third acquisition module 504 is specifically used to: extract the first keyword from the text information; obtain the role tag information corresponding to the first keyword according to the pre-established mapping relationship between the keyword and the role tag information, and determine the role tag information corresponding to the first keyword as the first role tag information of the first multimedia data.

[0111] The detection module 505 is used to detect whether the interactive object is the expected role based on the first role tag information and the first biometric vector.

[0112] In one possible implementation, the detection module 505 is specifically used to: determine a reference biometric vector and the mapping relationship between the reference biometric vector and reference role label information; calculate the similarity between the first biometric vector and the reference biometric vector to obtain the similarity between the first biometric vector and the reference biometric vector; in response to the similarity being greater than or equal to a preset threshold, determine the first reference role label information corresponding to the reference biometric vector; and in response to the first role label information being the same as the first reference role label information, determine that the interaction object is the expected role.

[0113] In some embodiments of this application, the detection module 505 may also be used to: determine that the interaction object is an unexpected role in response to the similarity being less than a preset threshold, and / or the first role label information being different from the first reference role label information.

[0114] Optionally, in some embodiments of this application, such as Figure 6 As shown, the device may further include a binding module 606. The first acquisition module 601 is further configured to acquire the second multimedia data of the interactive object during the initial question-and-answer session of the dual-recording quality inspection scenario. The feature extraction module 602 is further configured to extract biometric features from the second multimedia data; the second acquisition module 603 is further configured to acquire the second biometric vector of the interactive object based on the biometric features in the second multimedia data, and determine the second biometric vector as a reference biometric vector; the third acquisition module 604 is further configured to acquire the text information of the second multimedia data, acquire the second role tag information of the second multimedia data based on the text information, and determine the second role tag information as reference role tag information. The binding module 606 is configured to bind the reference biometric vector with the reference role tag information to establish a mapping relationship between the reference biometric vector and the reference role tag information. Figure 6 601-605 and Figure 5 The 501-505 series have the same function and structure.

[0115] Optionally, in some embodiments of this application, such as Figure 7As shown, the device may further include an authentication module 707. The authentication module 707 is used to obtain valid identity information of the interaction object, perform role-based authentication on the interaction object based on the valid identity information, and determine whether the interaction object has passed the role-based authentication. Figure 7 701-706 and Figure 6 The 601-606 series have the same function and structure.

[0116] Optionally, in some embodiments of this application, such as Figure 8 As shown, the device may further include a warning notification module 808. The warning notification module 808 is used to issue a warning notification in a preset manner when the detection module detects that the interactive object is an unexpected role. Figure 8 801-807 and Figure 7 The Chinese 701-707 have the same function and structure.

[0117] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0118] According to embodiments of this application, this application also provides an electronic device and a readable storage medium.

[0119] like Figure 9 The diagram shown is a block diagram of an electronic device for a role detection method according to an embodiment of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0120] like Figure 9As shown, the electronic device includes one or more processors 901, a memory 902, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 9 Take the 901 processor as an example.

[0121] The memory 902 is the non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the role detection method provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions for causing a computer to perform the role detection method provided in this application.

[0122] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the role detection method in the embodiments of this application. The processor 901 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 902, thereby implementing the role detection method in the above method embodiments.

[0123] The memory 902 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 902 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 902 may optionally include memory remotely located relative to the processor 901, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0124] The electronic device may also include an input device 903 and an output device 904. The processor 901, memory 902, input device 903, and output device 904 can be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.

[0125] Input device 903 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as touch screens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. Output device 904 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The display device may include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device may be a touch screen.

[0126] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0127] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0129] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0130] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0131] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0132] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A role detection method, comprising: Obtain the first multimedia data of the interactive objects within the question-and-answer time period of the current question-and-answer session; Extract biometric features from the first multimedia data, and obtain the first biometric feature vector of the interactive object based on the biometric features in the first multimedia data; Obtain the text information of the first multimedia data, and obtain the first role tag information of the first multimedia data based on the text information; Based on the first role tag information and the first biometric vector, detecting whether the interactive object is the expected role includes: Determine the reference biometric vector and the mapping relationship between the reference biometric vector and the reference role label information; The similarity between the first biometric vector and the reference biometric vector is calculated to obtain the similarity between the first biometric vector and the reference biometric vector. In response to the similarity being greater than or equal to a preset threshold, first reference role label information corresponding to the reference biometric vector is determined; In response to the fact that the first role tag information is the same as the first reference role tag information, the interaction object is determined to be the expected role within the question-and-answer time interval.

2. The method as described in claim 1, wherein, The first multimedia data includes voice data, and the first biometric vector includes a voiceprint feature vector; and / or, The first multimedia data includes video data, and the first biometric vector includes lip movement feature vectors.

3. The method as described in claim 1, wherein, The step of obtaining the first role tag information of the first multimedia data based on the text information includes: Extract the first keyword from the text information; Based on the pre-established mapping relationship between keywords and role tag information, the role tag information corresponding to the first keyword is obtained, and the role tag information corresponding to the first keyword is determined as the first role tag information of the first multimedia data.

4. The method of claim 1, wherein, The step of detecting whether the interactive object is the expected role based on the first role tag information and the first biometric vector further includes: In response to the similarity being less than the preset threshold, and / or the first role tag information being different from the first reference role tag information, the interaction object is determined to be an unexpected role.

5. The method as described in claim 1 or 4, wherein, The reference biometric vector, and the mapping relationship between the reference biometric vector and the reference role tag information, are obtained in advance in the following manner: During the initial question-and-answer session in a dual-recording quality inspection scenario, the second multimedia data of the interactive object is acquired within the question-and-answer time interval of the initial question-and-answer session. Extract biometric features from the second multimedia data, obtain a second biometric vector of the interactive object based on the biometric features in the second multimedia data, and determine the second biometric vector as a reference biometric vector; Obtain the text information of the second multimedia data, and obtain the second role tag information of the second multimedia data based on the text information of the second multimedia data, and determine the second role tag information as reference role tag information; The reference biometric vector is bound to the reference role tag information to establish a mapping relationship between the reference biometric vector and the reference role tag information.

6. The method of claim 5, wherein before binding the reference biometric vector to the reference role tag information, the method further comprises: Obtain the valid identity information of the interaction object; The interaction object is authenticated based on its role using the valid identity information. Determine that the interaction object has passed the role-based authentication.

7. The method of claim 1, further comprising: In response to detecting that the interaction object is an unexpected role, a warning prompt is issued in a preset manner.

8. A character detection device, comprising: The first acquisition module is used to acquire the first multimedia data of the interactive objects within the question-and-answer time period of the current question-and-answer session; The feature extraction module is used to extract biological features from the first multimedia data; The second acquisition module is used to acquire the first biometric vector of the interactive object based on the biometric features in the first multimedia data; The third acquisition module is used to acquire the text information of the first multimedia data and acquire the first role tag information of the first multimedia data based on the text information; The detection module is used to detect whether the interactive object is the expected role based on the first role tag information and the first biometric vector, including: Determine the reference biometric vector and the mapping relationship between the reference biometric vector and the reference role label information; The similarity between the first biometric vector and the reference biometric vector is calculated to obtain the similarity between the first biometric vector and the reference biometric vector. In response to the similarity being greater than or equal to a preset threshold, first reference role label information corresponding to the reference biometric vector is determined; In response to the fact that the first role tag information is the same as the first reference role tag information, the interaction object is determined to be the expected role.

9. The apparatus of claim 8, wherein, The first multimedia data includes voice data, and the first biometric vector includes a voiceprint feature vector; and / or, The first multimedia data includes video data, and the first biometric vector includes lip movement feature vectors.

10. The apparatus of claim 8, wherein, The third acquisition module is specifically used for: Extract the first keyword from the text information; Based on the pre-established mapping relationship between keywords and role tag information, the role tag information corresponding to the first keyword is obtained, and the role tag information corresponding to the first keyword is determined as the first role tag information of the first multimedia data.

11. The apparatus of claim 8, wherein, The detection module is also used for: In response to the similarity being less than the preset threshold, and / or the first role tag information being different from the first reference role tag information, the interaction object is determined to be an unexpected role.

12. The apparatus of claim 8 or 11, wherein, The first acquisition module is also used to acquire the second multimedia data of the interactive object during the question-and-answer time interval of the first question-and-answer session in the dual-recording quality inspection scenario. The feature extraction module is also used to extract biological features from the second multimedia data; The second acquisition module is further configured to acquire the second biometric vector of the interactive object based on the biometric features in the second multimedia data, and determine the second biometric vector as a reference biometric vector; The third acquisition module is further configured to acquire the text information of the second multimedia data, acquire the second role tag information of the second multimedia data based on the text information of the second multimedia data, and determine the second role tag information as reference role tag information; The device further includes: The binding module is used to bind the reference biometric vector to the reference role tag information to establish a mapping relationship between the reference biometric vector and the reference role tag information.

13. The apparatus of claim 12, further comprising: An authentication module is used to obtain valid identity information of the interaction object, perform role-based authentication on the interaction object based on the valid identity information, and determine that the interaction object passes the role-based authentication.

14. The apparatus of claim 8, further comprising: The warning module is used to issue a warning in a preset manner when the detection module detects that the interactive object is an unexpected role.

15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice role segmentation method and device, computer equipment and storage medium

    CN113192516A

  • Double-recording file quality inspection method and device, computer equipment and storage medium

    CN114579523A