Anchei detection method, device and computer readable storage medium

By collecting and matching voiceprint information, and combining it with the dispersion of tone information sets for dual verification, the accuracy problem of employee cheating detection is solved, the company's understanding of employees' mastery of communication skills is improved, and training costs are reduced.

CN116030834BActive Publication Date: 2026-04-21VOICEAI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VOICEAI TECH CO LTD
Filing Date
2022-12-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

During employee training and assessment, employees may obtain high scores through cheating methods (such as taking exams for others), making it impossible for companies to accurately grasp the true level of employees' proficiency in communication skills.

Method used

By collecting the audio of the user to be detected, the voiceprint information is extracted and matched with the preset voiceprint information. If the match is successful, the tone information set is extracted, and cheating detection is performed based on the dispersion of the tone information set. The voiceprint information is then used for double verification.

Benefits of technology

It improved the accuracy of cheating detection, ensured the authenticity of employee assessments, reduced the company's human resource training costs, and eliminated internal malpractices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030834B_ABST
    Figure CN116030834B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cheating detection method, device and computer readable storage medium. The application embodiment extracts the voiceprint information of the audio to be detected of the user to be detected by collecting the audio to be detected;Voiceprint information is matched with preset voiceprint information;When detecting that voiceprint information and preset voiceprint information match successfully, the pitch information set of the audio to be detected is extracted;According to the dispersion of pitch information set, the cheating detection of the user to be detected is carried out. To this end, after verification by voiceprint information, double verification can also be carried out according to the dispersion of the extracted pitch information set, effectively ensuring the authenticity of each employee cheating examination, greatly improving the accuracy of cheating detection, so that enterprises can more accurately understand the real skill situation of employees.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction, specifically to a cheating detection method, device, and computer-readable storage medium. Background Technology

[0002] With social development, my country's labor-intensive and service-oriented enterprises are booming. For positions in these enterprises where communication is the primary application scenario (such as customer service, debt collection, and telemarketing), employees need to be proficient in basic communication skills to provide good service to customers. In order to shorten the training cycle and facilitate enterprises' understanding of employee training progress, some companies are currently using online training systems to conduct systematic training and assessment for new employees.

[0003] However, in actual performance evaluations, employees may cheat (e.g., taking exams for others) to pass the evaluation and secure a job. Therefore, how to accurately detect cheating during evaluations so that companies can understand the true level of employees' proficiency in sales scripts is a problem that needs to be solved. Summary of the Invention

[0004] This application provides a cheating detection method, apparatus, and computer-readable storage medium, which can improve the accuracy of cheating detection.

[0005] To address the aforementioned technical problems, the embodiments of this application provide the following technical solutions:

[0006] A cheating detection method includes:

[0007] Collect the audio data of the user to be tested;

[0008] Extract the voiceprint information of the audio to be detected, and match the voiceprint information with preset voiceprint information;

[0009] When it is detected that the voiceprint information matches the preset voiceprint information, the pitch information set of the audio to be detected is extracted;

[0010] Based on the dispersion of the tone information set, cheating detection is performed on the user to be detected.

[0011] A cheating detection device, comprising:

[0012] The acquisition unit is used to acquire the audio of the user to be tested;

[0013] A matching unit is used to extract the voiceprint information of the audio to be detected and match the voiceprint information with preset voiceprint information;

[0014] The extraction unit is used to extract the pitch information set of the audio to be detected when the voiceprint information is detected to match the preset voiceprint information.

[0015] The detection unit is used to detect cheating by the user to be detected based on the dispersion of the tone information set.

[0016] In some embodiments, the extraction unit is configured to:

[0017] When the voiceprint information is detected to match the preset voiceprint information, the pitch information of the audio to be detected is extracted sequentially according to the audio time sequence and the preset acquisition rules to obtain the pitch information set.

[0018] In some embodiments, the detection unit includes:

[0019] The first determining subunit is used to determine the corresponding dispersion based on the distribution state of multiple tone information in the tone information set;

[0020] The comparison subunit is used to compare the dispersion with a preset dispersion.

[0021] The preset dispersion is obtained by performing dispersion statistics based on historical audio information;

[0022] The second determining subunit is used to determine the cheating detection result of the user to be detected based on the dispersed comparison result.

[0023] In some embodiments, the first determining subunit is configured to:

[0024] Calculate the variance of multiple pitch information in the pitch information set;

[0025] The step of comparing the dispersion with a preset dispersion includes:

[0026] Calculate the degree of difference between the variance value and the preset variance value;

[0027] The corresponding dispersion comparison results are determined based on the degree of difference.

[0028] In some embodiments, the second determining subunit is configured to:

[0029] When the detected difference is greater than a preset threshold, the dispersed comparison result is determined to be a discrete result, and the user to be detected is determined to be cheating.

[0030] When the detected difference is not greater than a preset threshold, the dispersed comparison result is determined to be the centralized result, and the user to be tested is determined to be a non-cheating result.

[0031] In some embodiments, the apparatus further includes a verification unit for:

[0032] After confirming that the user being tested has cheated, lock the screen;

[0033] Start the camera equipment and capture facial images;

[0034] The screen is unlocked when a face image is detected to match a preset face image.

[0035] In some embodiments, the apparatus further includes a user matching unit for:

[0036] When it is detected that the voiceprint information does not match the preset voiceprint information, the voiceprint information will be matched with each historical voiceprint information in the database.

[0037] Identify the target historical voiceprint information that matches the voiceprint information;

[0038] Obtain the target user corresponding to the target's historical voiceprint information.

[0039] A computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the steps in the cheating detection method described above.

[0040] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the cheat detection method provided above.

[0041] A computer program product or computer program includes computer instructions stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium and executes the computer instructions, causing the computer device to perform the steps in the cheating detection method provided above.

[0042] This application embodiment collects the audio of the user to be tested; extracts the voiceprint information of the audio and matches it with preset voiceprint information; when a successful match is detected, extracts the pitch information set of the audio; and performs cheating detection on the user based on the dispersion of the pitch information set. Thus, after verification using voiceprint information, a double verification can be performed based on the dispersion of the extracted pitch information set, effectively ensuring the authenticity of each employee's cheating assessment and greatly improving the accuracy of cheating detection, enabling companies to more accurately understand the true proficiency of their employees' speech skills. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram of a cheating detection system provided in an embodiment of this application.

[0045] Figure 2 This is a flowchart illustrating the cheating detection method provided in the embodiments of this application;

[0046] Figure 3 This is another schematic flowchart of the cheating detection method provided in the embodiments of this application;

[0047] Figure 4 This is a schematic diagram of the cheating detection device provided in the embodiments of this application;

[0048] Figure 5 This is a schematic diagram of the terminal structure provided in the embodiments of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] This application provides a cheating detection method, apparatus, and computer-readable storage medium.

[0051] Please see Figure 1 , Figure 1 This is a schematic diagram of a cheating detection system provided in an embodiment of this application, including: a terminal and a server. The terminal and server can be connected via a communication network, which includes wireless and wired networks. The wireless network includes one or more combinations of wireless wide area networks, wireless local area networks, wireless metropolitan area networks, and wireless personal networks. Network entities such as routers and gateways are included in the network, but are not shown in the diagram. The terminal can interact with the server through the communication network; for example, the terminal can collect the audio of the user to be detected and send it to the server.

[0052] This terminal can be a tablet computer, mobile phone, laptop computer, desktop computer, or any other device with a storage unit and a microprocessor that enables computing power, used to collect the audio data of the user to be tested.

[0053] The cheating detection system may include a cheating detection device, which may be integrated into a server. This server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The system is used to collect audio from the user to be detected via a terminal; extract the voiceprint information of the audio to be detected and match it with preset voiceprint information; when a successful match is detected, extract the pitch information set of the audio to be detected; and perform cheating detection on the user to be detected based on the dispersion of the pitch information set.

[0054] It should be noted that, Figure 1 The schematic diagram of the cheating detection system shown is merely an example. The cheating detection system and scenario described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of cheating detection systems and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0055] The following sections will provide detailed explanations.

[0056] This application provides a cheating detection method, which can be executed by a client on a terminal or by a server. This application embodiment uses execution by a client on a terminal as an example for illustration.

[0057] It should be noted that while some processes described in the specification, claims, and drawings include multiple steps appearing in a specific order, it should be clearly understood that these steps may not be performed in the order they appear herein, or may be performed in parallel. The step numbers are merely used to distinguish different steps and do not themselves represent any execution order. Furthermore, descriptions such as "first" and "second" in this document are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0058] Please see Figure 2 , Figure 2 This is a flowchart illustrating the cheating detection method provided in an embodiment of this application. The cheating detection method includes:

[0059] In step 101, the audio of the user to be tested is collected.

[0060] With social development, my country's labor-intensive and service-oriented enterprises are booming. For positions in these enterprises where communication is the primary application scenario (such as customer service, debt collection, and telemarketing), employees need to be proficient in basic communication skills to provide excellent customer service. Therefore, newly recruited employees often need to be trained by experienced staff to familiarize them with customer communication processes and the appropriate communication techniques for different situations. This training method is inefficient in knowledge distribution and has a long training cycle, resulting in slow onboarding for new employees. Furthermore, since one experienced employee needs to train multiple new employees, it's difficult to ensure each new employee's learning progress is adequately monitored, consuming significant human resources, and the company cannot accurately track employee training progress.

[0061] To address this situation, companies are now using online training systems for new employee training and assessment. However, during the actual assessment process, employees may cheat (e.g., by having someone else take the test) to obtain high scores and quickly pass the training and assessment. Therefore, how to detect cheating in assessments so that companies can accurately assess employees' true proficiency in communication skills is a pressing issue that needs to be resolved.

[0062] In order to solve the above problems, this application embodiment can collect the audio of the user to be tested through a sound acquisition device on the client during employee training. The sound acquisition device can be a microphone or microphone array, etc. The user to be tested is the user currently participating in the assessment, and the audio to be tested is the audio signal collected during the assessment process.

[0063] In one implementation, each enterprise user can bind a corresponding enterprise account, which pre-stores the user's relevant information. When the user is training, they need to log in with the corresponding enterprise account to record the assessment process and results of the enterprise account, which facilitates enterprise management.

[0064] In one implementation, the client can prompt the user to begin training or assessment with a voice message after the training or assessment has started, such as "Assessment begins. Please answer according to the following scenario." After the voice prompt, the client turns on the microphone to collect the audio of the user to be tested.

[0065] In step 102, the voiceprint information of the audio to be detected is extracted and matched with preset voiceprint information.

[0066] To better illustrate the embodiments of this application, voiceprint is explained as a sound wave spectrum carrying speech information, displayed using electroacoustic instruments. Modern scientific research shows that voiceprints not only possess specificity but also relative stability. After adulthood, a person's voice can remain relatively stable over a long period.

[0067] In this way, voiceprint information can be extracted from the audio to be detected for accurate cheating detection. The preset voiceprint information is the voiceprint information pre-stored by the user to be detected. The preset voiceprint information can be bound to the account information of the user to be detected. In this way, the client can perform similarity matching based on the voiceprint information of the audio to be detected and the corresponding bound preset voiceprint information.

[0068] Due to the uniqueness of voiceprint information, the similarity between a user's voiceprint and a preset voiceprint can be extremely high, reaching over 97%. Therefore, a similarity threshold of 95% can be set. When the similarity between the voiceprint and the preset voiceprint is greater than the threshold, it is determined that the voiceprint and preset voiceprint are successfully matched, and step 103 is executed. Conversely, when the similarity between the voiceprint and the preset voiceprint is not greater than the threshold, it is determined that the voiceprint and preset voiceprint are not successfully matched, indicating that the user participating in the training or assessment is not the actual person and that there is proxy testing, and the user is directly identified as a cheating user.

[0069] In some implementations, the method further includes:

[0070] (1) When it is detected that the voiceprint information does not match the preset voiceprint information, the voiceprint information is matched with each historical voiceprint information in the database in turn;

[0071] (2) Determine the target historical voiceprint information that matches the voiceprint information;

[0072] (3) Obtain the target user corresponding to the target's historical voiceprint information.

[0073] When the detected voiceprint information fails to match the preset voiceprint information, it indicates that the user participating in the training or assessment is not the actual person and that there is a proxy test-taking phenomenon. The user is directly identified as a cheater. The cheating voiceprint information can be sequentially matched with each historical voiceprint information stored in the background to find the target historical voiceprint information that matches the current cheating voiceprint information (i.e., the most similar). The target historical voiceprint information is used to determine the target user of the cheating, for example, the target user of the cheating is identified as "Zhang San". In this way, a warning or punishment can be generated for the target user, thereby increasing the monitoring of cheating and eliminating bad practices within the enterprise.

[0074] In step 103, when the voiceprint information is detected to match the preset voiceprint information, the pitch information set of the audio to be detected is extracted.

[0075] When the similarity between the detected voiceprint information and the preset voiceprint information is greater than the similarity threshold, it is determined that the voiceprint information and the preset voiceprint information are successfully matched, indicating that the probability of the current user participating in the training or assessment being the user himself is extremely high.

[0076] As cheating methods become increasingly sophisticated in the market, voiceprint imitators have also emerged. Some cheating software can cheat by imitating voiceprints. Therefore, the method of verifying voiceprint information alone is clearly insufficient to meet current requirements. Thus, this application's embodiments creatively propose that when a successful match between voiceprint information and preset voiceprint information is detected, the pitch information set of the audio to be detected should be extracted.

[0077] Pitch refers to the highness or lowness of a sound. The pitch is related to the structure of the vocal body, as the structure of the vocal body affects the frequency of the sound. The unit of pitch can be melodic (mel). Since the structure of the human vocal body is fixed, the range of pitch is also within a certain range. Therefore, the set of pitch information of the audio to be detected can be extracted as a pitch information set. This pitch information set consists of multiple pitch information, and each pitch information can be generated for one pronunciation.

[0078] In some implementations, the extraction of the pitch information set of the audio to be detected includes: extracting the pitch information of the audio to be detected sequentially according to the audio time sequence and preset acquisition rules to obtain the pitch information set.

[0079] The preset acquisition rule can be to acquire one tone information for each pronunciation. In this way, the tone information of the audio to be detected can be extracted according to the audio time sequence and the preset acquisition rule to obtain a tone information set. This tone information set can reflect the dispersion of the tone produced by the user, that is, the tone pattern.

[0080] In step 104, cheating detection is performed on the user to be detected based on the dispersion of the tone information set.

[0081] It is understandable that since the vocal structure of the same user does not change, the pattern of the emitted pitch information will be within a certain range. Therefore, it is possible to collect the rules of pitch information emitted by the user under historical audio information and generate a preset dispersion. This preset dispersion represents the central pattern of the pitch information emitted by the user under test, which can also be understood as the degree of deviation of each pitch information from the mathematical expectation (mean).

[0082] Based on this, the system determines whether the user is currently logged into the company account by comparing the difference between the dispersion of each tone information in the tone information set and the preset dispersion. When the difference is close, it indicates that the tone information set of the current speech is from the user, and the user is considered to be undergoing normal training or assessment. Conversely, when the difference is large, it indicates that the tone information set of the current speech is not from the user, and the user is considered to be cheating.

[0083] In some implementations, the cheating detection of the user to be detected is based on the dispersion of the pitch information set, including:

[0084] (1) Determine the corresponding dispersion based on the distribution of multiple tone information in the tone information set;

[0085] (2) Compare the dispersion with the preset dispersion;

[0086] The preset dispersion is obtained by performing dispersion statistics based on historical audio information;

[0087] (3) Based on the results of the distributed comparison, determine the cheating detection results of the user to be detected.

[0088] The dispersion reflects the regularity of the distribution among multiple tone information. The more concentrated the distribution of multiple tone information, the smaller the dispersion and the more regular the distribution. Conversely, the less concentrated the distribution of multiple tone information, the larger the dispersion and the more irregular the distribution.

[0089] However, since the vocal structure of the same user is constant, the distribution of the pitch information of the same user must be regular. Therefore, dispersion statistics can be performed based on historical audio information. The dispersion of multiple historical pitch information in historical audio information can be statistically analyzed to obtain a preset dispersion. This preset dispersion reflects the regularity of the historical pitch information of the user's vocalization. The regularity of the pitch information emitted by the user in the future is close to the regularity of the preset dispersion.

[0090] In this way, the dispersion of the distribution of multiple pitch information in the current audio pitch information set can be compared with the preset dispersion. When the difference between the two is close, it indicates that the pitch information set of the current speech is the voice of the person being tested, and the user being tested is judged to be undergoing normal training or assessment. Conversely, when the difference between the two is large, it indicates that the pitch information set of the current speech is not the voice of the person being tested, and the user being tested is judged to be cheating.

[0091] Based on this, the embodiments of this application, in addition to voiceprint verification, also perform multiple authentications through the distribution patterns of tone information, which greatly improves the accuracy of cheating detection.

[0092] In some implementations, determining the corresponding dispersion based on the distribution of multiple pitch information in the pitch information set includes:

[0093] (1.1) Calculate the variance of multiple pitch information in the pitch information set;

[0094] The step of comparing the dispersion with a preset dispersion includes:

[0095] (1.2) Calculate the degree of difference between the variance value and the preset variance value;

[0096] (1.3) Determine the corresponding dispersion comparison results based on the degree of difference.

[0097] Variance, in probability theory and statistics, measures the dispersion of a random variable or a set of data. In probability theory, variance measures the degree of deviation between a random variable and its expected value (mean). In statistics, variance (sample variance) is the average of the squared differences between each sample value and the mean of all sample values. Studying variance, or the degree of deviation, is of great significance in many practical problems.

[0098] In this embodiment, the variance of multiple pitch information in the pitch information set can be calculated using a variance calculation formula to determine the degree of deviation of the pitch information set. It should be noted that this preset variance value is the average of the variances of multiple historical pitch information in the statistical historical audio information, which can reflect the pitch patterns of the historical audio information. Therefore, the difference between the current variance value and the preset variance value can be used to measure whether they belong to the same user.

[0099] The difference can be used to determine the corresponding dispersion comparison result. When the difference is close, it means that the tone information set of the voice is the user's own voice, and the user under test is judged to have undergone normal training or assessment. Conversely, when the difference is large, it means that the tone information set of the voice is not the user's own voice, and the user under test is judged to have cheated.

[0100] In some implementations, determining the cheating detection result for the user to be detected based on the dispersion comparison result includes:

[0101] (2.1) When the difference is detected to be greater than the preset threshold, the dispersed comparison result is determined to be a discrete result, and the user to be detected is determined to be a cheating result;

[0102] (2.2) When the difference is detected to be no greater than the preset threshold, the dispersed comparison result is determined to be the centralized result, and the user to be detected is determined to be a non-cheating result.

[0103] The preset threshold is a critical value used to determine whether the tone information set conforms to a pattern. When the detected difference exceeds the preset threshold, it indicates that the central pattern of the current tone information set differs significantly from the central pattern of historical tone information. Therefore, the dispersed comparison result is determined to be a discrete result, i.e., it does not conform to a central pattern, and the user under test is identified as cheating. Conversely, when the detected difference is not greater than the preset threshold, it indicates that the central pattern of the current tone information set is close to the central pattern of historical tone information. Therefore, the dispersed comparison result is determined to be a central result, and the user under test is identified as the person who underwent normal training or assessment.

[0104] In some implementations, the method further includes:

[0105] (3.1) After confirming that the cheating detection result of the user to be detected is a cheating result, lock the screen;

[0106] (3.2) Start the camera equipment and capture facial images;

[0107] (3.3) When a face image is detected to match a preset face image, the screen is unlocked.

[0108] When a user is detected to be cheating, the screen can be locked to prevent further cheating, and a camera can be activated in real time to capture the current facial image. This facial image can be used as evidence for subsequent cheating and to implement subsequent cheating penalties.

[0109] Furthermore, the screen can only be unlocked and operated when the detected face image matches the preset face image bound to the logged-in account information; otherwise, the client will remain locked, preventing cheating users from taking the exam on behalf of others.

[0110] In other embodiments, the method further includes:

[0111] (4.1) After confirming that the cheating detection result of the user to be detected is a cheating result, display the voice verification code on the display screen and lock the screen;

[0112] (4.2) When the voice verification information entered by the user is detected to match the voice verification code, the screen is unlocked.

[0113] In this system, once a user is detected as cheating, the screen can be locked to prevent further cheating. A voice verification code will be displayed, which will indicate the corresponding text, such as "I am Zhang San, ID 9527". The client will then activate its microphone to collect the user's voice verification information. The screen will only be unlocked and the error resolved if the text indicated by the voice verification information matches the text indicated by the voice verification code. In this way, in the event of cheating, verification can be performed through voice verification codes to prevent bots from taking the exam and to improve the diversity of cheating detection.

[0114] As described above, this embodiment of the application collects the audio of the user to be tested; extracts the voiceprint information of the audio and matches it with preset voiceprint information; when a successful match is detected, extracts the pitch information set of the audio; and performs cheating detection on the user based on the dispersion of the pitch information set. Thus, after verification using voiceprint information, double verification can be performed based on the dispersion of the extracted pitch information set, effectively ensuring the authenticity of each employee's cheating assessment and greatly improving the accuracy of cheating detection, enabling companies to more accurately understand the true proficiency of their employees' speech skills.

[0115] In this embodiment, the cheating detection device will be specifically integrated into the terminal as an example for explanation. Please refer to the following description for details.

[0116] Please see Figure 3 , Figure 3 Another schematic diagram of the cheating detection method provided in this application embodiment. The method flow may include:

[0117] In step 201, the terminal collects the audio of the user to be detected.

[0118] Users can log in to the client terminal through a registered corporate account, such as "Zhang San", to start training or assessment. The corporate account can pre-store relevant information of the user to be tested, such as voiceprint information and historical audio information.

[0119] After the training or assessment begins, the client can use voice prompts to guide the user to start the training or assessment. For example, the voice prompt may say, "Assessment begins. Please answer according to the following scenarios." After the voice prompt, the client will turn on the microphone to collect the audio of the user to be tested.

[0120] In step 202, the terminal extracts the voiceprint information of the audio to be detected and matches the voiceprint information with preset voiceprint information.

[0121] The preset voiceprint information is the voiceprint information pre-stored by the user to be detected. The preset voiceprint information can be bound to the account information of the user to be detected. In this way, the client can perform similarity matching between the voiceprint information of the audio to be detected and the corresponding bound preset voiceprint information.

[0122] In step 203, when the terminal detects that the voiceprint information matches the preset voiceprint information, the pitch information of the audio to be detected is extracted sequentially according to the audio time sequence and the preset acquisition rules to obtain the pitch information set.

[0123] When the similarity between the detected voiceprint information and the preset voiceprint information is greater than the similarity threshold, it is determined that the voiceprint information and the preset voiceprint information are successfully matched, indicating that the probability of the current user participating in the training or assessment being the user himself is extremely high.

[0124] With the increasing variety of cheating methods on the market, voiceprint imitators have also emerged. Some cheating software can cheat by imitating voiceprints. Therefore, the means of verifying solely based on voiceprint information are clearly insufficient to meet current requirements. The preset acquisition rule can collect pitch information for each pronunciation. Therefore, this application's embodiment creatively proposes that when a successful match between voiceprint information and preset voiceprint information is detected, the pitch information of the audio to be detected can be extracted according to the audio time sequence and the preset acquisition rule to obtain a pitch information set. This pitch information set can reflect the dispersion of the pitch emitted by the user, i.e., the pitch pattern.

[0125] For example, if the audio to be detected consists of the sentence "Hello, I am customer service Zhang San", and each syllable is a tone information, then we can obtain 8 tone information formed by 8 syllables, forming a tone information set.

[0126] In step 204, the terminal calculates the variance of multiple pitch information in the pitch information set, calculates the degree of difference between the variance and the preset variance, and determines the corresponding dispersion comparison result based on the degree of difference.

[0127] The variance of multiple pitch information values ​​in the pitch information set can be calculated using a variance calculation formula to determine the degree of deviation of the pitch information set. It should be noted that this preset variance value is the average of the variances of multiple historical pitch information values ​​in the statistical historical audio information. This reflects the pitch patterns of the historical audio information. Therefore, the difference between the current variance value and the preset variance value can be used to determine the corresponding dispersion comparison result and assess whether the two belong to the same user.

[0128] In step 205, when the terminal detects that the difference is not greater than a preset threshold, the dispersed comparison result is determined to be a centralized result, and the user to be detected is determined to be a non-cheating result.

[0129] The preset threshold is a critical value used to determine whether the tone information set conforms to a pattern. When the detected difference is not greater than the preset threshold, it indicates that the concentration pattern of the current tone information set is close to the concentration pattern of historical tone information. The result of the dispersion comparison is determined to be a concentration result, and the user to be tested is confirmed to be the person undergoing normal training or assessment.

[0130] In step 206, when the terminal detects a difference greater than a preset threshold, it determines that the dispersed comparison result is a discrete result and determines that the user to be detected is cheating.

[0131] When the detected difference exceeds a preset threshold, it indicates that the concentration pattern of the current tone information set differs significantly from the concentration pattern of historical tone information. The result of the dispersion comparison is determined to be a discrete result, which does not conform to the concentration pattern, and the user to be detected is determined to be cheating.

[0132] Based on voiceprint verification, this application's embodiments innovatively propose multiple authentication methods based on the distribution patterns of tone information, which greatly improves the accuracy of cheating detection and reduces the cost of human resource training for enterprises.

[0133] In step 207, the terminal locks the screen, starts the camera device, and captures a face image. When the face image is detected to match a preset face image, the screen is unlocked.

[0134] In this system, once the detection result of the user being tested is found to be cheating, the screen can be locked to prevent further cheating. Since the current situation has been determined to be a cheating scenario, the camera can be activated in real time to obtain the current facial image. This facial image can be used as evidence for subsequent cheating and to implement subsequent cheating penalties.

[0135] Furthermore, the screen can only be unlocked and operated when the detected face image matches the preset face image bound to the logged-in account information; otherwise, the client will remain locked, preventing cheating users from taking the exam on behalf of others.

[0136] In step 208, when the terminal detects that the voiceprint information does not match the preset voiceprint information, it sequentially matches the voiceprint information with each historical voiceprint information in the database to determine the target historical voiceprint information that matches the voiceprint information and obtain the target user corresponding to the target historical voiceprint information.

[0137] When a voiceprint is detected that fails to match a preset voiceprint, it indicates that the user participating in the training or assessment is not the actual person and that someone else is taking the test on their behalf. This user is directly identified as a cheater. The cheating voiceprint can then be matched sequentially with each historical voiceprint stored in the backend to find the target historical voiceprint that matches the current cheating voiceprint (i.e., has the highest similarity). The target user for cheating can be identified through this target historical voiceprint, for example, the target user "Li Si". Warnings or penalties can then be generated for the target user "Li Si", thereby increasing the monitoring of cheating and eliminating unhealthy practices within the company.

[0138] As described above, this embodiment of the application collects the audio of the user to be tested; extracts the voiceprint information of the audio and matches it with preset voiceprint information; when a successful match is detected, extracts the pitch information set of the audio; and performs cheating detection on the user based on the dispersion of the pitch information set. Thus, after verification using voiceprint information, double verification can be performed based on the dispersion of the extracted pitch information set, effectively ensuring the authenticity of each employee's cheating assessment and greatly improving the accuracy of cheating detection, enabling companies to more accurately understand the true proficiency of their employees' speech skills.

[0139] Furthermore, by monitoring the identity of those who take exams on behalf of others, penalties for cheating can be increased, raising the cost of cheating and eliminating unhealthy practices within companies.

[0140] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of the cheating detection device provided in the embodiments of this application. The cheating detection device may include a collection unit 401, a matching unit 402, an extraction unit 403, and a detection unit 404, etc.

[0141] Acquisition unit 401 is used to acquire the audio of the user to be detected.

[0142] The matching unit 402 is used to extract the voiceprint information of the audio to be detected and match the voiceprint information with preset voiceprint information.

[0143] In some embodiments, the apparatus further includes a user matching unit for:

[0144] When it is detected that the voiceprint information does not match the preset voiceprint information, the voiceprint information is matched with each historical voiceprint information in the database.

[0145] Identify the target historical voiceprint information that matches this voiceprint information;

[0146] Obtain the target user corresponding to the target's historical voiceprint information.

[0147] The extraction unit 403 is used to extract the pitch information set of the audio to be detected when the voiceprint information is detected to match the preset voiceprint information.

[0148] In some embodiments, the extraction unit 403 is used for:

[0149] When the voiceprint information is detected to match the preset voiceprint information, the pitch information of the audio to be detected is extracted sequentially according to the audio time sequence and the preset acquisition rules to obtain the audio pitch set.

[0150] The detection unit 404 is used to detect cheating by the user to be detected based on the dispersion of the tone information set.

[0151] In some embodiments, the detection unit 404 includes:

[0152] The first determining subunit is used to determine the corresponding dispersion based on the distribution of multiple tone information in the tone information set;

[0153] The comparison sub-unit is used to compare the dispersion with a preset dispersion.

[0154] The preset dispersion is obtained by performing dispersion statistics based on historical audio information;

[0155] The second determining subunit is used to determine the cheating detection result of the user to be detected based on the dispersed comparison result.

[0156] In some embodiments, the first determining subunit is configured to:

[0157] Calculate the variance of multiple pitch information in this pitch information set;

[0158] The step of comparing the dispersion with a preset dispersion includes:

[0159] Calculate the degree of difference between this variance value and the preset variance value;

[0160] The corresponding dispersion comparison results are determined based on this degree of difference.

[0161] In some embodiments, the second determining subunit is configured to:

[0162] When the detected difference is greater than the preset threshold, the dispersion comparison result is determined to be a discrete result, and the user to be detected is determined to be cheating.

[0163] When the detected difference is not greater than the preset threshold, the dispersed comparison result is determined to be the centralized result, and the user to be tested is determined to be a non-cheating result.

[0164] In some embodiments, the device further includes a verification unit for:

[0165] After confirming that the user being tested has cheated, lock the screen;

[0166] Start the camera equipment and capture facial images;

[0167] The screen unlocks when a face image is detected that matches a preset face image.

[0168] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.

[0169] This application also provides a computer device, which can be a terminal, such as... Figure 5 As shown, it illustrates the structural diagram of the terminal involved in the embodiments of this application, specifically:

[0170] The computer device may include radio frequency (RF) circuitry 501, a memory 502 including one or more computer-readable storage media, an input unit 503, a display unit 504, a sensor 505, audio circuitry 506, a wireless fidelity (WiFi) module 507, a processor 508 including one or more processing cores, and a power supply 509, among other components. Those skilled in the art will understand that... Figure 5 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0171] RF circuit 501 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 508 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 501 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 501 can also communicate wirelessly with networks and other devices. Wireless communication can use any communication standard or protocol, including but not limited to GSM, GPRS, CDMA, WCDMA, LTE, email, and SMS.

[0172] The memory 502 can be used to store software programs and modules. The processor 508 executes various functional applications and cheat detection by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the terminal (such as audio data, phone book, etc.). In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide access to the memory 502 for the processor 508 and the input unit 503.

[0173] Input unit 503 can be used to receive input numerical or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, in one embodiment, input unit 503 may include a touch-sensitive surface and other input devices. A touch-sensitive surface, also known as a touch display or touchpad, can collect user touch operations on or near it (e.g., user operations using fingers, styluses, or any suitable object or accessory on or near the touch-sensitive surface) and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface may include a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 508, and can receive and execute commands from the processor 508. Furthermore, various types of touch-sensitive surfaces, such as resistive, capacitive, infrared, and surface acoustic wave, can be used. In addition to the touch-sensitive surface, input unit 503 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0174] Display unit 504 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 504 may include a display panel, optionally configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar form. Furthermore, a touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it transmits the information to processor 508 to determine the type of touch event. Subsequently, processor 508 provides corresponding visual output on the display panel according to the type of touch event. Although in Figure 5 In this context, the touch-sensitive surface and the display panel are two separate components for implementing input and output functions. However, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve both input and output functions.

[0175] The terminal may also include at least one sensor 505, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel according to the ambient light level, and the proximity sensor can turn off the display panel and / or backlight when the terminal is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that the terminal may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0176] Audio circuitry 506, a speaker, and a microphone provide an audio interface between the user and the terminal. Audio circuitry 506 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 506, converted back into audio data, and processed by processor 508. The processed data is then transmitted via RF circuitry 501 to, for example, another terminal, or output to memory 502 for further processing. Audio circuitry 506 may also include an earphone jack to facilitate communication between a peripheral headset and the terminal.

[0177] WiFi is a short-range wireless transmission technology. Terminals using the WiFi module 507 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 5 WiFi module 507 is shown, but it is understood that it is not a necessary component of the terminal and can be omitted as needed without changing the essence of the invention.

[0178] The processor 508 is the control center of the terminal, connecting various parts of the phone via various interfaces and lines. It executes software programs and / or modules stored in the memory 502, and calls data stored in the memory 502 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 508 may include one or more processing cores; preferably, the processor 508 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 508.

[0179] The terminal also includes a power supply 509 (such as a battery) to power various components. Preferably, the power supply can be logically connected to the processor 508 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 509 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0180] Although not shown, the terminal may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 508 in the terminal will load the executable files corresponding to the processes of one or more applications into the memory 502 according to the following instructions, and the processor 508 will run the applications stored in the memory 502 to realize various functions:

[0181] Collect the audio data of the user to be tested;

[0182] Extract the voiceprint information of the audio to be detected and match the voiceprint information with preset voiceprint information;

[0183] When the voiceprint information is detected to match the preset voiceprint information, the pitch information set of the audio to be detected is extracted;

[0184] Based on the dispersion of the tone information set, cheating detection is performed on the user to be detected.

[0185] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the cheating detection method above, which will not be repeated here.

[0186] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0187] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the cheating detection methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0188] Collect the audio data of the user to be tested;

[0189] Extract the voiceprint information of the audio to be detected and match the voiceprint information with preset voiceprint information;

[0190] When the voiceprint information is detected to match the preset voiceprint information, the pitch information set of the audio to be detected is extracted;

[0191] Based on the dispersion of the tone information set, cheating detection is performed on the user to be detected.

[0192] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0193] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0194] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0195] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the cheating detection methods provided in the embodiments of this application, the beneficial effects that any of the cheating detection methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0196] The foregoing has provided a detailed description of a cheating detection method, apparatus, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A cheating detection method, characterized in that, The method includes: Collect the audio data of the user to be tested; Extract the voiceprint information of the audio to be detected, and match the voiceprint information with preset voiceprint information; When it is detected that the voiceprint information matches the preset voiceprint information, the pitch information set of the audio to be detected is extracted; Cheating detection is performed on the user to be detected based on the dispersion of the pitch information set. This includes: determining the corresponding dispersion based on the distribution of multiple pitch information items in the pitch information set; comparing the dispersion with a preset dispersion to obtain a dispersion comparison result; wherein the preset dispersion is obtained by statistically analyzing dispersion based on historical audio information; and determining the cheating detection result of the user to be detected based on the dispersion comparison result. Determining the corresponding dispersion based on the distribution of multiple pitch information items in the pitch information set includes: calculating the variance of multiple pitch information items in the pitch information set; wherein the step of comparing the dispersion with the preset dispersion includes: calculating the difference between the variance and the preset variance; and determining the corresponding dispersion comparison result based on the difference.

2. The cheating detection method according to claim 1, characterized in that, The extraction of the pitch information set of the audio to be detected includes: According to the audio time sequence and preset acquisition rules, the pitch information of the audio to be detected is extracted sequentially to obtain a pitch information set.

3. The cheating detection method according to claim 1, characterized in that, The step of determining the cheating detection result of the user to be detected based on the dispersion comparison result includes: When the detected difference is greater than a preset threshold, the dispersed comparison result is determined to be a discrete result, and the user to be detected is determined to be cheating. When the detected difference is not greater than a preset threshold, the dispersed comparison result is determined to be the centralized result, and the user to be tested is determined to be a non-cheating result.

4. The cheating detection method according to any one of claims 1 to 3, characterized in that, The method also includes After confirming that the user being tested has cheated, lock the screen; Start the camera equipment and capture facial images; The screen is unlocked when a face image is detected to match a preset face image.

5. The cheating detection method according to any one of claims 1 to 3, characterized in that, The method further includes: When it is detected that the voiceprint information does not match the preset voiceprint information, the voiceprint information is matched with each historical voiceprint information in the database in turn; Identify the target historical voiceprint information that matches the voiceprint information; Obtain the target user corresponding to the target's historical voiceprint information.

6. A cheating detection device, characterized in that, include: The acquisition unit is used to acquire the audio of the user to be tested; A matching unit is used to extract the voiceprint information of the audio to be detected and match the voiceprint information with preset voiceprint information; The extraction unit is used to extract the pitch information set of the audio to be detected when the voiceprint information is detected to match the preset voiceprint information. A detection unit is configured to perform cheating detection on the user to be detected based on the dispersion of the pitch information set. The cheating detection based on the dispersion of the pitch information set includes: determining a corresponding dispersion based on the distribution of multiple pitch information items in the pitch information set; comparing the dispersion with a preset dispersion to obtain a dispersion comparison result; wherein the preset dispersion is obtained by dispersion statistics based on historical audio information; and determining the cheating detection result of the user to be detected based on the dispersion comparison result. The determination of the corresponding dispersion based on the distribution of multiple pitch information items in the pitch information set includes: calculating the variance value of the multiple pitch information items in the pitch information set; wherein the step of comparing the dispersion with the preset dispersion includes: calculating the difference between the variance value and the preset variance value; and determining the corresponding dispersion comparison result based on the difference.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the cheating detection method according to any one of claims 1 to 5.

8. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the cheating detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Oral exam cheating detection method and device

    CN105931632A

  • Voiceprint recognition method and device

    CN111199729A