Voiceprint recognition method, model construction method and system for authority verification
By using voiceprint recognition systems for random verification and real-time dynamic monitoring in complex environments such as operating rooms, the problem of low accuracy of traditional voiceprint recognition technology in high noise and high pressure environments is solved, and the full monitoring of the identity and status of personnel in the operating scenario is realized, improving safety and standardization.
Patent Information
- Application Number
- CN202510588270.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional voiceprint recognition technology is in a complex dynamic environment with high noise, high pressure, and multi-person collaboration, especially in the operating room environment, with low accuracy, making it difficult to achieve real-time and dynamic monitoring and permission confirmation of the continuous operation process after entering the scene.
The voiceprint recognition system is adopted, including the entrance voiceprint recognition device and the internal voiceprint recognition device, and dynamic verification is achieved through random verification instructions, real-time identity and status acoustic feature extraction, context information acquisition and predefined behavior rules, and ensure the accuracy and security of permission verification.
It improves the accuracy and security of voiceprint recognition, prevents recording and playback attacks, reduces the risk of identity impersonation or abuse of permissions, realizes full monitoring of personnel activities in the operation scenario, and improves operational normativeness and security.
Smart Images

Figure CN120236590A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of user authentication technology, and in particular to a voiceprint recognition method, model building method and system for authority verification. Background Art
[0002] In many scenarios that require strict permission control and operational specifications, such as financial transactions, confidential information access, and especially high-risk environments such as medical operations, it is crucial to ensure the accuracy of the operator's identity, qualification compliance, and behavioral compliance. Traditional identity authentication methods, such as passwords, keys, and IC cards, are prone to loss, theft, or lending.
[0003] In recent years, biometric technology has been widely used due to its uniqueness and difficulty in forgery, including fingerprint recognition, face recognition, iris recognition and voiceprint recognition. Fingerprint recognition and face recognition are currently the more mainstream biometric recognition methods. However, these technologies encounter limitations in certain specific scenarios. For example, in a medical operating room environment, medical staff usually need to wear protective equipment such as masks, gloves and surgical caps, which makes it difficult to fully capture facial features and fingerprint collection becomes inconvenient or unhygienic. In addition, these verification methods are usually used for one-time verification at the entrance, and it is difficult to conduct real-time and dynamic monitoring and permission confirmation of the continuous operation process after entering the scene.
[0004] Voiceprint recognition is a technology that uses acoustic parameters in speech waveforms that reflect the speaker's physiological and behavioral characteristics for identity recognition. It has the advantages of being contactless and easy to collect, and is not easily affected by factors such as wearing masks. Therefore, it is widely used in certain scenarios. Existing voiceprint recognition technology is mainly used for identity confirmation, such as access control systems and remote identity authentication. These systems usually extract voiceprint features and compare them with pre-registered voiceprint models in a relatively quiet environment by having the user read out fixed or random text content to complete identity verification.
[0005] However, in a complex, dynamic, and high-risk environment such as an operating room, it is far from enough to perform a one-time identity verification at the entrance. During the operation, there is a lot of interference from environmental noise (such as the sound of medical equipment running and people walking around), multiple medical staff may speak at the same time or alternately, and the surgical process is complicated. Different stages require personnel with specific qualifications (such as the surgeon) to issue key instructions or perform key operations. There is a risk that even if the entrance verification is passed, the person who actually performs the key operation or issues the key instruction may not be the scheduled or qualified person (for example, an inexperienced assistant performs an operation beyond the surgeon's authority instead of the surgeon). In addition, if an emergency occurs during the operation, the voice characteristics of the personnel may change under high pressure and high load, thus affecting the accuracy of voiceprint recognition.
[0006] Therefore, traditional voiceprint recognition technology has the defect of low accuracy in complex dynamic environments with high noise, high pressure, and multi-person collaboration, especially in the operating room environment. Summary of the Invention
[0007] Embodiments of the present application provide a voiceprint recognition method, a model construction method, and a system for permission verification, which can effectively improve the accuracy of voiceprint recognition in complex dynamic environments with high noise, high pressure, and multi-person collaboration, especially in the operating room environment.
[0008] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0009] In a first aspect, a voiceprint recognition method for permission verification is provided, which is applied to a voiceprint recognition system. The voiceprint recognition system includes an entrance voiceprint recognition device, an internal voiceprint recognition device, and an electronic device. The electronic device stores an identity reference model and a corresponding state baseline model. Each person corresponds to an identity reference model and a state baseline model. The method includes:
[0010] In response to an entry request sent by a person in a fixed entrance area of a predetermined operation scenario, a random verification instruction is sent through the entrance voiceprint recognition device;
[0011] Collect the response voice of the person to the random verification instruction, and extract the entrance identity acoustic features from the response voice;
[0012] Compare the entrance identity acoustic features with the identity reference model to obtain a comparison result, and determine whether to authorize the person to enter the predetermined operation scenario based on the comparison result and a preset entrance verification threshold;
[0013] When authorizing a person to enter the predetermined operation scenario, record the entry status of the person;
[0014] Inside the predetermined operation scenario, continuously capture the environmental sound through the internal voiceprint recognition device, and extract the target voice segment from the environmental sound;
[0015] Extract the real-time identity acoustic features and real-time state acoustic features from the target voice segment;
[0016] Compare the real-time identity acoustic features with the identity reference model to identify the speaker and their role who issued the target voice segment;
[0017] Obtain the real-time context information related to the predetermined operation scenario;
[0018] Determine the corresponding predefined behavior rules according to the identified role and real-time context information;
[0019] Perform dynamic verification on the real-time status acoustic features based on predefined behavior rules to obtain a dynamic verification result;
[0020] Generate output information according to the dynamic verification result.
[0021] In a possible implementation manner of the first aspect, compare the entrance identity acoustic features with the identity reference model to obtain a comparison result, and determine whether to authorize a person to enter a predetermined operation scenario based on the comparison result and a preset entrance verification threshold, including:
[0022] Obtain the identity reference models of all authorized persons associated with the predetermined operation scenario;
[0023] Calculate the similarity scores between the entrance identity acoustic features and each identity reference model;
[0024] Determine the highest score in the similarity scores and the corresponding identity reference model, and judge whether the highest score is greater than or equal to the preset entrance verification threshold;
[0025] When the highest score is greater than or equal to the entrance verification threshold, determine that the person corresponding to the authorized identity reference model is allowed to enter the predetermined operation scenario.
[0026] In another possible implementation manner of the first aspect, extract a target speech segment from the ambient sound, including:
[0027] Collect multi-channel audio signals through the microphone array configured by the internal voiceprint recognition device;
[0028] Apply a sound source localization algorithm to the multi-channel audio signals to determine the direction of the speech source;
[0029] Based on the direction of the speech source, apply a beamforming algorithm to process the multi-channel audio signals to obtain an enhanced single-channel speech signal;
[0030] Perform voice activity detection on the enhanced single-channel speech signal to segment and extract the target speech segment.
[0031] In another possible implementation manner of the first aspect, extract real-time identity acoustic features and real-time status acoustic features from the target speech segment, including:
[0032] Apply a first preset algorithm to extract deep voiceprint features for uniquely identifying the speaker's identity from the target speech segment as real-time identity acoustic features;
[0033] Apply a second preset algorithm set to extract real-time status acoustic features from the target speech segment, where the real-time status acoustic features include at least one of fundamental frequency statistics, jitter value, micro-vibration value, energy feature, and speech rate feature.
[0034] In another possible implementation of the first aspect, the real-time identity acoustic features are compared with the identity reference model to identify the speaker who uttered the target speech segment and their role, including:
[0035] Obtain the identity reference models of all authorized personnel whose current record is in the entry state and the corresponding role information;
[0036] Calculate the similarity score between the real-time identity acoustic features and each identity reference model;
[0037] Determine the highest score in the similarity scores and the corresponding identity reference model;
[0038] Judge whether the highest score is greater than or equal to a preset recognition threshold;
[0039] When the highest score is greater than or equal to the preset recognition threshold, identify the speaker who uttered the target speech segment as the authorized personnel associated with the identity reference model corresponding to the highest score, and determine the role of the speaker.
[0040] In another possible implementation of the first aspect, the predefined behavior rules include the instruction permissions and the expected state range of the role. According to the identified role and the real-time context information, determine the corresponding predefined behavior rules, including:
[0041] Obtain the current stage information of the predetermined operation scenario as the real-time context information;
[0042] Use the identified role and the current stage information as query conditions to retrieve, in the preset database, the set of instruction types authorized for the role under the current stage information; and
[0043] Retrieve the statistical range of the expected acoustic state features of the role under the current stage information;
[0044] Use the set of instruction types as the instruction permissions and the statistical range of the expected acoustic state features as the expected state range.
[0045] In another possible implementation of the first aspect, based on the predefined behavior rules, perform dynamic verification on the real-time state acoustic features to obtain a dynamic verification result, including:
[0046] Compare each feature value in the real-time state acoustic features with the expected state range, and count the number of features that exceed the expected state range;
[0047] Based on the number of features, determine whether the real-time state of the speaker is normal or abnormal, and use the real-time state of the speaker as the dynamic verification result;
[0048] Among them, when the number of features is greater than a preset number threshold, the real-time state of the speaker is determined to be abnormal, and when the number of features is not greater than the preset number threshold, the real-time state of the speaker is determined to be normal.
[0049] In a second aspect, the present application provides a model construction method, which is applied to an electronic device. The electronic device stores an identity reference model and a state baseline model. The construction method of the identity reference model includes:
[0050] Obtain authorized personnel information;
[0051] For each authorized person in the authorized personnel information, collect a voice sample set in a variety of vocalization scenarios;
[0052] Extract the deep voiceprint features of each authorized person from the voice sample set;
[0053] Perform an aggregation process on the deep voiceprint features to generate an identity reference model corresponding to each authorized person;
[0054] The construction method of the state baseline model includes:
[0055] Obtain authorized personnel information;
[0056] For each authorized person in the authorized personnel information, collect a voice sample set in a variety of vocalization scenarios;
[0057] Select voice segments representing the normal working state from the voice sample set;
[0058] Extract a set of state evaluation acoustic features of each authorized person from the voice segments. Among them, the set of state evaluation acoustic features includes one or more of fundamental frequency statistics, jitter value, micro-vibration value, energy feature, and speech rate feature;
[0059] Calculate the statistical distribution parameters of each set of state evaluation acoustic features in the normal working state. The statistical distribution parameters include mean and / or standard deviation;
[0060] Based on the statistical distribution parameters, determine a normal value range for each set of state evaluation acoustic features, and generate a state baseline model corresponding to each authorized person.
[0061] In a third aspect, the present application provides an electronic device, including:
[0062] A memory configured to store instructions; and
[0063] A processor configured to call the instructions from the memory and be able to implement the above model construction method when executing the instructions.
[0064] In a fourth aspect, the present application provides a voiceprint recognition system, including:
[0065] Entry voiceprint recognition device;
[0066] Internal voiceprint recognition device;
[0067] The electronic device is connected to both the entrance voiceprint recognition device and the internal voiceprint recognition device.
[0068] The above technical solutions greatly enhance security compared with the traditional single-point, static identity authentication method. Specifically, the random instructions at the entrance effectively prevent recording playback attacks, while the internal continuous monitoring makes up for the risk of identity fraud or abuse of authority that may occur after a single verification. Secondly, by introducing roles and real-time context information, the authority verification and behavior norm judgment are more refined and intelligent, and the rules can be dynamically adjusted according to the actual situation, which improves the flexibility and accuracy of management. The addition of dynamic verification of the speaker's real-time state acoustic characteristics can timely detect operational risks that may be caused by abnormal states such as fatigue and tension, providing an additional layer of security for high-risk environments. In addition, the use of microphone arrays, sound source localization and beamforming technologies improves the ability to extract and recognize speech in complex noise and multi-person speaking environments, and ultimately achieves full monitoring of personnel activities in the predetermined operation scenario. It not only improves the standardization and safety of operations, reduces the possibility of human errors or violations, but also provides detailed acoustic evidence and status records for post-event tracing and auditing, which has important application value for scenarios that require high trust and strict process control.
[0069] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A flow chart of a voiceprint recognition method for authority verification provided in an embodiment of the present application;
[0071] Figure 2 A schematic diagram of the structure of a voiceprint recognition system provided in an embodiment of the present application;
[0072] Figure 3 A flowchart of a model building method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of this application, and are not used to limit the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0074] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of this application, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0075] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0076] The application scenario of the embodiments of this application is a complex dynamic environment with high noise, high pressure, and multi-person collaboration. The following takes a hospital operating room as an example for illustration.
[0077] In the embodiments of the present application, in order to accurately bind the medical staff participating in the operation to their unique acoustic characteristics and construct a reference model for subsequent real-time comparison. In principle, first, an authorization source is required to confirm which personnel are eligible to participate in a specific operation or enter a specific operation area. This can be achieved through the operation scheduling system or the operation application and approval process in the electronic medical record system of the hospital, which includes the list of participating personnel and their roles (surgeon, assistant, anesthesiologist, nurse, etc.) confirmed by multiple parties such as the surgeon and the patient. After obtaining the authorized list, it is necessary to collect voiceprint data for each authorized person. Different from traditional voiceprint recognition that only focuses on identity discrimination, this solution not only needs to collect voiceprint features that can uniquely identify identity, but also needs to collect acoustic features that can reflect their physiological and psychological states such as emotions and stress. Therefore, during the collection process, in a relatively quiet environment, the user needs to be guided to read the specified text (text-related), and free conversations (text-unrelated) are also included. More importantly, different scenarios need to be simulated, such as normal speech rate, fast directive speech rate, and vocalization under mild stress (such as answering timed questions), to capture a wider range of acoustic variations.
[0078] After the collected raw audio data is preprocessed (such as silence excision, normalization), two types of features are extracted: one is the deep voiceprint feature for identity recognition, such as x-vector, which maps variable-length speech segments to fixed-dimensional vectors through a deep neural network and can effectively distinguish different speakers; the other is the set of acoustic features for state evaluation, including fundamental frequency (F0) and its statistics (mean, standard deviation, range), jitter (tiny changes in the fundamental frequency period), shimmer (tiny changes in amplitude), short-time energy, speech rate, formant frequency and other prosody and voice quality parameters. Using the extracted features, two models are constructed for each authorized person: one is the identity recognition model, and the other is the state baseline model, which defines the statistical distribution range (such as mean, standard deviation, normal value interval) of each state acoustic feature of the person in the "normal" working state. All models and personnel identity information (such as employee number), role information are securely stored in a dedicated voiceprint database, and data security is ensured through encryption and access control.
[0079] Specifically, Figure 1 Schematically shows a flowchart of a voiceprint recognition method for permission verification according to an embodiment of the present application. As Figure 1 shown, the embodiments of the present application provide a voiceprint recognition method for permission verification, which is applied to a voiceprint recognition system. The voiceprint recognition system includes an entrance voiceprint recognition device, an internal voiceprint recognition device, and an electronic device. The electronic device stores an identity reference model and a corresponding state baseline model. Each person corresponds to an identity reference model and a state baseline model. The method may include the following steps.
[0080] S101. In response to an entry request issued by a person in a fixed entry area of a predetermined operation scenario, send a random verification instruction through an entry voiceprint recognition device;
[0081] S102. Collect the response voice of the person to the random verification instruction, and extract the entry identity acoustic features from the response voice;
[0082] S103. Compare the entry identity acoustic features with an identity reference model to obtain a comparison result, and based on the comparison result and a preset entry verification threshold, determine whether to authorize the person to enter the predetermined operation scenario;
[0083] S104. When authorizing a person to enter the predetermined operation scenario, record the entry status of the person;
[0084] S105. Inside the predetermined operation scenario, continuously capture the ambient sound through an internal voiceprint recognition device, and extract a target voice segment from the ambient sound;
[0085] S106. Extract real-time identity acoustic features and real-time status acoustic features from the target voice segment;
[0086] S107. Compare the real-time identity acoustic features with the identity reference model to identify the speaker who issued the target voice segment and their role;
[0087] S108. Obtain real-time context information related to the predetermined operation scenario;
[0088] S109. Determine the corresponding predefined behavior rules according to the identified role and real-time context information;
[0089] S110. Based on the predefined behavior rules, perform dynamic verification on the real-time status acoustic features to obtain a dynamic verification result;
[0090] S111. Generate output information according to the dynamic verification result.
[0091] In this embodiment, the electronic device may be a device with a processor such as a tablet computer, desktop, laptop, handheld computer, wearable device, notebook computer, ultra-mobile personal computer (UMPC), netbook, etc. Of course, the electronic device may also be a server. The specific form of the electronic device in this application embodiment is not particularly limited.
[0092] Refer to Figure 2The voiceprint recognition system of an embodiment of the present application may include an entrance voiceprint recognition device deployed at the entrance of a predetermined operating scene (for example, an operating room), at least one internal voiceprint recognition device deployed inside the predetermined operating scene, and an electronic device. The electronic device is responsible for storing and processing data, including a pre-registered identity reference model and a corresponding state baseline model for each person. Each person (such as a doctor, nurse, anesthesiologist, intern, etc.) corresponds to a unique identity reference model and a state baseline model. The entrance voiceprint recognition device typically includes a microphone and a speaker for entrance verification interaction. The internal voiceprint recognition device typically includes a microphone array for continuously capturing internal sounds.
[0093] The identity reference model is a set of acoustic features that can uniquely identify the person based on a large number of voice samples extracted from the person. The state baseline model describes the typical range of acoustic parameters when the person speaks in normal working conditions, such as normal speaking speed, pitch, and smoothness of the voice. Together, these two models constitute a comprehensive description of the voice of the authorized person for subsequent identity confirmation and state assessment. The predetermined operation scenario can refer to any physical space or logical environment that requires strict authority control and process monitoring, such as an operating room, a financial trading room, a data center computer room, a confidential laboratory, etc.
[0094] In a fixed entrance area of a predetermined operation scenario, when a person issues an entry request (such as by pressing an access button, touching a screen, or triggering a proximity sensor), the entrance voiceprint recognition device will immediately respond and issue a random verification instruction. The random verification instruction is a voice prompt randomly selected from a pre-set instruction library, which can be a combination of numbers (such as "Please read 1-3-5-7"), a phrase (such as "Please say what day of the week it is today"), or a specific sentence (such as "Please say your name and position").
[0095] After a person makes a voice response to the random verification instruction, the entrance voiceprint recognition device collects this response voice through a built-in microphone. During the collection process, the gain control can be automatically adjusted to adapt to the volume differences of different people and environmental noise conditions. After the collection is completed, the voice signal is first preprocessed, including steps such as noise suppression, endpoint detection, and signal enhancement, to improve the accuracy of subsequent feature extraction. Then, the entrance identity acoustic features are extracted from the preprocessed response voice. Feature extraction uses deep neural network models, such as x-vector or d-vector, etc., to extract high-dimensional feature vectors with a dimension of 512 from the voice signal, so as to effectively capture individual differences such as the vocal tract characteristics and vocalization habits of the speaker. During the feature extraction process, special attention is paid to voice features that are not easily imitated and remain stable under different recording conditions, such as the distribution pattern of vocal tract formants and harmonic structure, etc. This is because the above features are closely related to the physiological structure of the speaker and have high distinctiveness. In this way, an acoustic feature representation that can uniquely identify the current speaker, that is, the entrance identity acoustic features, can be obtained.
[0096] The identity reference models of all authorized personnel associated with the predetermined operation scenario are pre-established during the initialization or personnel registration phase and stored in the secure database of the electronic device. By calculating the similarity score between the entrance identity acoustic features and each identity reference model, a match can be made with the personnel. Similarity calculation uses algorithms such as cosine similarity.
[0097] After the calculation is completed, the highest score in the similarity scores and the corresponding identity reference model are determined, and the highest score is compared with the preset entrance verification threshold. The entrance verification threshold is preset, usually between 0.7 - 0.9, and the higher the value, the stricter the verification. If the highest score is greater than or equal to the entrance verification threshold, it is determined that the verification is passed, and it is confirmed that the personnel identity matches the identity reference model corresponding to the highest score, and the person is authorized to enter the predetermined operation scenario. If the score is lower than the threshold, the entrance request is rejected, and a security alarm may be triggered or other verification methods may be required. In actual applications, the verification threshold will be dynamically adjusted according to different operation scenarios and security levels. For example, for high-risk operation areas, a higher threshold may be set to require a stricter matching degree.
[0098] After an authorized person enters the predetermined operation scenario, the entry status of the person is recorded in the database, including key data such as personnel identity information, entry time, verification score, etc. The recording of the entry status is managed using a state machine model, and each authorized person has a clear status identifier in it, such as "not entered", "entered", "left", etc.
[0099] When a person passes the entrance verification, their status changes from "not entered" to "entered", triggering the corresponding permission activation process. According to the person's role and permission level, the scope of operations that can be performed within the scenario can be automatically configured. At the same time, information such as the person's expected operation duration and expected activity area is recorded, providing a reference basis for subsequent behavior monitoring. The record of the entry status is not only used for access control but also bound to various devices and permissions within the scenario to ensure that only the currently present and authorized personnel can operate specific devices or access sensitive information.
[0100] Inside the predetermined operation scenario, the internal voiceprint recognition device continuously captures ambient sounds through a distributed microphone array. The microphone array consists of at least 4 high-sensitivity microphones and can achieve 360-degree omnidirectional sound collection. First, apply a sound source localization algorithm, such as the TDOA (Time Difference of Arrival) or GCC-PHAT (Generalized Cross-Correlation Phase Transform) algorithm, to the collected multi-channel audio signals to accurately calculate the sound source direction.
[0101] After determining the voice source direction, apply a beamforming algorithm (such as the MVDR algorithm) to process the multi-channel audio signals. By aligning the phases and adjusting the gains, enhance the sound from the target direction while suppressing the noise and interference from other directions, obtaining a single-channel voice signal with a significantly improved signal-to-noise ratio.
[0102] Perform voice activity detection (VAD) on the enhanced single-channel voice signal to segment and extract the target voice segment.
[0103] From the extracted target voice segment, two types of feature extractions are performed in parallel: real-time identity acoustic features and real-time status acoustic features. For real-time identity acoustic features, apply a first preset algorithm, which is a deep learning-based feature extraction method, such as the ResNet or ECAPA-TDNN network architecture, to extract the deep voiceprint features used to uniquely identify the speaker's identity from the target voice segment. Specifically, the real-time identity acoustic features are high-dimensional vectors of 512 or 1024 dimensions, capable of capturing the subtle differences in the speaker's vocal tract structure and pronunciation habits.
[0104] Meanwhile, apply the second preset algorithm set to extract real-time state acoustic features from the target speech segment. These features include: (1) Fundamental frequency statistics: Extract the fundamental frequency contour of the speech through autocorrelation method or spectral analysis method, and calculate statistics such as its mean, standard deviation, range, and change rate; (2) Jitter value: Quantify the short-term variation of the vocal cord vibration period, and the calculation method is the average absolute value of the difference between consecutive fundamental frequency periods; (3) Shimmer value: Measure the mid-term modulation (4 - 12 Hz) of the fundamental frequency, reflecting the stability of vocal cord muscle control; (4) Energy features: Include overall energy level, energy change rate, and frequency band energy distribution, etc.; (5) Speech rate feature: Calculate the pronunciation rate per unit time through syllable or phoneme recognition. The above state acoustic features are highly sensitive to the physiological and psychological states of the speaker and can reflect state changes such as fatigue, stress, and tension. The feature extraction process uses a sliding window technique, calculates once every 20 - 30 milliseconds, and obtains the feature representation of the entire speech segment through statistical aggregation to ensure capturing the dynamic changes of the speaker's state.
[0105] Obtain the identity reference models of all authorized personnel whose current record is in the entry state and the corresponding role information, and calculate the similarity scores between the real-time identity acoustic features and each identity reference model, using the same similarity calculation method as the entry verification. Then, determine the highest score in the similarity scores and the corresponding identity reference model, and judge whether the highest score is greater than or equal to the preset recognition threshold. The recognition threshold is usually set between 0.65 - 0.85, slightly lower than the entry verification threshold, to adapt to possible noise interference and changes in speaking styles within the scenario.
[0106] If the highest score is greater than or equal to the preset recognition threshold, identify the speaker of the target speech segment as the authorized personnel associated with the identity reference model corresponding to the highest score. Meanwhile, retrieve the role information of this authorized personnel from the database, such as "surgeon", "anesthesiologist", "circulating nurse", etc. The role information is associated with specific operation permissions and responsibility scopes and is an important basis for subsequent permission verification. In a multi-person collaboration scenario, multiple speakers can be tracked and identified simultaneously, and a dynamically updated scenario personnel status map is maintained to record the position, identity, and recent activities of each person. When a new voice input is detected, the identity recognition will be quickly completed, and the result will be integrated with the scenario status map to ensure continuous and accurate monitoring of each person within the scenario.
[0107] Obtain real-time context information related to a predetermined operation scenario to understand the current operation environment and judge the compliance of behaviors. The acquisition of context information adopts a multi-source fusion method, including: time information, which can be determined by recording the current precise time point and combining it with the time schedule of the predetermined operation scenario to determine which operation stage or time window it is currently in; spatial information, which can be determined by sound source localization or integration with other positioning methods to determine the specific position of the speaker in the scenario; operation status information, which is obtained by exchanging data with other devices (such as medical devices, consoles) in the scenario to obtain the current device operation status, executed operation steps, etc.; environmental parameters, which can be achieved by monitoring environmental factors such as noise level and personnel density. Particularly important is to obtain the current stage information of the predetermined operation scenario as the core real-time context information. For example, in the operating room scenario, it will identify whether it is currently in the "anesthesia preparation stage", "surgical operation stage", or "recovery stage". Stage identification can be achieved based on a preset time schedule, key event detection, or direct integration with surgical management. Different stages correspond to different permission configurations and behavior specifications, which directly affect the verification decision. Adopt a context-aware computing model to integrate various types of information obtained into a structured context representation, providing a comprehensive decision-making basis for subsequent behavior rule matching.
[0108] Based on the identified role and current stage information, perform a retrieval operation. First, use the identified role and current stage information as query conditions to retrieve in a preset database the set of instruction types authorized for the role under the current stage information. For example, for the "surgeon" role in the "surgical operation stage", the allowed instruction types may include "surgical instrument request", etc. Second, retrieve the statistical range of the expected acoustic state characteristics of the role under the current stage information. These ranges are predefined based on historical data and describe the acoustic feature distribution of a specific role in a specific stage under normal working conditions. For example, the fundamental frequency range, speech rate change, and energy level of the surgeon during critical surgical steps.
[0109] After the retrieval is completed, use the set of instruction types as instruction permissions and the statistical range of the expected acoustic state characteristics as the expected state range. The two together constitute the predefined behavior rules. These rules not only consider static permission allocation but also include dynamic behavior specifications, forming a comprehensive behavior evaluation framework. In this embodiment, a rule-based inference engine can be used to manage the above behavior rules and can dynamically adjust the application priority of the rules according to changes in the scenario.
[0110] Compare each feature value in the real-time state acoustic characteristics with the expected state range one by one to check whether there is a situation beyond the range. The comparison process uses the Z-score normalization method to convert each feature value into a standard deviation unit relative to the expected distribution:
[0111]
[0112] Among them, x is the real-time eigenvalue, and μ and σ are the mean and standard deviation of the expected range respectively. If |Z| > 2 (that is, the eigenvalue exceeds the expected range by about 95% confidence interval), then the feature is considered abnormal. Count the number of features that exceed the expected state range, and determine the real-time state of the speaker based on this number.
[0113] The specific judgment criterion is as follows: If the number of abnormal features is greater than the preset number threshold (usually set to 20% - 30% of the total number of features), determine the real-time state of the speaker as abnormal; otherwise, determine it as normal. In practical applications, different weights will be assigned to different features to reflect the differences in their importance in state judgment. Through this evaluation method, it is possible to accurately distinguish normal individual differences and true abnormal states, output the real-time state of the speaker as a dynamic verification result, and provide a key basis for subsequent decision-making.
[0114] According to the dynamic verification result, generate corresponding output information to achieve effective control of the operation scenario. When the dynamic verification result is normal, operation authorization information will be generated according to the identified role and instruction content, allowing relevant devices to perform corresponding operations. For example, confirm the "start surgery" instruction issued by the surgeon in a normal state and trigger the corresponding surgical equipment startup process. When the dynamic verification result is "abnormal", different levels of responses will be generated according to the degree and type of abnormality: for mild abnormalities, generate warning information to remind other on-site personnel to pay attention to the state of the speaker; for moderate abnormalities, require secondary confirmation or additional authorization; for severe abnormalities, the operation can be paused or an emergency intervention process can be triggered.
[0115] The generation of output information supports multiple output forms, including visual cues (such as status indicators on the display screen), audio cues (such as warning sounds or voice announcements), and control signals (such as device enable / disable commands).
[0116] In summary, this embodiment not only performs identity verification at the entrance, but also continuously monitors and verifies the identity and status of the operator inside the scenario, effectively preventing the risk of unauthorized personnel performing critical operations. By combining voiceprint recognition technology and status monitoring technology, it is possible to detect abnormal states of the operator (such as fatigue, nervousness, or excessive stress), intervene in potential risky operations in a timely manner, and prevent human errors. The context-aware feature makes the verification process more intelligent, capable of dynamically adjusting the permission configuration according to different operation stages and roles. In addition, there is no need for the operator to wear additional devices or perform specific actions, which does not interfere with the normal work process, and at the same time provides the ability of full-process operation recording and audit tracking.
[0117] Compared with the traditional single-point, static identity authentication method, this embodiment greatly enhances security. Specifically, the random instructions at the entrance effectively prevent the recording playback attack, and the internal continuous monitoring makes up for the risk of identity fraud or abuse of authority that may occur after a single verification. Secondly, by introducing roles and real-time context information, the authority verification and behavior norm judgment are more refined and intelligent, and the rules can be dynamically adjusted according to the actual situation, which improves the flexibility and accuracy of management. The addition of dynamic verification of the speaker's real-time state acoustic characteristics can timely discover the operational risks that may be caused by abnormal states such as fatigue and tension of personnel, providing an additional layer of security for high-risk environments. In addition, the use of microphone arrays, sound source localization and beamforming and other technologies improves the ability of voice extraction and recognition in complex noise and multi-person speaking environments, and finally realizes the full monitoring of personnel activities in the predetermined operation scene. It not only improves the standardization and safety of operations, reduces the possibility of human errors or violations, but also provides detailed acoustic evidence and status records for post-event tracing and auditing, which has important application value for scenarios that require high trust and strict process control.
[0118] In one implementation of this embodiment, the entrance identity acoustic feature is compared with the identity reference model to obtain a comparison result, and based on the comparison result and a preset entrance verification threshold, it is determined whether the authorized person enters the predetermined operation scene, including the following steps:
[0119] S201, obtaining identity reference models of all authorized personnel associated with a predetermined operation scenario;
[0120] S202, calculating a similarity score between the entry identity acoustic feature and each identity reference model;
[0121] S203, determining the highest score among the similarity scores and the corresponding identity reference model, and judging whether the highest score is greater than or equal to a preset entry verification threshold;
[0122] S204: When the highest score is greater than or equal to the entry verification threshold, determine whether the person corresponding to the identity reference model is authorized to enter the predetermined operation scenario.
[0123] Before performing authentication, it is first necessary to obtain the identity reference models of all authorized personnel associated with the predetermined operation scenario. The identity reference models are stored in the security database of the electronic device and are pre-constructed during the initialization or personnel registration phase. Specifically, during implementation, first query the scenario-person mapping table according to the identifier of the predetermined operation scenario (such as "Operating Room A") to obtain the list of all authorized personnel IDs associated with this scenario. According to the obtained list of personnel IDs, retrieve the corresponding identity reference models from the identity model database. Each identity reference model is a high-dimensional feature vector or a set of feature vectors, usually with dimensions between 512 and 1024, representing the voiceprint feature distribution of this authorized personnel.
[0124] After obtaining all relevant identity reference models, calculate the similarity scores between the entrance identity acoustic features and each identity reference model. Similarity calculation is the core link of voiceprint recognition and directly affects the recognition accuracy. In actual implementation, it is achieved by calculating the angular similarity between feature vectors using cosine similarity.
[0125] After completing the calculation of all similarity scores, determine the highest score and its corresponding identity reference model, and make a verification decision. First, sort all the calculated similarity scores to find the highest score value and its corresponding identity reference model. In practical applications, not only the highest score is concerned, but also the gap (score interval) between the highest score and the second-highest score is calculated as an auxiliary basis for judgment. A larger score interval indicates a higher credibility of the recognition result. Then, normalize the highest score so that it is within the range of [-1, 1], and compare the highest score with the preset entrance verification threshold. The entrance verification threshold is usually between 0.7 and 0.9, and the specific value is dynamically set according to the security requirements of the operation scenario.
[0126] When it is determined that the highest similarity score is greater than or equal to the preset entrance verification threshold, it is confirmed that the identity of this person matches the identity reference model corresponding to the highest score, and the person is authorized to enter the predetermined operation scenario. The authorization process includes multiple key steps. First is identity confirmation, which finally checks the recognition result against the records in the authorized personnel database to ensure that this person indeed has the permission to enter the current operation scenario. And check the permission status of this person to confirm whether their authorization is within the validity period and whether there are any temporary restrictions or special conditions. After confirmation, generate an authorization signal, which is reflected in the form of unlocking an electronic lock, releasing access control, activating an access card, etc.
[0127] At the same time, record detailed authorization information in the security log, including personnel identity, verification score, authorization time, entrance location, etc., to provide a complete record for subsequent security audits. At the same time as the authorization is completed, feedback will be provided to the authorized personnel.
[0128] This embodiment provides a good user experience while ensuring high security through advanced voiceprint feature extraction and matching technologies, combined with a dynamic threshold adjustment strategy. Compared with traditional identity verification methods, voiceprint recognition is applicable to scenarios where quick verification is required and the hands may be occupied (such as medical staff). The multi-model comparison mechanism of this method ensures the accuracy of recognition, and can quickly locate the correct identity even in complex scenarios with a large number of authorized users. Through threshold control and security policy configuration, the verification strictness can be flexibly adjusted according to the security requirements of different scenarios, achieving the best balance between security and convenience.
[0129] In one implementation of this embodiment, extracting a target voice segment from ambient sound includes the following steps:
[0130] S301: Collect multi-channel audio signals through the microphone array configured by the internal voiceprint recognition device;
[0131] S302: Apply a sound source localization algorithm to the multi-channel audio signals to determine the direction of the voice source;
[0132] S303: Based on the direction of the voice source, apply a beamforming algorithm to process the multi-channel audio signals to obtain an enhanced single-channel voice signal;
[0133] S304: Perform voice activity detection on the enhanced single-channel voice signal to segment and extract the target voice segment.
[0134] Inside a predetermined operation scenario, multi-channel audio signals are collected through the microphone array configured by the internal voiceprint recognition device. The microphone array consists of multiple high-precision microphones, usually arranged in a circular or linear pattern to achieve omnidirectional sound collection.
[0135] In actual implementation, the microphone array can include 4 to 8 omnidirectional microphones, with a spacing usually of 5 - 15 cm. This configuration can avoid spatial aliasing effects while ensuring spatial resolution. Each microphone is equipped with a high-quality preamplifier and an analog-to-digital converter, with a sampling rate set to 16 kHz or higher and a quantization precision of 24 bits to capture the full spectral details of the voice.
[0136] To cope with different operating environments, the microphone array adopts adaptive gain control technology, which can automatically adjust the sensitivity according to the ambient noise level to ensure clear voice signals can still be obtained in noisy environments. During the signal acquisition process, the precise timestamps of each microphone are recorded synchronously. Multichannel signal acquisition uses time-division multiplexing or parallel acquisition architectures to ensure the synchronization of signals in each channel, and the time synchronization error is controlled at the microsecond level. To reduce electromagnetic interference and improve the signal-to-noise ratio, the microphone array uses shielded cables and differential signal transmission, and a low-noise filtering circuit is added to the signal path. Finally, multichannel audio signals are acquired.
[0137] After obtaining the multichannel audio signals, a sound source localization algorithm is applied to these signals to accurately determine the direction of the voice source. In actual implementation, a method based on time delay estimation is adopted, such as the Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm, to calculate the time difference of arrival of sound between different microphone pairs. The GCC-PHAT algorithm is implemented through the following steps: perform a short-time Fourier transform on each pair of microphone signals, calculate the cross-power spectrum, apply phase transform weighting, then obtain the cross-correlation function through inverse transform, and finally find the peak position of the cross-correlation function to determine the time delay. The mathematical expression is:
[0138]
[0139] where, X i (f) and are the signal spectra of microphones i and j respectively, and τ is the time delay. After obtaining the time differences of arrival for multiple pairs of microphones, hyperbolic positioning or the least squares method is applied to solve for the sound source position.
[0140] After determining the direction of the voice source, based on this direction information, a beamforming algorithm is applied to process the multichannel audio signals to obtain an enhanced single-channel voice signal. Beamforming is a spatial filtering technology that forms an acoustic focus pointing in a specific direction by adjusting the phase and gain of each microphone signal, enhancing the sound in the target direction while suppressing interference and noise in other directions. In actual implementation, an adaptive beamforming algorithm is adopted, which can dynamically adjust the parameters according to changes in the acoustic environment. The most basic Delay-and-Sum Beamforming (DS) is achieved by aligning the time delays of each channel signal and then summing them. Its mathematical expression is:
[0141]
[0142] where, y(t) is the output signal, x i (t) is the signal of the i-th microphone, τ i is the time delay calculated according to the target direction, and M is the number of microphones.
[0143] Perform voice activity detection (VAD) on the enhanced single-channel voice signal. VAD aims to accurately distinguish voice segments from non-voice segments (such as background noise, silence, or non-verbal sounds), providing a pure voice input for subsequent voiceprint analysis. In actual implementation, the enhanced voice signal is processed frame by frame. Each frame is typically 20 - 30 milliseconds, and the frame shift is 10 - 15 milliseconds to ensure smooth feature extraction. For each frame, multiple acoustic features are calculated, including: short-time energy (STE), zero-crossing rate (ZCR), spectral entropy, Mel-frequency cepstral coefficients (MFCC), spectral flux (SF), etc. These features describe the difference between voice and non-voice from different perspectives. For example, voice segments usually have higher energy, lower zero-crossing rate, and more complex spectral structures. A deep learning-based VAD model, such as a convolutional neural network (CNN) or a long short-term memory network (LSTM), is used to take these features as inputs and output the probability that each frame belongs to voice. After detecting consecutive voice frames, an endpoint refinement algorithm is applied to accurately locate the start and end points of the voice, usually including a buffer of 50 - 100 milliseconds before and after the voice to ensure that the start and end parts of the voice are not lost.
[0144] For special scenarios, such as a conference environment where multiple people take turns speaking, speaker segmentation technology is also combined to separately process the voice segments of different speakers. Finally, according to the VAD results, the target voice segments are segmented and extracted from the original signal, forming a series of voice segments with clearly defined time markers. Each segment contains the continuous voice content of a single speaker. Through this precise voice activity detection, high-quality target voice segments can be extracted from the continuous audio stream.
[0145] This embodiment realizes the acquisition and processing of high-quality voice signals in complex and noisy environments, laying a solid foundation for subsequent voiceprint recognition and state analysis. The stereo field information is collected through a multi-channel microphone array, and the speaker position is accurately determined by combining a sound source localization algorithm, effectively solving the performance limitations of traditional single microphones in multi-person speaking and high-noise environments. The adaptive beamforming technology realizes acoustic focusing on the target speaker, significantly improving the signal-to-noise ratio of the voice signal. Even in complex environments where medical equipment is operating and multiple people are moving simultaneously, clear voice signals can be extracted. The voice activity detection algorithm based on multi-feature fusion accurately segments the effective voice segments, avoiding misprocessing of non-voice signals, and improving the processing efficiency and recognition accuracy. The entire processing flow is completely automated without manual intervention, capable of real-time processing of continuous audio streams and supporting continuous monitoring of the operation scenario.
[0146] In one implementation of this embodiment, real-time identity acoustic features and real-time state acoustic features are extracted from the target voice segment, including the following steps:
[0147] S401. Apply a first preset algorithm to extract deep voiceprint features for uniquely identifying the speaker's identity from the target speech segment as real-time identity acoustic features;
[0148] S402. Apply a second preset algorithm set to extract real-time state acoustic features from the target speech segment, where the real-time state acoustic features include at least one of fundamental frequency statistics, jitter value, micro-vibration value, energy feature, and speech rate feature.
[0149] From the extracted target speech segment, apply a first preset algorithm to extract deep voiceprint features for uniquely identifying the speaker's identity as real-time identity acoustic features. The first preset algorithm is based on a deep learning architecture and uses a neural network model to extract high-dimensional and robust voiceprint representations from speech signals.
[0150] In actual implementation, the algorithm first preprocesses the target speech segment, including pre-emphasis, framing, and windowing operations. Pre-emphasis is achieved through a first-order high-pass filter with a transfer function of H(z) = 1 - αz (-1) , where α is taken as 0.97, aiming to compensate for the natural attenuation of the high-frequency part of the speech. Framing divides the continuous speech signal into short frames of 25 milliseconds with a frame shift of 10 milliseconds to ensure sufficient overlap between adjacent frames. Each frame is applied with a Hamming window function to reduce the spectral leakage effect. After preprocessing, the algorithm extracts acoustic features from each frame, such as Mel-frequency cepstral coefficients (MFCC) or filter bank energy features (FBANK).
[0151] Among them, the calculation process of the filter bank energy feature includes: performing a fast Fourier transform on each frame of the signal, calculating the power spectrum, and applying a Mel filter bank (usually 40 triangular filters) to obtain the filter output energy. These low-level features are used as the input of the deep neural network. The deep neural network adopts a residual network (ResNet) or a time-delay neural network (TDNN) architecture, which includes multiple convolutional layers, pooling layers, and fully connected layers. It is pre-trained with a large-scale speech dataset to learn and extract abstract features related to the speaker's identity. In the feature extraction stage, the speech signal propagates forward through the network, and the output of the last hidden layer (usually a 512- or 1024-dimensional vector) is the deep voiceprint feature.
[0152] Meanwhile, applying the second preset algorithm set, real-time state acoustic features are extracted from the target speech segment, which can reflect the physiological and psychological state changes of the speaker. The second preset algorithm set contains multiple specially designed acoustic feature extraction algorithms, each targeting different types of state indicators. First is the extraction of fundamental frequency statistics. The fundamental frequency (F0) reflects the vocal cord vibration frequency and is closely related to the emotional state. The algorithm combines the autocorrelation method (ACF) and the spectral peak picking method to extract the fundamental frequency contour. The basic principle of the autocorrelation method is to calculate the similarity between the signal and its time-shifted version, and the fundamental frequency period corresponds to the first main peak of the autocorrelation function. The mathematical expression is:
[0153]
[0154] where x(n) is the speech signal, τ is the time shift amount, and N is the frame length. After extracting the fundamental frequency contour, its statistics are calculated, including the mean (reflecting the overall pitch level), standard deviation (reflecting the pitch variation range), variation rate (reflecting the pitch change speed), and distribution shape parameters (such as skewness and kurtosis).
[0155] Next, the jitter value is extracted to quantify the short-term variation of adjacent fundamental frequency periods. The calculation formula is:
[0156]
[0157] where T i is the duration of the i-th fundamental frequency period, and N is the total number of periods. A high jitter value usually indicates unstable vocal cord control and is used to reflect a tense or fatigued state. The shimmer value measures the amplitude change of adjacent fundamental frequency periods. The calculation formula is similar, but the period duration is replaced by the period amplitude. The energy features include the short-time energy mean, variation rate, and frequency band energy distribution, which reflect the speaker's vitality level and emotional intensity. The speech rate features are calculated through syllable or phoneme recognition techniques, including the pronunciation rate (number of syllables per second), pause frequency, and duration distribution. In addition, the algorithm also extracts auxiliary features such as the harmonic-to-noise ratio and spectral tilt to comprehensively characterize the acoustic properties of the sound. All the above features use the sliding window technique, which is calculated every 20 - 30 milliseconds, and the feature representation of the entire speech segment is obtained through statistical aggregation. Through this multi-dimensional extraction of state acoustic features, the subtle changes in the speaker's state can be comprehensively captured, providing rich acoustic evidence for subsequent state assessment.
[0158] This embodiment realizes a comprehensive acoustic characterization of the speaker's identity and state. By using deep learning technology to extract high-dimensional and abstract voiceprint features, it can accurately capture the speaker's vocal tract characteristics and pronunciation habits, and still maintain good recognition performance even under different speech contents, emotional states, and recording conditions. At the same time, by extracting multi-dimensional state acoustic features, it can sensitively detect subtle changes in the speaker's state, such as fatigue, stress, tension, etc. The two feature extraction processes are executed in parallel, realizing the simultaneous progress of identity verification and state monitoring, without additional speech acquisition or processing steps, improving the real-time performance and efficiency, realizing double protection of the operator's identity and state, and providing more comprehensive and in-depth technical support for the safety control of high-risk scenarios.
[0159] In one implementation of this embodiment, the real-time identity acoustic features are compared with the identity reference model to identify the speaker who emits the target speech segment and their role, including the following steps:
[0160] S501. Obtain the identity reference models of all authorized personnel whose current record is in the entry state and the corresponding role information;
[0161] S502. Calculate the similarity scores between the real-time identity acoustic features and each identity reference model;
[0162] S503. Determine the highest score among the similarity scores and the corresponding identity reference model;
[0163] S504. Judge whether the highest score is greater than or equal to the preset recognition threshold;
[0164] S505. When the highest score is greater than or equal to the preset recognition threshold, identify the speaker who emits the target speech segment as the authorized personnel associated with the identity reference model corresponding to the highest score, and determine the role of the speaker.
[0165] Before performing real-time identity recognition, it is first necessary to obtain the identity reference models of all authorized personnel whose current record is in the entry state and the corresponding role information, which is realized by querying the real-time status database that continuously tracks and updates the entry and exit status of all personnel.
[0166] In specific implementation, first access the personnel status table and use a query statement to filter out all authorized personnel records with a status flag of "entered" and not marked as "left". The records contain information such as the unique identifier of the personnel, the entry timestamp, and the last activity time. According to the obtained list of personnel identifiers, retrieve the corresponding identity reference models from the identity model database. At the same time, retrieve the role information of these personnel from the role permission database, such as "surgeon", "anesthesiologist", "circulating nurse", etc. The role information is associated with specific operation permissions and responsibility scopes and is an important basis for subsequent permission verification. To improve the retrieval efficiency, a caching mechanism is adopted to load frequently accessed models and role information into memory.
[0167] After obtaining all relevant identity reference models, calculate the similarity score between the real-time identity acoustic features and each identity reference model. Similarity calculation is the core link of voiceprint recognition and directly affects the accuracy of recognition. In actual implementation, the cosine similarity is used to calculate the angular similarity between feature vectors as the similarity score.
[0168] The value range of the cosine similarity is [-1, 1], and the closer the value is to 1, the more similar the directions of the two vectors are.
[0169] For each identity reference model, if the model contains multiple feature vectors, the similarity between the real-time feature and each reference feature will be calculated, and then the final score will be obtained through a weighted average or maximum selection strategy. To cope with possible noise interference and changes in speaking styles in the scenario, the features will be normalized before calculating the similarity, such as length normalization and adaptive feature normalization.
[0170] After completing the calculation of all similarity scores, it is necessary to determine the highest score and its corresponding identity reference model as the basis for identity recognition. First, sort all the calculated similarity scores to find the highest score value and its corresponding identity reference model. In actual applications, not only the highest score is concerned, but also the gap (score interval) between the highest score and the second-highest score will be calculated as an auxiliary basis for judgment. A larger score interval indicates a higher credibility of the recognition result, while a smaller interval may indicate a risk of identity confusion. To improve the reliability of the recognition, the highest score will be normalized so that it is comparable under different speakers and different environmental conditions. The normalization method can be Z-score normalization.
[0171] Compare the determined highest similarity score with a preset recognition threshold, where the recognition threshold is set between 0.65 - 0.85 (for scores normalized to the [-1, 1] range).
[0172] When it is determined that the highest similarity score is greater than or equal to the preset recognition threshold, the speaker of the target voice segment is identified as the authorized person associated with the identity reference model corresponding to the highest score, and the role of the speaker is determined. The determination of the recognition result includes not only the identity information of the speaker, but also the role definition associated with this identity. The role information is extracted from the role database obtained in the previous steps, reflecting the responsibilities and scope of authority of the person in the current operation scenario. For example, in an operating room environment, the roles may include "surgeon", "anesthesiologist", "circulating nurse", etc., and each role has a clearly defined scope of operation authority and responsibilities. The recognition result will be recorded in the activity log together with the role information, including detailed information such as the recognition time, location, score value, and confidence level, providing a basis for subsequent auditing and analysis.
[0173] At the same time, update the real-time personnel status map, marking the latest activity time and location of the speaker. In a multi-person collaboration environment, this real-time identity and role recognition is crucial for ensuring operation safety. It can prevent unauthorized personnel from performing operations beyond their authority scope, and at the same time ensure that critical operations are carried out by personnel with corresponding qualifications.
[0174] This embodiment realizes continuous and accurate verification of the speaker's identity in a complex operation environment, significantly improving security and usability. Different from traditional one-time identity verification, this method can continuously monitor and verify the identity of the person issuing the instruction throughout the operation process, ensuring that each critical operation is carried out by authorized personnel. By only comparing with the authorized personnel models currently recorded as present, the computational complexity and potential identity confusion risk are greatly reduced, improving the accuracy and efficiency of recognition. It not only identifies the speaker's identity but also determines their role simultaneously. This identity-role mapping mechanism provides an accurate basis for subsequent permission control. In addition, the full-process verification method does not interfere with the normal operation process, improving the practicality and user acceptance. In one implementation manner of this embodiment, the predefined behavior rules include the instruction permissions and expected status ranges of the roles. According to the identified role and real-time context information, the corresponding predefined behavior rules are determined, including the following steps:
[0175] S601. Obtain the current stage information of the predetermined operation scenario as the real-time context information;
[0176] S602. Use the identified role and the current stage information as query conditions to retrieve, in the preset database, the set of instruction types authorized for the role to execute under the current stage information; and
[0177] S603. Retrieve the statistical range of the expected acoustic state characteristics of the role under the current stage information;
[0178] S604. Use the set of instruction types as the instruction permission and the statistical range of the expected acoustic state characteristics as the expected state range.
[0179] Obtain the current stage information of a predetermined operation scenario. Taking the operating room as an example, the surgical process can be divided into stages such as the anesthesia preparation stage, the surgical incision stage, the core operation stage, the suture stage, and the postoperative observation stage, etc. The current stage information can be obtained in various ways: It can directly obtain the current surgical progress information through the real-time data interface with the operating room management; it can also identify the current surgical state by analyzing the image data of the indoor camera in combination with a deep learning model; in addition, it can also infer the current stage by analyzing the usage of medical equipment in the operating room. When obtaining the stage information, a timestamp is recorded simultaneously to facilitate subsequent analysis of whether the stage duration is abnormal. The current stage information includes not only the stage name but also metadata such as the expected duration and standard operation process of this stage. These information together constitute the real-time context information.
[0180] Use the identified role and the current stage information as query conditions to retrieve the set of instruction types that the role is authorized to execute under the current stage information in a preset database. In a medical scenario, different roles have different operation permissions at different stages of the operation. The detailed role-stage-permission mapping relationship is stored in the preset database. When performing the query, first construct an SQL query statement or NoSQL query conditions, and use the identified role identifier (such as "surgeon") and the current stage identifier (such as "core operation stage") as exact match conditions. The query result returns a set of instruction types.
[0181] In this embodiment, the qualification level of the role is also checked. For example, some high-risk operations may only be executable by senior surgeons. If the query result is empty, it means that the role has no operation permission at the current stage. At this time, this abnormal situation is recorded and an alarm is triggered. This dynamic permission allocation mechanism based on roles and stages ensures that only the appropriate personnel perform the appropriate operations at the appropriate time, greatly improving the operation safety.
[0182] The statistical range of the expected acoustic state characteristics of the role under the current stage information can be used to determine whether the personnel state is normal. In a high-risk operation environment, the psychological and physiological states of the operators directly affect the operation safety. As a non-invasive physiological and psychological state indicator, the voice characteristics can effectively reflect the stress level, emotional state, and cognitive load of the personnel. The normal acoustic state characteristic statistical ranges of each role at different stages are stored in the preset database, and these ranges are obtained by analyzing a large amount of historical data and simulated training data.
[0183] In specific implementation, a query is constructed to retrieve the statistical range of the expected acoustic state characteristics of the identified role at the current stage from the database based on the role and the current stage. The retrieval results include the normal value ranges of multiple acoustic characteristics. For example, the mean range of the fundamental frequency statistic is [120 Hz, 150 Hz] and the standard deviation range is [10 Hz, 25 Hz]; the normal range of the jitter value is [0.005, 0.015]; the normal range of the shimmer value is [0.02, 0.06]; the normal distribution parameters of the energy characteristic are {mean: 65 dB, standard deviation: 5 dB}; the normal range of the speech rate characteristic is [3.5 syllables per second, 5.5 syllables per second]. Additionally, a personalized state baseline model is established for each authorized person to make the state assessment more accurate. In this way, it is possible to distinguish the acoustic changes caused by normal operating pressure from those caused by abnormal states (such as extreme tension or fatigue), providing a scientific basis for subsequent state judgment.
[0184] The set of instruction types is used as the instruction permission, and the statistical range of the expected acoustic state characteristics is used as the expected state range to construct a complete predefined behavior rule. In actual implementation, the above-obtained set of instruction types and the statistical range of the expected acoustic state characteristics are integrated to form a complete behavior rule object. This object consists of two main parts: The instruction permission part defines all the legal instruction types that the role can execute at the current stage. The subsequent recognized voice instructions will be matched with this set, and only the instructions that match successfully will be accepted and executed; The expected state range part defines the normal acoustic state parameter ranges of the role at the current stage. The real-time extracted acoustic characteristics will be compared with these ranges to evaluate whether the state of the speaker is normal.
[0185] The behavior rule object can also contain additional metadata, such as rule priority, rule expiration date, etc. For example, in an emergency, some high-priority rules may temporarily override the regular rules. The constructed behavior rule object can be cached in memory for quick access during subsequent real-time verification processes. At the same time, the rule application records can be written to the log for subsequent auditing and analysis. This dual-verification mechanism that combines instruction permission and state range not only ensures the legality of the instructions but also monitors the abnormal states of the operators, providing a more comprehensive security guarantee than traditional single-factor authentication.
[0186] This embodiment significantly enhances the security and reliability of permission verification. Compared with the traditional method that only relies on identity recognition, it can dynamically adjust permission allocation according to the real-time stage of the operation scenario to ensure that instructions are issued by the correct personnel at the correct time. At the same time, by monitoring the acoustic state characteristics of the speaker, it can timely detect the abnormal state of the operator and prevent high-risk operations from being performed in an inappropriate state. This multi-dimensional verification mechanism is particularly applicable to high-risk operation environments such as operating rooms and nuclear power plant control rooms, which can effectively reduce the risk of human errors and improve operation safety. In addition, its non-invasive feature does not interfere with the normal operation process, while the refined permission management and status monitoring provide detailed data support for post-event auditing and analysis, contributing to the continuous improvement of operation specifications and training programs.
[0187] In one implementation of this embodiment, based on predefined behavior rules, dynamic verification is performed on the real-time state acoustic characteristics to obtain a dynamic verification result, including the following steps:
[0188] S701. Compare each feature value in the real-time state acoustic characteristics with the expected state range, and count the number of features that exceed the expected state range;
[0189] S702. Based on the number of features, determine whether the real-time state of the speaker is normal or abnormal, and use the real-time state of the speaker as the dynamic verification result;
[0190] Among them, when the number of features is greater than the preset number threshold, it is determined that the real-time state of the speaker is abnormal, and when the number of features is not greater than the preset number threshold, it is determined that the real-time state of the speaker is normal.
[0191] In actual implementation, the real-time state acoustic characteristics are a set of acoustic parameters extracted from the speaker's speech, and these parameters can reflect the current physiological and psychological state of the speaker. During the comparison process, first obtain the expected state range of the current role at the current stage from the predefined behavior rules. The expected state range is expressed as the normal value interval of each feature. For example, the normal range of the average fundamental frequency may be [120Hz, 150Hz]. Then, each real-time extracted feature value is compared one by one to check whether it falls within the corresponding expected range. During specific implementation, the following formula can be used to calculate whether the feature value exceeds the range:
[0192]
[0193] where f i represents the real-time value of the i-th feature, min_range i and max_range irespectively represent the lower and upper limits of the expected range of the feature. In this embodiment, the Z-score can be used to measure the degree to which the feature value deviates from the center of the normal distribution:
[0194]
[0195] where μ i and σ i are the expected mean and standard deviation of feature i respectively. If Z i is greater than a certain threshold (usually 2 or 3), then the feature value is considered abnormal. Finally, count the number of features that exceed the expected range.
[0196] Based on the number of features, determine whether the real-time state of the speaker is normal or abnormal, and use the real-time state of the speaker as the dynamic verification result.
[0197] In this embodiment, if the number of features that exceed the expected range is greater than the preset number threshold, then determine that the real-time state of the speaker is abnormal; otherwise, determine it as normal.
[0198] Compared with the traditional method that only relies on identity recognition, this embodiment can detect state abnormalities caused by factors such as fatigue, stress, mood swings, or drug effects. Even for the same authorized person, it can distinguish between their normal and abnormal states. This dynamic verification mechanism is particularly suitable for high-risk operation environments such as operating rooms and nuclear power plant control rooms, and can effectively prevent human errors caused by the poor state of operators.
[0199] Referring to Figure 3 , this application embodiment also provides a model construction method, which is applied to an electronic device. The electronic device stores an identity reference model and a state baseline model. The construction method of the identity reference model includes the following steps:
[0200] S11. Obtain authorized person information;
[0201] S12. For each authorized person in the authorized person information, collect a voice sample set in multiple voice scenarios;
[0202] S13. Extract the deep voiceprint features of each authorized person from the voice sample set;
[0203] S14. Perform aggregation processing on the deep voiceprint features to generate an identity reference model corresponding to each authorized person;
[0204] The construction method of the state baseline model includes:
[0205] S21. Obtain authorized person information;
[0206] S22. For each authorized person in the authorized person information, collect a voice sample set in multiple voice scenarios;
[0207] S23. Select voice segments representing the normal working state from the voice sample set;
[0208] S24. Extract the state evaluation acoustic feature set of each authorized person from the voice segments, where the state evaluation acoustic feature set includes one or more of fundamental frequency statistics, jitter value, microtremor value, energy feature, and speech rate feature;
[0209] S25. Calculate the statistical distribution parameters of each state evaluation acoustic feature set in the normal working state, and the statistical distribution parameters include mean and / or standard deviation;
[0210] S26. Based on the statistical distribution parameters, determine a normal value range for each state evaluation acoustic feature set, and generate a state baseline model corresponding to each authorized person.
[0211] In actual implementation, the information of authorized persons is usually obtained from human resource management or permission management in specific scenarios. Such information includes but is not limited to: the unique identifier of the person (such as employee number, ID card number), name, position, department, qualification level, professional field, etc. In a medical scenario, it may also include specific information such as qualifications. The acquisition process can directly pull data through an API interface or import a pre-prepared personnel information form file.
[0212] For each authorized person in the authorized person information, collect a voice sample set in multiple voice production scenarios. In actual implementation, a comprehensive voice acquisition plan needs to be designed to cover various possible voice production scenarios and conditions.
[0213] First, it is necessary to determine the types of acquisition scenarios, including quiet environments (such as a dedicated recording studio), simulated working environments (such as a simulated operating room), and actual working environments (such as a real operating room), etc. Second, it is necessary to design the acquisition content, including preset texts (such as specific sentences or paragraphs), free conversations, professional terms, and instructions, etc. During the acquisition process, the authorized person is required to speak at different speech rates, volumes, and emotional states (such as calm, nervous, fatigued, etc.) to capture the natural changes in the voice. In terms of acquisition equipment, a high-quality microphone array should be used, and recordings should be made at different distances and angles to simulate various situations in the actual usage environment. To enhance the anti-noise ability of the model, different types and intensities of background noise can also be introduced during the acquisition process. Each authorized person needs to provide a sufficient number of voice samples, usually at least 30 minutes of effective voice data, distributed in different scenarios and conditions. The acquired voice data needs to be subject to quality inspection to eliminate samples with too low quality or obvious interference. The finally formed voice sample set will contain diverse voice data of each authorized person under various conditions, providing the basic materials for subsequent feature extraction and model construction.
[0214] Extracting the deep voiceprint features of each authorized person from the voice sample set is a process of converting the original voice data into a digital representation that can be used for identity recognition. In actual implementation, first, the original voice is preprocessed, including operations such as noise reduction, frame segmentation, and windowing, to improve the quality of subsequent feature extraction. Then, a deep learning model is applied to extract deep voiceprint features from the preprocessed voice. The deep learning model can include a deep neural network (DNN), a convolutional neural network (CNN), a long short-term memory network (LSTM), or a combination of them.
[0215] The above deep learning model is usually pre-trained on a large-scale voice dataset and can learn the deep-level features related to the speaker's identity in the voice. The specific feature extraction process is as follows: First, convert the voice signal into a time-frequency representation such as a spectrogram or Mel-frequency cepstral coefficients; then input these representations into the pre-trained deep neural network; finally, extract the output from the penultimate layer of the network as the deep voiceprint feature. The extracted features are usually high-dimensional vectors (such as 512-dimensional or 1024-dimensional) that contain the voice characteristic information that can uniquely identify the speaker. For each authorized person, multiple deep voiceprint feature vectors need to be extracted from all their voice samples, and these vectors together constitute the voiceprint feature set of this person.
[0216] Aggregating the deep voiceprint features to generate an identity reference model corresponding to each authorized person is a process of integrating multiple feature vectors into a unified model. In actual implementation, since each authorized person has multiple voice samples, multiple deep voiceprint feature vectors will be extracted. The above vectors need to be aggregated to generate a unified model that can represent the voice characteristics of this person. The aggregation process includes the following methods: average pooling method, that is, calculating the mean of all feature vectors as the final model; Gaussian mixture model, fitting the distribution of feature vectors through multiple Gaussian distributions; support vector data description, constructing a hypersphere to enclose the feature vectors; using a neural network encoder to encode multiple feature vectors into a unified representation. When selecting the aggregation method, it is necessary to consider the balance of computational complexity, storage requirements, and recognition performance. The finally generated identity reference model usually includes: feature vectors or vector sets, model parameters (such as the mean and covariance matrix of the GMM), and metadata (such as model creation time, update records, etc.).
[0217] In the process of constructing a state baseline model, in addition to basic identity information, attributes related to personnel state assessment also need to be concerned. Attributes can include: age group, gender, health status, work experience, assessment results of stress tolerance, etc. This information helps to understand and explain the differences in acoustic characteristics of different personnel in the normal state, providing a personalized reference baseline for subsequent state assessment. In a medical scenario, information such as the work intensity, shift schedule, and professional background of medical staff also needs to be obtained, as these factors may affect their voice performance during work.
[0218] The acquisition process can be achieved by docking with the human resources and health management systems. The authorized personnel information obtained will serve as the basic data for constructing the state baseline model, be used to organize subsequent voice collection and feature extraction work, and provide the necessary context information for the personalized state assessment model.
[0219] For each authorized person in the authorized personnel information, collecting a set of voice samples in multiple vocalization scenarios is the basic data collection step for constructing the state baseline model. Similar to the voice collection of the identity reference model, but in the construction of the state baseline model, more attention is paid to capturing the voice changes in different states.
[0220] In actual implementation, collection needs to be carried out in different working states, including normal rest state, mild workload state, moderate workload state, and state after high-intensity work, etc. Secondly, different psychological states can be induced by simulating different situations. The collection content should include standard texts and natural conversations (reflecting the voice characteristics in real work). To obtain more accurate state information, other physiological indicators (such as heart rate, blood pressure) can be recorded simultaneously during voice collection or subjective feelings (such as fatigue level, stress level) can be evaluated through questionnaires. Each authorized person needs to provide sufficient voice samples in different states. Usually, at least 10 minutes of effective voice data is required for each state. The collected data needs to be labeled to clarify the state category and intensity level corresponding to each segment of voice. These detailedly labeled voice sample sets will provide the necessary data support for subsequent feature extraction and state model construction, enabling the constructed state baseline model to accurately reflect the voice feature changes of each authorized person in different states.
[0221] Selecting voice segments representing the normal working state from the voice sample set is a crucial step in constructing the state baseline model, as the purpose of the state baseline model is to establish an acoustic feature reference standard for personnel in the normal state. In actual implementation, it is first necessary to clearly define the criteria for the normal working state, which generally refers to the state where personnel are moderately alert, without obvious fatigue, have stable emotions, and have a moderate cognitive load. From the collected voice sample set, voice segments that meet this criterion need to be screened. The screening can be based on multiple pieces of information: one is the state annotation recorded during collection, and select the segments marked as the normal state; the other is the physiological index data collected synchronously, and select the segments where indicators such as heart rate and blood pressure are within the normal range. Each authorized person should select a sufficient number of voice segments in the normal state, usually with a total duration of not less than 30 minutes, to ensure the reliability of subsequent statistical analysis. The selected voice segments will be marked and stored as the original data for constructing the state baseline model of this authorized person.
[0222] Extracting the state assessment acoustic feature set of each authorized person from the voice segments is a process of converting the original voice data into digital features that can be quantitatively analyzed. In actual implementation, the state assessment acoustic features mainly focus on those voice parameters that can reflect the physiological and psychological states of the speaker. First of all, the fundamental frequency statistics are one of the most basic features, including the mean, standard deviation, range, change rate, etc. of the fundamental frequency. These parameters reflect the basic characteristics of vocal cord vibration and may be affected by states such as stress and fatigue. Secondly, the jitter value measures the short-term change in the length of adjacent acoustic wave cycles and is an important indicator of voice stability, usually quantified by calculating the relative change in the duration of adjacent cycles. The shimmer value measures the short-term change in voice amplitude and reflects the stability of vocal cord vibration intensity. The energy features include the average energy, energy change rate, energy distribution, etc. of the voice, which can reflect the vitality level and emotional state of the speaker. The speech rate features are quantified by calculating the number of syllables, pause frequency, and duration per unit time, etc. These features may be affected by the cognitive load and fatigue state.
[0223] In the feature extraction process, first preprocess the voice, including noise reduction, framing (usually with a frame length of 20 - 30 ms), and windowing. Then apply signal processing algorithms to extract the above features. For example, use the autocorrelation function method to extract the fundamental frequency, use the waveform analysis method to calculate the jitter value and shimmer value, use the short-time energy analysis method to extract the energy features, and use the speech activity detection and syllable segmentation algorithms to extract the speech rate features. For each authorized person, extract these features from all their normal state voice segments to form a comprehensive state assessment acoustic feature set.
[0224] Calculating the statistical distribution parameters of each state assessment acoustic feature set in the normal working state can quantitatively describe the distribution law of the acoustic features of authorized personnel in the normal state.
[0225] In actual implementation, for the set of state - evaluation acoustic features extracted for each authorized person, statistical analysis needs to be carried out to understand the distribution characteristics of these features under normal conditions. First, calculate the mean of each feature, which is the arithmetic average of all sample values, and this reflects the central tendency of the feature. For example, the mean fundamental frequency of a certain authorized person under normal conditions may be 120 Hz, which will serve as the reference value for their fundamental - frequency feature. Second, calculate the standard deviation, which measures the degree of dispersion of the feature values and reflects the natural range of variation of the feature. For example, the standard deviation of the fundamental frequency of the same authorized person may be 10 Hz, indicating that their fundamental frequency usually fluctuates around the mean under normal conditions but does not deviate too much. Finally, each authorized person will obtain a set of statistical distribution parameters for each acoustic feature, and these parameters together describe the distribution law of the sound characteristics of this authorized person under normal working conditions, providing a scientific basis for determining the normal - value range in the follow - up.
[0226] Based on the statistical distribution parameters, determine a normal - value range for each set of state - evaluation acoustic features and generate a state - baseline model corresponding to each authorized person. In actual implementation, according to the above - calculated statistical distribution parameters, a reasonable normal - value range needs to be defined for each acoustic feature. An interval can be constructed based on the mean and standard deviation. For example, use "mean ± 2 times the standard deviation" as the normal range. For the fundamental - frequency feature, if the mean is 120 Hz and the standard deviation is 10 Hz, the normal range may be set as [100 Hz, 140 Hz].
[0227] After determining the normal - value range for each feature, pack these ranges together with the corresponding feature descriptions and statistical parameters to form a complete state - baseline model. Each authorized person will have an exclusive state - baseline model, which reflects their individual characteristics. These models will be stored in an electronic device and associated with the corresponding identity reference model to form a complete personnel - model library. Through this personalized state - baseline model, it is possible to accurately evaluate whether the real - time state of each authorized person is abnormal, improving the accuracy and reliability of state monitoring.
[0228] This embodiment establishes a two-layer model system for each authorized person: an identity reference model and a status baseline model. The two-layer model achieves more comprehensive security than traditional voiceprint recognition. The identity reference model ensures high accuracy and robustness of identity recognition by collecting voice samples in diverse scenarios and extracting deep voiceprint features, and can adapt to changes in different environments and speaking styles. The status baseline model extracts and statistically analyzes acoustic features reflecting physiological and psychological states by selecting voice segments in normal working states, and establishes a personalized status evaluation standard for each authorized person. This personalized modeling method takes into account individual differences among people, avoids misjudgments that may be caused by general standards, and improves the accuracy of status monitoring. It is particularly applicable to high-risk operation environments such as operating rooms and nuclear power plant control rooms, and can detect abnormal states of operators in a timely manner while ensuring identity security, and prevent potential security risks.
[0229] This embodiment of the present application further provides an electronic device, including:
[0230] A memory configured to store instructions; and
[0231] A processor configured to call the instructions from the memory and be capable of implementing the above-mentioned model construction method when executing the instructions.
[0232] Referring to Figure 2 , this embodiment of the present application further provides a voiceprint recognition system, including:
[0233] An entrance voiceprint recognition device;
[0234] An internal voiceprint recognition device;
[0235] An electronic device connected to both the entrance voiceprint recognition device and the internal voiceprint recognition device.
[0236] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0237] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0238] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0239] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0240] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0241] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0242] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0243] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0244] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A voiceprint recognition method for authority verification, characterized in that: Applied to a voiceprint recognition system, the voiceprint recognition system includes an entry voiceprint recognition device, an internal voiceprint recognition device and an electronic device, the electronic device stores an identity reference model and a corresponding state baseline model, each person corresponds to an identity reference model and a state baseline model, the method includes: In response to an entry request issued by a person at a fixed entrance area of a predetermined operation scenario, a random verification instruction is issued through an entrance voiceprint recognition device; Collecting the personnel's response voice to the random verification command, and extracting the entrance identity acoustic features from the response voice; Compare the entrance identity acoustic features with the identity reference model to obtain a comparison result, and determine whether to authorize the person to enter the predetermined operation scene based on the comparison result and a preset entrance verification threshold; When an authorized person enters a predetermined operation scene, the entry status of the person is recorded; In the predetermined operation scenario, the internal voiceprint recognition device continuously captures the ambient sound and extracts the target voice segment from the ambient sound; extracting real-time identity acoustic features and real-time state acoustic features from the target speech segment; Compare the real-time identity acoustic features with the identity reference model to identify the speaker and role of the target speech segment; Obtain real-time contextual information related to the intended operation scenario; Determine the corresponding predefined behavior rules based on the identified roles and real-time context information; Based on predefined behavior rules, dynamic verification is performed on the real-time state acoustic characteristics to obtain dynamic verification results; Generate output information based on dynamic verification results.
2. The method according to claim 1, characterized in that The entrance identity acoustic features are compared with the identity reference model to obtain a comparison result, and based on the comparison result and the preset entrance verification threshold, it is determined whether the authorized person enters the predetermined operation scene, including: Obtain identity reference models of all authorized personnel associated with a predetermined operation scenario; Calculate the similarity score between the entry identity acoustic features and each identity reference model; Determine the highest score among the similarity scores and the corresponding identity reference model, and determine whether the highest score is greater than or equal to a preset entry verification threshold; When the highest score is greater than or equal to the entry verification threshold, it is determined that the person corresponding to the identity reference model is authorized to enter the predetermined operation scenario.
3. The method according to claim 1, characterized in that Extracting the target speech segment from the ambient sound, including: Collect multi-channel audio signals through the microphone array configured by the internal voiceprint recognition device; Applying sound source localization algorithms to multi-channel audio signals to determine the direction of speech sources; Based on the direction of the speech source, the beamforming algorithm is applied to process the multi-channel audio signal to obtain an enhanced single-channel speech signal; Voice activity detection is performed on the enhanced single-channel speech signal to segment and extract the target speech segment.
4. The method according to claim 1, characterized in that Extract real-time identity acoustic features and real-time state acoustic features from the target speech segment, including: Applying a first preset algorithm, extracting a deep voiceprint feature for uniquely identifying a speaker from a target speech segment as a real-time identity acoustic feature; A second preset algorithm set is applied to extract real-time state acoustic features from the target speech segment, wherein the real-time state acoustic features include at least one of fundamental frequency statistics, jitter value, tremor value, energy feature and speech rate feature.
5. The method according to claim 1, characterized in that Compare the real-time identity acoustic features with the identity reference model to identify the speaker and role of the target speech segment, including: Obtain the identity reference model and corresponding role information of all authorized personnel currently recorded as entering the state; Calculate the similarity score between the real-time identity acoustic features and each identity reference model; Determine the highest score among the similarity scores and the corresponding identity reference model; Determine whether the highest score is greater than or equal to a preset recognition threshold; When the highest score is greater than or equal to a preset recognition threshold, the speaker who utters the target speech segment is identified as an authorized person associated with the identity reference model corresponding to the highest score, and the role of the speaker is determined.
6. The method according to claim 1, characterized in that Predefined behavior rules include the role's instruction authority and expected state range. According to the identified role and real-time context information, the corresponding predefined behavior rules are determined, including: Obtaining current stage information of a predetermined operation scenario as real-time context information; Using the identified role and current stage information as query conditions, searching in a preset database to obtain a set of instruction types that the role is authorized to execute under the current stage information; and Retrieve the statistical range of the expected acoustic state characteristics of the character under the current stage information; The instruction type set is used as the instruction authority, and the statistical range of the expected acoustic state characteristics is used as the expected state range.
7. The method according to claim 6, characterized in that Based on predefined behavior rules, dynamic verification is performed on the real-time state acoustic characteristics to obtain dynamic verification results, including: Compare each feature value in the real-time state acoustic feature with the expected state range, and count the number of features that exceed the expected state range; Based on the number of features, determine whether the real-time state of the speaker is normal or abnormal, and use the real-time state of the speaker as a dynamic verification result; When the number of features is greater than a preset threshold, the real-time state of the speaker is determined to be abnormal; when the number of features is not greater than the preset threshold, the real-time state of the speaker is determined to be normal.
8. A model building method, characterized in that: Applied to electronic devices, the electronic devices store an identity reference model and a state baseline model, and the method for constructing the identity reference model includes: Obtain authorized personnel information; For each authorized person in the authorized person information, voice sample sets in multiple speech scenarios are collected; Extracting deep voiceprint features of each authorized person from the voice sample set; Aggregate the deep voiceprint features to generate an identity reference model corresponding to each authorized person; The construction method of the state baseline model includes: Obtain authorized personnel information; For each authorized person in the authorized person information, voice sample sets in multiple speech scenarios are collected; Selecting a speech segment representing a normal working state from the speech sample set; Extracting a state assessment acoustic feature set of each authorized person from the voice clip, wherein the state assessment acoustic feature set includes one or more of fundamental frequency statistics, jitter value, tremor value, energy feature and speech rate feature; Calculate the statistical distribution parameters of each state assessment acoustic feature set under normal working conditions, where the statistical distribution parameters include a mean value and / or a standard deviation; Based on the statistical distribution parameters, a normal value range is determined for each state assessment acoustic feature set, and a state baseline model corresponding to each authorized person is generated.
9. An electronic device, characterized in that: include: a memory configured to store instructions; as well as A processor is configured to call the instructions from the memory and implement the model building method according to claim 8 when executing the instructions.
10. A voiceprint recognition system, characterized in that: The voiceprint recognition method for authority verification applied to any one of claims 1 to 7 comprises: Entry voiceprint recognition device; Internal voiceprint recognition device; The electronic device is connected to both the entrance voiceprint recognition device and the internal voiceprint recognition device.
Citation Information
Patent Citations
Voiceprint recognition method, voiceprint recognition device, voiceprint recognition equipment and storage medium
CN116564315A
Doctor-patient identity checking method, system and device based on voiceprint recognition and storage medium
CN116612762A
Data processing method and device, computer equipment, storage medium and product
CN116975823A
Intelligent business management and scheduling decision-making method of heat supply system based on voice interaction
CN118607959A
Voice state detection system and method based on voice interaction
CN119626229A
Cited By
Building construction access control method and system based on voice recognition
CN120690201A
Zero-configuration adaptive speaker recognition method and system
CN120708626A
A zero-configuration adaptive speaker recognition method and system
CN120708626B
Online speaker affiliation method and system based on voiceprint recognition
CN121641032A
Online speaker diarization method and system based on voiceprint recognition
CN121641032B