Patient identity verification and semantic archiving method based on voiceprint recognition

By employing state-aware voiceprint recognition and multi-speaker separation technology, the accuracy and privacy issues of voiceprint recognition in medical scenarios have been resolved, enabling efficient patient identity verification and semantic medical record archiving, thereby improving the level of medical informatization.

CN120853579APending Publication Date: 2025-10-28MINAMI ACOUSTICS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510710574.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing voiceprint recognition systems have low accuracy in medical settings, cannot adapt to changes in a patient's voice when they are not in good health, and suffer from issues such as voice privacy leaks and confusion caused by multiple speakers. Furthermore, voice medical record archiving is inefficient.

Method used

A state-aware voiceprint recognition method is adopted. By acquiring the patient's current state information to construct a state vector, and inputting it together with the speech data into the voiceprint recognition model, a state-aware voiceprint vector is generated. Similarity matching is performed and the pronunciation path residual is calculated. Combined with semantic analysis and multi-speaker separation technology, identity verification and semantic archiving are achieved.

Benefits of technology

It improves the adaptability and accuracy of voiceprint recognition in medical scenarios, enhances the reliability and security of recognition, protects voice privacy, and enables automatic archiving and traceability of structured medical records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853579A_ABST
    Figure CN120853579A_ABST
Patent Text Reader

Abstract

The invention discloses a patient identity verification and semantic filing method based on voiceprint recognition, and relates to the technical field of medical information processing, and the method comprises the following steps: obtaining voice data in a medical scene; inputting the voice data into a voiceprint recognition model; voiceprint vectors are extracted, current state information of a patient is obtained, and state vectors are constructed; jointly inputting the state vector and the voice data into a voiceprint recognition model to generate a state sensing voiceprint vector; performing similarity matching on the state sensing voiceprint vector and a template vector in a voiceprint database, and generating a synthetic voice; calculating a pronunciation path residual error between the synthetic voice and the original voice; and if the pronunciation path residual error is lower than a preset threshold value, determining that the initial identification identity is valid, and extracting the medical record information and filing the medical record information to a database corresponding to the patient identity. The technical problems of abnormal patient state, concurrence of multiple speakers, voice privacy protection, semantic archiving and the like in a medical scene are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information processing technology, and more specifically, to a patient authentication and semantic archiving method based on voiceprint recognition. Background Technology

[0002] With the deepening application of artificial intelligence technology in the medical field, voice-based intelligent recognition and information processing methods are gradually becoming an important part of medical support systems. Especially in patient identity management and medical record information collection, traditional manual recording and verification methods have many problems, such as high workload, low accuracy, high dependence on patients, and susceptibility to subjective interference. Therefore, more and more hospitals are beginning to try using voiceprint recognition technology to assist in patient identity verification, in order to achieve contactless automatic identification and voice archiving.

[0003] Existing voiceprint recognition systems are mostly based on voice modeling under healthy conditions, and are widely used in general speech recognition, security identification, and financial identity verification scenarios. However, in medical settings, patients are often in a state of post-operative illness, fever, sore throat, or other physical discomfort, leading to significant changes in vocalization patterns, timbre structure, and speech rate and rhythm. Traditional voiceprint models are less adaptable to these variations, resulting in a significant decrease in recognition accuracy. Furthermore, medical settings often involve multiple speakers in the same room, such as doctors interacting with multiple patients during ward rounds, or family members assisting in answering patient questions. In these situations, voice data is often mixed and intertwined, easily causing confusion regarding speaker attribution.

[0004] On the other hand, medical voice data is highly sensitive to privacy. Current technologies often directly perform speech recognition, recording, and playback after collecting patient voice data, lacking necessary information anonymization mechanisms, thus posing a certain risk of privacy leakage. Especially in public medical areas, remote consultations, or multi-user shared voice systems, once a patient's original voice is listened to or extracted, it is highly likely to expose their individual identity, pathological characteristics, and emotional state, causing patients hidden stress and even legal risks.

[0005] Meanwhile, current voice medical record archiving still mainly relies on doctors' dictation or manual input. Even when a voice recognition system is introduced, it usually only completes text transcription and fails to automatically classify structured medical record fields, resulting in low efficiency in subsequent review, analysis, and data mining.

[0006] To address the aforementioned issues, there is an urgent need to propose a robust, multi-speaker adaptable, verifiable, secure, and controllable voiceprint recognition and information processing solution suitable for medical contexts, with semantic archiving capabilities, to support the intelligent voice information flow and medical record construction needs in medical scenarios. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a patient authentication and semantic archiving method based on voiceprint recognition, so as to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A patient authentication and semantic archiving method based on voiceprint recognition includes:

[0010] Acquire voice data in medical settings;

[0011] The voice data is input into the voiceprint recognition model to extract the voiceprint vector;

[0012] When extracting the voiceprint vector, the patient's current state information is obtained and a state vector is constructed;

[0013] Based on the state vector and the speech data, a state-aware voiceprint vector is generated by jointly inputting the voiceprint recognition model.

[0014] The state-aware voiceprint vector is matched with the template vector in the voiceprint database to obtain the initial identification identity;

[0015] Synthetic speech is generated based on the voiceprint template corresponding to the initial identified identity.

[0016] Calculate the pronunciation path residual between the synthesized speech and the original speech;

[0017] If the pronunciation path residual is lower than a preset threshold, the initial identification identity is confirmed to be valid, and the original speech is subjected to semantic analysis to extract medical record information and archive it into the database under the corresponding patient identity.

[0018] In some embodiments, the state vector includes at least one symptom state label, which includes any one of pharyngeal inflammation, postoperative weakness, crying state, speech delay, and confusion.

[0019] In some embodiments, the pronunciation path residual is calculated using the following formula:

[0020]

[0021] Where Δ represents the pronunciation path residual, and T represents the number of speech frames. This represents the raw speech Mel spectrum value of frame t. Represents the Mel spectrum value of the synthesized speech in frame t;

[0022] If the residual of the pronunciation path is higher than the preset threshold, the recognition fails, the system abandons the assignment and requests confirmation from the doctor or waits for the next round of voice input for identity recognition.

[0023] In some embodiments, the method further includes: generating standard speech data based on the voiceprint template corresponding to the initial identification identity, the standard speech data being used to replace the original speech in doctor-patient communication, wherein the standard speech data is generated by removing emotional features, speech rate features and physiological state features from the original speech to protect patient privacy.

[0024] In some embodiments, the speech data includes multi-speaker speech signals, further comprising the following steps:

[0025] Perform speaker separation on the speech data to obtain multiple single-speaker speech segments;

[0026] The state-aware voiceprint extraction, identity matching, and pronunciation path residual calculation steps are performed on the multiple speech segments respectively.

[0027] In some embodiments, the speaker separation operation employs a multi-source separation model based on an attention mechanism.

[0028] In some embodiments, the candidate set of initial identification identities is limited based on the medical scenario context, and the limitation rules include:

[0029] Current room number, doctor's ward round route, duty log, and patient bed information;

[0030] Before performing similarity matching, the candidate identities in the voiceprint database are restricted to the candidate set.

[0031] In some embodiments, during the semantic analysis process, each clause after semantic segmentation is assigned to a matching medical record field and categorized according to the following fields:

[0032] Chief complaint, present illness, past medical history, medication history, allergy history, family medical history.

[0033] In some embodiments, the following steps are further included before generating the state vector:

[0034] Acoustic features such as speech rate, fundamental frequency change rate, pitch span, pause length, and energy fluctuation are extracted from speech data.

[0035] Based on the acoustic features, a multi-class neural network model is used to predict the corresponding symptom state labels, and the state labels are used to construct the state vector.

[0036] In some embodiments, the semantic analysis is followed by the following steps:

[0037] Construct a four-tuple structure that includes speaker identity, semantic fragments, timestamps, and pronunciation path residuals;

[0038] The quadruple structure is stored in a semantic archive database for medical record review, accountability tracing, and voice replay.

[0039] The advantage of this invention over existing technologies lies in its introduction of a state-aware voiceprint recognition mechanism, which solves the problem of vocal feature drift in patients under unhealthy conditions. During voiceprint recognition, the system inputs not only the patient's voice itself but also a state vector constructed from symptom states. This allows the voiceprint model to dynamically adjust its recognition criteria based on different pathological states, enabling stable identity recognition even under abnormal conditions such as postoperative weakness, pharyngitis, crying, or confusion, thus improving the adaptability and accuracy of voiceprint recognition in medical scenarios.

[0040] Furthermore, this invention introduces a two-way verification mechanism based on speech synthesis and path residual calculation after the initial recognition is completed. The system generates synthesized speech based on the recognition result and compares it with the original speech at the frame level to determine the difference in pronunciation paths. The identity is confirmed only when the residual is lower than a preset threshold. This closed-loop verification process significantly enhances the reliability of recognition, and is particularly suitable for multi-person medical environments with high identity similarity, effectively avoiding misidentification and impersonation.

[0041] In multi-speaker scenarios, this invention utilizes a multi-source separation model based on an attention mechanism to extract speech segments from multiple speakers, and performs state-aware modeling and identity verification separately, ensuring the parsability of speech data in multi-person concurrent scenarios and enhancing the actual deployment capability of the system.

[0042] To address patient voice privacy concerns, this invention generates standard voice data after successful identity verification to replace the original voice playback. This standard voice data does not carry identifiable elements such as emotion, speech rate, or physiological state, thus preventing patient identity guessing, disease assessment, or psychological identification due to the exposure of the original voice. This provides an executable technical path for voice privacy protection.

[0043] In terms of semantic archiving, after analyzing the voice content, this invention not only extracts the description of the illness, but also assigns the content to standardized medical record fields such as chief complaint, present illness history, and medication history based on semantic features. At the same time, by constructing a semantic quadruple structure that includes identity, time, and residuals, it achieves complete traceable voice archiving, which facilitates subsequent auditing and review and improves the level of medical informatization. Attached Figure Description

[0044] Figure 1 This is the overall flowchart of the method of the present invention;

[0045] Figure 2 This is a flowchart of the voiceprint and state vector extraction process of the present invention;

[0046] Figure 3This is a flowchart of the multi-speaker signal parsing and semantic archiving process. Detailed Implementation

[0047] The specific embodiments of the present invention will now be described with reference to the accompanying drawings.

[0048] In medical settings, this invention implements a patient authentication and semantic archiving method based on voiceprint recognition to ensure that medical staff can accurately identify patients and effectively record their voice information. This method is applicable to scenarios such as routine hospital care, ward rounds, and multi-patient consultations, and is particularly suitable for situations where patients' vocalizations are unstable, their voice content is complex, and their privacy is sensitive.

[0049] like Figure 1 The diagram shown is an overall flowchart of the present invention, which includes the following steps:

[0050] Acquire voice data in medical settings;

[0051] The speech data is input into the voiceprint recognition model to extract the voiceprint vector; while extracting the voiceprint vector, the patient's current state information is obtained and a state vector is constructed.

[0052] Based on the state vector and the speech data, a state-aware voiceprint vector is generated by jointly inputting the voiceprint recognition model.

[0053] The state-aware voiceprint vector is matched with the template vector in the voiceprint database to obtain the initial identification identity;

[0054] Synthetic speech is generated based on the voiceprint template corresponding to the initial identified identity, and the pronunciation path residual between the synthesized speech and the original speech is calculated.

[0055] If the pronunciation path residual is lower than a preset threshold, the initial identification identity is confirmed to be valid, and the original speech is subjected to semantic analysis to extract medical record information and archive it into the database under the corresponding patient identity.

[0056] like Figure 2 As shown, more specifically, in terms of voice data collection, microphones can be placed on doctors when they visit wards, or voice data in medical scenarios can be obtained through high-sensitivity microphone arrays set up in wards, examination rooms, or remote medical terminals.

[0057] Next, this invention employs a voiceprint recognition model based on a deep neural network structure to process the preprocessed speech data. Specifically, the system extracts voiceprint vectors using this voiceprint recognition model. These voiceprint vectors are high-dimensional feature vectors that can represent the speaker's identity, typically ranging from 128 to 512 dimensions. To improve the accuracy of voiceprint recognition, the voiceprint model can employ a pre-trained Transformer, ResNet, TDNN, or hybrid neural network architecture.

[0058] While extracting the voiceprint vector, to address specific vocal conditions of patients in medical settings (such as pharyngeal inflammation, postoperative weakness, crying, slow speech, or confusion), the system also acquires the patient's current state information and generates a state vector. Specifically, the state vector includes at least one symptom state label, which reflects the patient's possible abnormal vocal state, such as hoarseness caused by pharyngeal inflammation, weak vocalization due to postoperative weakness, or pitch changes during crying.

[0059] The state vector generation process is based on real-time analysis of multi-dimensional acoustic parameters in the original speech segments. During the speech frame processing stage, the system first extracts a series of numerical indicators describing the pronunciation features from each speech segment. These indicators include multiple dimensions such as frame-level speech rate, instantaneous rate of change of fundamental frequency, pitch range span, silence duration between speech frames, and overall signal energy fluctuation. These acoustic dimensions, after regularization, are combined into an initial pronunciation feature vector, which serves as the original input for state modeling.

[0060] In constructing the state recognition model, a lightweight multi-class neural network can be used as the pronunciation state discriminator. This network employs a hybrid structure of stacked convolutional and recurrent layers. The bottom layer consists of multiple one-dimensional convolutional layers, extracting stable but locally variable feature regions from the time series. The middle layer uses bidirectional gated recurrent units to model the temporal dynamics and non-linear rhythmic changes during the patient's pronunciation process. The top layer connects to a standard fully connected classification layer, outputting the multi-class probability distribution of the state labels. During training, the network incorporates a cross-entropy loss function and performs one-hot encoding supervision on the output state labels. The training data comes from a manually annotated medical speech dataset containing speech segments under various typical symptom states. Each speech segment has been manually annotated by doctors with its corresponding symptom, such as hoarseness in inflammation, weakness in postoperative speech, and frequency tremors during crying.

[0061] During the model inference phase, the system calculates the aforementioned acoustic feature vectors for each speech segment in real time and inputs them into the trained pronunciation state recognition network. The network outputs confidence probability vectors for multiple state categories. The system uses the state label corresponding to the highest probability term as the main symptom state of the current speech segment and embeds and encodes this label. Combined with the preceding probability information, it constructs the final state vector. This state vector serves as a crucial input to the voiceprint model, working in conjunction with the speech coding features to determine the generation process of the current voiceprint vector.

[0062] Throughout the process, the state recognition model and the voiceprint recognition model maintain a decoupled structure, are trained independently, and are jointly invoked in a modular manner during the recognition phase. This design preserves the independence and debuggability of state judgments while facilitating specific optimization of the representation space of the state vector. As a result, the entire system can maintain high voiceprint recognition accuracy and semantic attribution ability even under unstable, abnormal, and diverse patient speech conditions.

[0063] After constructing the state vector, the system does not treat it as an auxiliary label, but instead inputs it along with the original speech signal into the speaker recognition module for joint feature modeling. The core of this process is to build a deep speaker recognition network that can adapt to differences in pronunciation states at both the perception and expression levels, thereby solving the problem of decreased accuracy in traditional speaker systems when processing speech with unstable states.

[0064] To achieve deep fusion of state information, the voiceprint recognition model introduces a conditional vector adjustment mechanism into its overall architecture. This mechanism treats the state vector as a conditional control variable for each layer within the network, injecting it into the intermediate layers or attention calculation module of the model through a specific parameter path, rather than simply concatenating speech features and state features at the input stage.

[0065] Taking the Transformer architecture as an example, the state vector can be used to adjust the calculation method of attention weights in multi-head attention, so that the model has different perceptual weights for key regions in the speech frame sequence under different states. When using TDNN or ResNet architectures, the state vector can selectively enhance or suppress different acoustic feature dimensions through channel attention mechanisms, making it more consistent with the current pronunciation characteristics.

[0066] This joint modeling approach learns the "systematic influence of pronunciation state on voiceprint representation" during training, thus automatically adapting to differences in speech characteristics under different physical, emotional, or linguistic states during the inference phase. When the same patient pronounces in two states—one of high spirits and the other of postoperative weakness—although their physiological speech morphology changes, the model can accurately identify these changes as normal state shifts, rather than being caused by different individual pronunciations, through a state vector adjustment mechanism. The model does not rigidly rely on the absolute value of a fixed acoustic feature during recognition; instead, it learns judgment criteria based on a combination space of "speech + state," establishing an equivalent representation under state conditions in the latent voiceprint representation space, thereby effectively preventing voiceprint recognition mismatches caused by state drift.

[0067] During the inference process, the system generates the final voiceprint vector through a state-aware voiceprint model. This vector not only highly encapsulates the speaker's individual characteristics but also incorporates contextual information about the current pronunciation state into its internal encoding. This makes the vector more stable and adaptable for tasks such as identity matching, speech synthesis, and semantic attribution. For example, in the subsequent speech synthesis process, the state-aware voiceprint vector can be used to determine whether it is necessary to restore the patient's unique pronunciation state or to uniformly convert it into a standard speech style. Similarly, in semantic judgment, the system can combine state encoding to determine whether the current sentence might be affected by cognitive state, resulting in abnormal expression.

[0068] Through this joint embedding modeling method, the voiceprint recognition system has evolved from the traditional "static individual model" to a "state-condition model," which has the ability to explain the source of pronunciation changes, understand the dynamics of individual states, and dynamically correct voiceprint spatial offsets. This provides a key technological foundation for realizing context-adaptive voiceprint recognition in medical environments.

[0069] To further enhance recognition accuracy and reduce false recognition rate, this invention introduces a candidate identity screening strategy for medical scenarios. Before performing voiceprint template matching, the system constructs an identity constraint range based on the physical and procedural context of the current voice acquisition. This range is constructed by referencing multiple sources, including room number, bed binding information, doctor's ward round records, and nurse station sign-in records. The system dynamically generates an identity candidate pool through real-time interaction with the Hospital Information System (HIS) or Electronic Medical Record System (EMR). The candidate pool limits the recognition scope to the list of individuals who are likely to speak within the current scenario, effectively eliminating interference from a large number of irrelevant individuals in the database. This context-constrained strategy not only improves computational efficiency but also significantly reduces the probability of false recognition across wards and departments in practical applications.

[0070] After the state-aware voiceprint vector is extracted, the system does not directly output the recognition result, but enters the matching stage. The goal of this stage is to perform a high-precision comparison between the newly generated state-aware voiceprint vector and the vectors in the pre-established voiceprint template library in the system, thereby confirming the speaker's true identity.

[0071] The voiceprint template library is constructed from registered speech samples created for each patient in advance, containing multiple sets of voiceprint vectors for each patient in various speech states. Each template is not just an average vector, but a set of vectors, typically containing multiple representative speech samples collected from the patient on different dates and in different states (such as awake, fatigued, post-operative, etc.). To improve matching accuracy, each template data in the library also includes relevant additional information, such as the corresponding state label, speech length, and background noise level, for fine-grained filtering and weighting during matching.

[0072] During the matching process, the system first calculates the similarity between the candidate voiceprint vector and each voiceprint template vector in the template library. Similarity calculation typically uses cosine similarity or Euclidean distance, but a trained metric learning model can also be used to generate a discrimination score. Taking cosine similarity as an example, let the state-aware voiceprint vector be v. q A template vector is v i The similarity is defined as:

[0073]

[0074] The numerator is the dot product of the two vectors, and the denominator is the L2 norm product of the two vectors. The closer the similarity value is to 1, the closer the two vectors are in direction, meaning the more similar their voiceprints are.

[0075] To improve comparison efficiency, the system typically employs an approximate nearest neighbor search algorithm (such as the HNSW or IVF structure in the Faiss library) to quickly select several most similar template vectors from a large-scale vector database. After initial screening, the system filters candidate results based on contextual information (such as ward location, shift schedule, etc.), retaining only the voiceprints of individuals "likely to speak" in the current scenario. For the retained candidates, the system performs a final, precise full-scale similarity comparison to obtain the final score and ranking.

[0076] The candidate with the highest score is output as the initial identification result and enters the subsequent reversible verification mechanism. If multiple candidate scores are close and the system's judgment is uncertain, multiple candidate identities can be retained simultaneously, and the final identity can be further confirmed by generating speech residuals, semantic comparison, or context prediction.

[0077] After extracting the state-aware voiceprint vector and matching it with the voiceprint template library, the system obtains candidate identities. Although the identification result has a high similarity in the vector space, considering that patients in a medical environment may have similar voiceprints or noise interference, directly accepting this identification result may introduce risks. Therefore, to further verify whether the identity attribution is authentic and reliable, this invention introduces a bidirectional reversible identity verification mechanism, upgrading the identification process from "one-way comparison" to "closed-loop verification".

[0078] Specifically, after obtaining a candidate identity, the system retrieves the corresponding voiceprint vector from the voiceprint template library and inputs this vector into the speech generation module. This module uses a neural network structure to perform reverse speech reconstruction, generating speech sequences that are highly consistent with the acoustic features of the identity, conditioned on the voiceprint vector. The speech generation module can employ a Tacotron structure based on an attention mechanism, or use a more efficient FastSpeech-like non-autoregressive structure, significantly reducing inference latency while maintaining high speech quality.

[0079] The speech data output by the speech synthesis module is the "fitted speech," which should have a strong similarity to the original input speech in the spectral space. If the recognition result is accurate, the generated fitted speech should highly overlap with the original speech in key dimensions such as pronunciation path and spectral trend. To quantify this degree of overlap, this invention designs a residual calculation mechanism based on frame-level Mel-spectral differences.

[0080] Both the original and synthesized speech are processed into a series of consecutive Mel-spectral frames. Assuming each speech segment is divided into T frames, the system sequentially calculates the absolute value of the Mel-spectral difference in each frame. The overall difference is then summed using the frame-level absolute residuals and normalized to the frame number, forming the final pronunciation path residual Δ. The formula for calculating this residual is as follows:

[0081]

[0082] Where Δ represents the pronunciation path residual, and T represents the number of speech frames. This represents the raw speech Mel spectrum value of frame t. Let represent the Mel spectrum value of the synthesized speech in frame t.

[0083] The system sets a residual threshold based on empirical data and actual recognition scenarios. When Δ is less than the threshold, it indicates that the synthesized speech accurately reproduces the pronunciation trajectory of the original speech, which means that the recognition result is highly consistent with the template in terms of physical pronunciation path, and the recognized identity can be identified as real and valid.

[0084] If Δ exceeds the preset threshold, it indicates a significant deviation between the voice generated by the candidate identity and the original voice in the key spectral path, potentially leading to misidentification. In this case, the system will not directly assign the voice to the candidate identity, but will interrupt the recognition process and generate a system prompt suggesting that the doctor manually confirm the voice using other methods, or wait for the next round of voice input to complete the supplementary recognition.

[0085] By reverse-engineering the speech recognition results and comparing them with the original speech at the spectral level residual, this invention constructs a complete verification loop in the voiceprint recognition process, significantly improving the security and reliability of the recognition results. It is particularly suitable for high-risk medical scenarios such as the coexistence of patients with similar identities, complex pronunciation states, and unstable audio quality. This mechanism not only corrects accidental similarity misjudgments in the voiceprint vector space but also provides a speech-level credibility calculation method that does not rely on manual judgment, laying an accurate and reliable identity foundation for subsequent semantic archiving, speech synthesis, or information broadcasting.

[0086] Furthermore, in the actual operation of the healthcare system, voice interactions between patients and doctors do not always occur in a single, closed environment. Especially in scenarios such as ward rounds, free clinics, preoperative discussions, and group rehabilitation consultations, multiple patients and their families often participate in the conversation simultaneously. The voice data collected in these multi-person interactive environments often exhibits characteristics such as speaker overlap, speech crossover, and uneven speech rates, making speaker attribution a highly challenging aspect of voiceprint recognition. Directly inputting mixed voice data into a voiceprint model for recognition can easily lead to speaker misidentification and semantic confusion, affecting the accuracy of subsequent medical records and legal traceability.

[0087] Therefore, such as Figure 3 As shown, this invention introduces a multi-speaker signal parsing module into the basic voiceprint recognition process to handle overlapping speech trajectories in multi-speaker speech signals. The system first performs frame-level segmentation on the original speech signal and then uses an attention-based multi-source separation network for signal separation. This model utilizes context-aware mechanisms and cross-time-step sequence attention to capture the differences between different speakers in the frequency and time domains, combined with residual path structures to improve separation efficiency. During training, the model aligns the temporal spectral distribution between training samples and target speech sources, achieving accurate separation of multiple clearly identifiable speech streams from a single-channel input.

[0088] During the model inference phase, the system inputs each mixed speech segment into the sound source separation network, resulting in multiple independent speech trajectories. These separation results are structurally very close to single-speaker speech data, exhibiting strong speaker consistency and speech content integrity. Each separated speech trajectory is treated as an independent recognition unit and sequentially fed into the subsequent voiceprint recognition processing flow.

[0089] On each speech trajectory, the system performs the same state vector extraction, state-aware voiceprint recognition, and pronunciation path residual verification process as for single-person speech. The state recognition section determines whether the speaker is in a specific state for each speech track, such as pharyngitis, emotional excitement, or physical weakness, generating a corresponding state vector to ensure the voiceprint model accurately perceives the current vocal conditions. The voiceprint matching module then compares the voiceprint feature vector generated from this state-aware vector with templates in the voiceprint database and generates preliminary recognition results. The system then performs closed-loop verification of these results through the speech synthesis and path residual evaluation modules to ensure that each speech trajectory in complex multi-person scenarios obtains a clear and reliable speaker attribution.

[0090] Because multi-person voice interaction scenarios are often accompanied by uncontrollable factors such as background noise interference, differences in the amplitude of the speaker's own voice source, and overlapping speech rates, the multi-speaker recognition structure designed in this invention not only solves the "drift recognition" problem that occurs in traditional recognition systems in mixed speech, but also establishes a scalable, autonomous, and verifiable multi-speaker identity recognition link, further improving the accuracy and robustness of the system in processing voice data in real medical environments. This mechanism supports various voice scenarios such as single-person speaking, alternating speaking, and simultaneous speaking, and can output traceable voice attribution results, providing clear and structured input data for subsequent modules such as semantic archiving, synthesized speech construction, and medical liability determination.

[0091] After completing the state-aware voiceprint recognition of the patient's identity and confirming the authenticity and validity of the recognition result through a synthetic speech residual verification mechanism, this invention does not directly use the patient's original voice as subsequent communication or recording content. Instead, it further introduces a standardized voice replacement mechanism to effectively address the practical needs for voice privacy protection in medical scenarios. In traditional medical systems, patients' voice content is often directly used for remote communication, recording and storage, voice medical record retrieval, or automatic playback. However, the patient's original voice usually carries obvious individual characteristics, such as speech rate, tone of voice, pronunciation habits, accent differences, and even current physiological state (such as weakness, anxiety, or agitation). If this individual information is listened to, captured, or analyzed outside the medical environment, it may lead to the inference of the patient's identity, speculation about their condition, or even discriminatory judgments.

[0092] To address this risk, this invention designs a speech standardization processing module. Using a confirmed patient voiceprint template as input, a substitute speech is synthesized through a neural network speech generation model. This substitute speech does not contain the subjective emotional fluctuations of the original speech, does not reflect immediate physiological changes, and does not maintain the original speech rate and rhythm, thus achieving neutralization and standardization of vocal expression. The speech generation model structure can be based on non-autoregressive speech generation techniques such as FastSpeech, Glow-TTS, or Diffusion models, supplemented by a control vector encoding mechanism. This allows the system to selectively "suppress" certain expressive dimensions without altering the semantic content. For example, by adding a constant-value emotion smoothing parameter, intonation fluctuations are compressed to an average range; by setting a uniform speech rate ratio controller, fast-paced segments are output slowly while maintaining the semantic logical structure.

[0093] While generating speech, the system also removes high-frequency noise, blurred syllables, and unstructured components such as liaison and phrasing that reflect the patient's physiological state (fatigue, pain, anxiety, etc.) from the original audio. The output standardized speech has a highly predictable expressive structure, and its acoustic parameters are close to the "template voice patterns" set in the training library, such as neutral female or male voices. This is suitable for medical assistants to standardize the reading of patient information or for relaying speech content in remote call scenarios. This standardized audio, uniformly synthesized by the system, can completely preserve the actual semantics of the patient's speech without exposing the speaker's true identity, physiological state, and emotional characteristics.

[0094] At the application level, this standard voice can be directly used in various scenarios such as hospital broadcasting, voice output of intelligent triage systems, voice bridging of remote consultation platforms, playback modules of voice electronic medical record systems, and voice-based medical knowledge Q&A platforms. Especially in the context of remote voice calls, the system converts the patient's real-time voice content into standard audio in real time and sends it to the remote doctor or platform, ensuring that the original audio is never forwarded or stored during the communication process. Even if the call is intercepted or recorded, it is impossible to deduce the patient's identity from the voice, thus establishing a two-way secure communication mechanism based on structural semantic preservation and acoustic privacy stripping.

[0095] Through this standardized voice replacement mechanism, this invention not only provides a semantic transmission path in medical voice streams, but also achieves "de-identification processing" of individualized voiceprint information at the source, effectively blocking attack paths such as external eavesdropping, voice synthesis forgery, and patient identity profiling, and establishing basic privacy protection capabilities for the remote circulation, sharing, collaboration, and cross-platform use of medical voice data.

[0096] After completing the identity recognition and verification process, this invention does not stop at converting voice information into text records, but further performs deep semantic-level automated processing on the original voice content to achieve structured medical record archiving. This processing not only considers the lexical and syntactic information of the language itself, but also combines domain knowledge in the medical context to classify, index, and temporally organize semantic segments, ultimately transforming the entire voice data from freely expressed natural language into data units with medical semantic structures.

[0097] In terms of semantic analysis implementation, the system of this invention employs a language understanding model trained on a large-scale medical corpus. During the pre-training phase, this model incorporates large-scale electronic medical record texts and clinical question-and-answer dialogue pairs as training data sources. Enhanced by medical entity annotation and a terminology dictionary, it is able to identify key information units in medical discourse, such as symptom names, disease progression descriptions, past events, drug names, and treatment procedures. By processing the speech recognition text sentence by sentence, the system can decompose multiple parallel or nested information items in long sentences, and each decomposed semantic fragment is assigned to a standard medical field.

[0098] The classification of medical record fields follows standard clinical recording criteria, including but not limited to chief complaint, present illness, past medical history, medication history, allergy history, and family medical history. During classification, the system uses a soft classification strategy to associate uncertain statements with multiple fields based on keywords, contextual relationships, and the chronological order of events within the semantic content. For example, when the voice says "I used to be allergic to cephalosporins," the system not only classifies it into the allergy history field but also marks it as a past drug reaction event and associates it with the drug entity "cephalosporin." This semantic linkage capability ensures that the automatic archiving results not only have accurate fields but also reflect the inherent logical relationships between medical events.

[0099] To support subsequent medical record auditing, legal liability determination, and temporal tracing of voice content, the system binds each semantic unit to its source information after semantic extraction, constructing a four-tuple data structure. This four-tuple includes the speaker's identity identifier, the parsed semantic segment content, the timestamp position in the voice data, and the pronunciation path residual information obtained during voiceprint recognition. The timestamp information can locate the start and end positions of specific sentences in the speech, while the residual information serves as an auxiliary parameter for evaluating recognition reliability.

[0100] The four-tuple structure transforms speech recordings from simple sentence-by-sentence text recognition into structured medical record information with logical indexing, temporal localization, and source verification capabilities. The system can perform multi-dimensional indexing based on this structure. For example, in the event of a dispute or controversy, it can quickly locate a specific passage spoken by a particular patient within a certain time period, extract its semantic content, and verify it using voiceprint path residuals, thereby constructing a complete chain of auditory evidence. Furthermore, this structure provides a highly valuable foundational data model for future medical natural language research, automated follow-up modeling, and multi-round consultation mapping.

[0101] This invention transforms patients' original speech into semantically clear, clearly attributed, and structurally stable medical record information units, and introduces a four-tuple storage mechanism, realizing a complete leap from the perception layer to the knowledge layer of medical voice data. This makes voice medical records not only readable and writable, but also traceable, auditable, and verifiable.

[0102] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A patient authentication and semantic archiving method based on voiceprint recognition, characterized in that, include: Acquire voice data in medical settings; The voice data is input into the voiceprint recognition model to extract the voiceprint vector; When extracting the voiceprint vector, the patient's current state information is obtained and a state vector is constructed; Based on the state vector and the speech data, a state-aware voiceprint vector is generated by jointly inputting the voiceprint recognition model. The state-aware voiceprint vector is matched with the template vector in the voiceprint database to obtain the initial identification identity; Synthetic speech is generated based on the voiceprint template corresponding to the initial identified identity, and the pronunciation path residual between the synthesized speech and the original speech is calculated. If the pronunciation path residual is lower than a preset threshold, the initial identification identity is confirmed to be valid, and the original speech is subjected to semantic analysis to extract medical record information and archive it into the database under the corresponding patient identity.

2. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, The state vector includes at least one symptom state label, which includes any one of pharyngeal inflammation, postoperative weakness, crying state, slow speech, and confusion.

3. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, The pronunciation path residual is calculated using the following formula: Where Δ represents the pronunciation path residual, and T represents the number of speech frames. This represents the raw speech Mel spectrum value of frame t. Represents the Mel spectrum value of the synthesized speech in frame t; If the residual of the pronunciation path is higher than the preset threshold, the recognition fails, the system abandons the assignment and requests confirmation from the doctor or waits for the next round of voice input for identity recognition.

4. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, The method further includes: generating standard speech data based on the voiceprint template corresponding to the initial identification identity, wherein the standard speech data is used to replace the original speech in doctor-patient communication, and wherein the standard speech data removes emotional features, speech rate features and physiological state features from the original speech during generation to protect patient privacy.

5. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, The speech data includes multi-speaker speech signals and further includes the following steps: Perform speaker separation on the speech data to obtain multiple single-speaker speech segments; The state-aware voiceprint extraction, identity matching, and pronunciation path residual calculation steps are performed on the multiple speech segments respectively.

6. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 5, characterized in that, The speaker separation operation described herein employs a multi-source separation model based on an attention mechanism.

7. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, The candidate set for the initial identification is limited based on the medical scenario context, and the limitation rules include: Current room number, doctor's ward round route, duty log, and patient bed information; Before performing similarity matching, the candidate identities in the voiceprint database are restricted to the candidate set.

8. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, In the semantic analysis process, each clause after semantic segmentation is assigned to a matching medical record field and categorized according to the following fields: Chief complaint, present illness, past medical history, medication history, allergy history, family medical history.

9. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, Before generating the state vector, the following steps are further included: Acoustic features such as speech rate, fundamental frequency change rate, pitch span, pause length, and energy fluctuation are extracted from speech data. Based on the acoustic features, a multi-class neural network model is used to predict the corresponding symptom state labels, and the state labels are used to construct the state vector.

10. The patient authentication and semantic archiving method based on voiceprint recognition according to claim 1, characterized in that, Following semantic analysis, the following steps are further included: Construct a four-tuple structure that includes speaker identity, semantic fragments, timestamps, and pronunciation path residuals; The quadruple structure is stored in a semantic archive database for medical record review, accountability tracing, and voice replay.