Operation process recording and instruction interaction system based on multi-mode biological characteristic verification

The surgical record system, which uses a multimodal biometric verification system to collect data in real time and use encrypted signatures, solves the problems of surgical record delays and information distortion. It enables real-time and accurate recording of voice commands and generation of digital evidence during surgery, improving the timeliness and security of surgical records.

CN122024982APending Publication Date: 2026-05-12ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-01-29
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies rely on postoperative recollections for surgical records, leading to information distortion and making it difficult to capture physician instructions completely and in real time. Furthermore, traditional electronic signatures are not applicable to sterile surgical environments, resulting in delays in the chain of evidence.

Method used

A multimodal biometric verification system is adopted, which collects voice streams through a multi-channel microphone array, combines biometric sensors and heterogeneous on-chip systems to perform real-time speech-to-text conversion and natural language understanding, and performs hardware-level encryption and signing to generate tamper-proof signed data blocks.

Benefits of technology

It enables real-time and accurate recording of voice commands during surgery, improving the timeliness and security of surgical records, providing immediate and traceable digital evidence, and enhancing the compliance of medical procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024982A_ABST
    Figure CN122024982A_ABST
Patent Text Reader

Abstract

The invention discloses an operation process recording and instruction interaction system based on multi-mode biological characteristic verification. The system comprises a multi-channel microphone array used for collecting audio signals in an operating room environment; the biological characteristic sensor is used for collecting biological characteristic information of a physician; and a surgical interaction and verification processor. The processor is configured to: process the audio signal to generate a physician voice stream of high signal-to-noise ratio; converting the voice stream into a text stream in real time and carrying out natural language understanding so as to process and form one or more structured data entries; and when a submission trigger instruction is received or the biometric sensor is activated, performing non-repudiation verification and encrypted signature on the one or more structured data entries based on the biometric information to generate a signed data block. The technical problems of operation record delay, information distortion and weak evidence chain in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical information technology, and in particular to a surgical procedure recording and command interaction system based on multimodal biometric verification. Background Technology

[0002] In modern surgical practice, detailed and accurate surgical records are crucial for ensuring medical quality, meeting legal and regulatory requirements, and conducting subsequent clinical research. Currently, the generation of surgical records mainly relies on the surgeon writing them from memory post-surgery, or on the circulating nurse assisting in recording during the operation. This traditional model has several inherent technical problems: First, post-operative recall-based recording is prone to information distortion or omission of key details due to memory bias; second, manual recording during the operation can distract the assistants and makes it difficult to capture every verbal instruction and key decision given by the surgeon in real time and completely; finally, the confirmation and signing of the record are usually completed several hours or even longer after the operation, resulting in a significant delay in the chain of evidence. Summary of the Invention

[0003] This application aims to provide a surgical procedure recording and instruction interaction system based on multimodal biometric verification, in order to solve the technical problems of surgical records in the prior art, such as delays, high risk of information distortion, and weak legal evidence chains.

[0004] In a first aspect, this application provides a surgical procedure recording and instruction interaction system based on multimodal biometric verification, comprising: a multi-channel microphone array configured to acquire multiple raw audio streams within an operating room environment; a biometric sensor configured to acquire biometric data of the surgeon; and a surgical interaction and verification processor configured to: perform spatial filtering on the multiple raw audio streams to generate a high signal-to-noise ratio (SNR) physician speech stream; perform real-time speech-to-text processing and natural language understanding processing on the high SNR physician speech stream to distinguish the physician speech stream into text containing procedure descriptions and text containing executable instructions, and process the aforementioned text to form one or more data entries to be signed; and, in response to a submission trigger event, associate the one or more data entries to be signed with the biometric data acquired by the biometric sensor, and perform an encrypted signing operation to generate a timestamped, tamper-proof signed data block.

[0005] Optionally, the surgical interaction and verification processor is a heterogeneous system-on-a-chip, which physically integrates: a digital signal processor core configured to perform the spatial filtering process; a central processing unit core cluster configured to perform the real-time speech-to-text processing and the natural language understanding process; and a security coprocessor configured to perform the encryption signature operation in a physically isolated environment.

[0006] Optionally, the heterogeneous system-on-a-chip also integrates a state-gated asynchronous data bus, which is configured to connect the central processing unit core cluster and the security coprocessor.

[0007] Optionally, the state-gated asynchronous data bus is configured to operate in two mutually exclusive working states: a transcription state, in which the central processing unit core cluster is allowed to write one or more data entries to be signed into a shared temporary buffer; and a commit state, in which write access to the temporary buffer is locked by hardware logic, and a one-time, uninterruptible data block transfer is performed to transmit the entire contents of the temporary buffer as a data snapshot to the security coprocessor.

[0008] Optionally, the submission trigger event includes at least one of the following: the natural language understanding processing of the surgical interaction and verification processor recognizes a preset submission trigger voice command; or, the biometric sensor is activated by a specific physical action of the surgeon.

[0009] Optionally, the biometric sensor is an iris scanner or a microphone for collecting acoustic fingerprints.

[0010] Optionally, the security coprocessor is further configured to: compare and verify the biometric data collected in real time with a pre-stored biometric template before performing the cryptographic signature operation; and continue to perform the cryptographic signature operation only if the comparison and verification are successful.

[0011] Optionally, the signed data block includes a data content portion and a cryptographic signature portion; wherein the data content portion includes the one or more data entries to be signed, the timestamp, and a verifier identity identifier; the cryptographic signature portion is generated by applying a cryptographic hash algorithm to the data content portion.

[0012] Optionally, the system further includes a display unit configured to present, in real time, the text converted from the physician's voice stream and the one or more data entries to be signed, before the cryptographic signing operation is performed, for the surgeon to confirm immediately.

[0013] Optionally, the state switching of the state-gated asynchronous data bus is controlled by a hardware interrupt signal, which is generated by the submission trigger event.

[0014] Secondly, this application provides a surgical procedure recording and instruction interaction system based on multimodal biometric verification, comprising: a spatial filtering module configured to generate a high signal-to-noise ratio (SNR) physician speech stream based on multiple raw audio streams acquired from a multi-channel microphone array; a semantic understanding and structuring module configured to convert the high SNR physician speech stream into structured data entries, wherein the structured data entries are distinguished into process descriptive text and executable instruction text; and a secure signature and solidification module configured to, in response to a submission trigger event, perform an encrypted signature operation on the structured data entries based on the biometric data of the surgeon acquired from a biometric sensor, to generate a timestamped, tamper-proof signed data block.

[0015] Optionally, the secure signature and persistence module includes a hardware security coprocessor configured to perform the cryptographic signature operation in a physically isolated environment.

[0016] The technical solution provided in this application seamlessly integrates the traditionally separate "intraoperative procedures" and "postoperative documentation" stages into a synchronous process by capturing and structuring the physician's voice commands in real time during surgery and instantly signing and solidifying them using non-repudiable biometrics. This solution fundamentally eliminates the risk of information distortion caused by memory delays, greatly improving the accuracy and timeliness of surgical records. Simultaneously, by introducing hardware-level security isolation and encrypted signature mechanisms, it provides immediate and traceable digital evidence for key decisions and instructions during surgery, enhancing the security and compliance of the medical process. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the system architecture of a surgical procedure recording and command interaction system based on multimodal biometric verification, according to an embodiment of this application.

[0019] Figure 2 This is an internal structure block diagram of a surgical interaction and verification processor according to an embodiment of this application.

[0020] Figure 3 This is a flowchart of a surgical procedure recording and instruction interaction method according to an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0023] Example 1

[0024] This embodiment provides a surgical procedure recording and command interaction system based on multimodal biometric verification. In a specific implementation, this system deeply integrates real-time multimodal data acquisition, semantic understanding and structuring, and hardware-level secure signature embedding, thereby enabling real-time and accurate recording of voice commands and key operations during surgery. This method solves the technical problems of existing technologies, such as poor accuracy and easy information distortion caused by surgical records relying on postoperative recollection, and the cumbersome process of traditional electronic signatures, which are unsuitable for aseptic surgical environments. It achieves the beneficial effect of generating legally valid digital evidence in real time during the golden moment of surgery.

[0025] Reference Figure 1 The architecture of system 100 includes multiple cooperating functional entities. These entities together constitute a complete workflow for data acquisition, processing, verification, and archiving. System 100 includes a multi-channel microphone array 110, a biometric sensor 120, and a surgical interaction and verification processor 130 as the core of the system. Optionally, the system may also include a display unit 140 and an EMR system interface 150 for communicating with a hospital information system (HIS) or electronic medical record (EMR) system.

[0026] The multi-channel microphone array 110 serves as the system's auditory input front-end, configured to acquire multiple raw audio streams within the operating room environment. In the complex acoustic environment of an operating room, characterized by high noise and multiple conversations, the array's design is crucial to ensure accurate capture of the surgeon's voice. In one specific implementation, the multi-channel microphone array 110 comprises eight high-sensitivity microelectromechanical systems (MEMS) microphone units evenly distributed on a circular array substrate with a radius of 70 millimeters (mm). This circular geometry allows for comprehensive coverage of sound sources from a 360-degree horizontal plane. Each microphone unit converts the received sound pressure level signal into an independent digital audio signal stream, with a sampling rate set at 48 kHz and a quantization bit depth of 24 bits, ensuring high fidelity of the audio signal. These eight parallel raw audio streams are synchronously transmitted to the surgical interaction and verification processor 130 for further processing. This array can be integrated into the center of the operating room's shadowless lamp or designed as a lightweight wearable device for physicians to optimize the accuracy of sound source localization. Through multi-point acquisition at the physical level, it provides the necessary spatial information for subsequent sound source separation and noise suppression algorithms.

[0027] The biometric sensor 120 functions as a signature tool for authentication and operation confirmation within the system, and is configured to collect the biometric data of the surgeon. To accommodate the special requirements of the surgeon's hands being sterile and unable to make physical contact during surgery, this sensor employs non-contact biometric identification technology.

[0028] Exemplarily, in a preferred embodiment, the biometric sensor 120 is an iris scanner integrated into the eyepiece of a surgical microscope or the goggles worn by the physician. The scanner includes a near-infrared LED light source with a wavelength of 850 nanometers (nm) to illuminate the physician's iris, and an image sensor with a resolution of 1280x720 pixels to capture a high-resolution image of the iris texture. When the physician needs to sign a record, a brief gaze at a specific area in the eyepiece triggers the iris scan. The iris image data stream acquired by the sensor, such as a 300-kilobyte (KB) image file, is transmitted to the surgical interaction and verification processor 130 for identity verification. Alternatively, in another embodiment, the biometric sensor 120 can be a dedicated near-field voiceprint acquisition microphone optimized to pick up speech signals with specific frequency ranges and loudness characteristics (e.g., the physician speaking a preset passphrase), using the voiceprint as the basis for verification. This design ensures the convenience, speed, and seamless integration of the signing process with the surgical procedure.

[0029] The surgical interaction and verification processor 130 is a highly integrated computing hub configured to perform a full suite of complex tasks, from audio signal processing to the final generation of signed data blocks. (Refer to...) Figure 2 In one specific embodiment, the processor 130 is implemented as a heterogeneous system-on-a-chip (SoC) 200. This design greatly optimizes data transfer efficiency between modules and overall system power consumption by physically integrating computing units with different functions onto a single chip. The heterogeneous SoC 200 internally includes: a digital signal processor (DSP) core 210, a central processing unit (CPU) core cluster 220, and a security coprocessor 230. These three core units each perform their respective functions and exchange data efficiently and securely through a specially designed state-gated asynchronous data bus (SGAD-Bus) 240.

[0030] The digital signal processor (DSP) core 210 is configured to perform spatial filtering processing. The DSP core 210 receives eight raw audio streams from a multi-channel microphone array 110. An integrated hardware accelerator enables it to perform complex mathematical operations on these data streams with extremely low latency. Exemplarily, the DSP core 210 employs an adaptive beamforming algorithm, such as the Minimum Variance Distortionless Response (MVDR) algorithm, as a specific implementation of spatial filtering. This algorithm uses the phase difference of the signals received by each microphone unit to calculate the spatial location of the sound source and dynamically generates a spatial filter. This filter precisely points the main pickup beam towards the surgeon's mouth while creating nulls in other directions (e.g., the direction of monitor alarms or metallic instrument sounds), thereby suppressing noise and interference by more than 25 dB. After processing, the DSP core 210 outputs a clean, signal-to-noise ratio-enhanced stream of the surgeon's voice in a mono, 48 kHz sampling rate, 24-bit bit-depth pulse-code modulation (PCM) format.

[0031] The central processing unit (CPU) core cluster 220 is the part of the system responsible for running the operating system and performing complex logic processing. In one embodiment, it consists of four high-performance ARM Cortex-A76 cores. The CPU core cluster 220 is configured to perform real-time speech-to-text (ASR) processing and natural language understanding (NLU) processing. It receives a clean speech stream from the DSP core 210 and feeds it into a deep neural network ASR engine optimized for medical terminology. Optimization of the ASR and NLU engines is a crucial step that can be performed by those skilled in the art to ensure high recognition accuracy in highly specialized surgical environments. Exemplarily, this optimization process may include fine-tuning pre-trained acoustic and language models using publicly available, large-scale clinical text corpora (e.g., the MIMIC-IV dataset).

[0032] Furthermore, to handle the specific vocabulary of the operating room, a custom dictionary containing over 8,000 technical terms can be integrated into the ASR engine's decoder. This dictionary covers detailed anatomical structures (such as "hepatic hilum"), surgical instruments (such as "Harmonics knife"), drug names, and common instruction verbs. For the NLU module, at least fifteen classification templates for high-frequency operating room scenarios can be designed for its intent recognition model, such as [instrument request], [drug instruction], [vital sign query], [operation procedure description], and [signature submission], ensuring the system accurately understands each physician's instruction. The ASR engine can convert the speech stream into a sequence of text strings in real time. The NLU module then performs semantic analysis on these text strings. For example, the NLU module classifies and structures the text stream based on a predefined grammar rule and intent recognition model. For instance, the NLU module recognizes the text "freeing the round ligament of the liver" as "process description text." For the text "Please prepare No. 7 thread", the NLU module recognizes it as "executable instruction text" and further extracts structured data, such as {"type": "order", "action": "prepare_suture", "spec": {"size": "7-0"}}. Subsequently, this recognized and structured text, whether process description or instruction, will be processed together to form data entries to be signed. In addition, the NLU module can also recognize special trigger instructions, such as "confirm record," which will be used to initiate the subsequent signing process.

[0033] The security coprocessor 230 is a hardware module in the system responsible for performing security-critical tasks. In one embodiment, it is a physically isolated RISC-V microcontroller with its own independent, encrypted internal memory and a hardware encryption engine (e.g., supporting AES-256 and SHA-256 algorithms). The security coprocessor 230 operates in a physically isolated environment, meaning that the main CPU core cluster 220 cannot directly access its internal programs and data, thus ensuring absolute security of the signing process. It is configured to perform encrypted signing operations. Upon receiving the data to be signed from the CPU core cluster 220 via the SGAD-Bus 240 and the biometric data from the biometric sensor 120, it first performs authentication. For example, it compares real-time acquired iris image data with an iris template pre-stored in its secure memory. This comparison algorithm, such as a Hamming distance calculation, can be completed in less than 100 milliseconds (ms). Only after successful verification (e.g., the Hamming distance is less than a preset threshold of 0.32) will it proceed with the signing operation. The signing process includes: obtaining a high-precision timestamp, serializing the timestamp, verifier identity, and data to be signed, and then using the SHA-256 algorithm to calculate a 256-bit hash value as a digital digest. This process ensures the integrity, authenticity, and non-repudiation of the record.

[0034] A key innovation of the heterogeneous system-on-a-chip 200 is its integrated State-Gated Asynchronous Data Bus (SGAD-Bus) 240. This bus is specifically designed to resolve the conflict between real-time data stream processing and transactional safety operations. It is configured to connect the CPU core cluster 220 and the security coprocessor 230, and operates in two mutually exclusive operating states.

[0035] The first state is "transcription state". In this state, the SGAD-Bus 240 functions as a standard high-speed AXI bus, allowing the CPU core cluster 220 to freely and with low latency write structured data entries generated by the NLU module into a shared 32-kilobyte (KB) temporary buffer. This data is then synchronously sent to the display unit 140 for presentation.

[0036] The second type is the "commit state". When the system detects a commit trigger event (e.g., the NLU recognizes a "acknowledge record" instruction), a hardware interrupt signal triggers a state switch for the SGAD-Bus 240. Upon entering the commit state, the bus controller physically disconnects the CPU core cluster 220 from the temporary buffer via a set of hardware logic gates, achieving an instantaneous "freeze" of the data. Immediately afterwards, the bus controller performs a one-time, non-interruptible DMA (Direct Memory Access) transfer, completely transferring the entire contents of the temporary buffer—for example, a data block containing five records and a total size of 2.1KB—as an atomic data snapshot to the internal secure memory of the secure coprocessor 230.

[0037] This process is guaranteed to be atomic by hardware and takes less than 1 microsecond (µs). After the transmission is complete, the bus automatically switches back to the transcription state, and the CPU core cluster 220 can continue to process new voice input. This design, through hardware-level state switching, perfectly coordinates the continuity of the data flow and the atomicity of transaction operations. To enable those skilled in the art to implement this without difficulty, a feasible implementation of the state-gated asynchronous data bus 240 is described. The controller of the state-gated asynchronous data bus 240 can be implemented by a finite state machine (FSM). The FSM includes at least a 'transcription state' and a 'commit state'.

[0038] When the CPU core cluster 220 detects a commit trigger event, it sends an interrupt signal to the bus controller via a dedicated interrupt line. In response to this interrupt signal, the FSM transitions from the 'transcription state' to the 'commit state'. In this state, the FSM's output logic directly acts on the bus arbitrator of the shared temporary buffer, forcibly asserting a write lock signal, thereby physically preventing any write requests from the CPU core cluster 220. Simultaneously, the FSM issues a single-block transfer instruction to an integrated Direct Memory Access (DMA) controller, containing the base address of the temporary buffer and the data length. After the DMA transfer is complete, the DMA controller sends a completion signal back to the FSM, which then automatically resets to the 'transcription state' and releases the write lock. This hardware-based FSM and DMA implementation ensures the atomicity and uninterruptibility of state transitions and data snapshot transfers.

[0039] Reference Figure 3 This embodiment also provides a surgical procedure recording and command interaction method based on multimodal biometric verification, which is executed by the aforementioned system.

[0040] S100: Perform spatial filtering to process the multiple raw audio streams, thereby generating a single high signal-to-noise ratio physician speech stream.

[0041] This step is performed by the digital signal processor (DSP) core 210 within the surgical interaction and verification processor 130. After system startup, the multi-channel microphone array 110 deployed in the operating room begins to continuously collect ambient sound. Assuming the array consists of 8 microphone units, at time t, the system will acquire 8 parallel digital audio signal samples. These signals are a mixture of the surgeon's voice, conversations with other personnel, instrument noise, and ambient reverberation. The DSP core 210 receives these eight signals and applies a spatial filtering algorithm. As a preferred implementation, this algorithm can be adaptive beamforming. The algorithm first estimates the direction of arrival (DOA) of the surgeon's sound source by calculating the cross-correlation function between the signals. For example, it calculates the azimuth angle of the sound source relative to the array normal. Pitch angle The algorithm then uses this DOA information to calculate a complex weight for each signal. The goal of these weights is to minimize the total power of the output signal (i.e., suppress noise and interference from other directions) while maintaining a gain of 1 (no distortion) for the signal in the target direction (i.e., the doctor's speech). Finally, the DSP core 210 sums the weighted signals to obtain the clean output speech signal. ,in yes The conjugate of . This output signal This results in a high signal-to-noise ratio physician speech stream, with background noise levels reduced by approximately 25 dB compared to the original signal, significantly improving the accuracy of subsequent speech recognition. This process is performed in real-time, frame-by-frame, ensuring continuous and clear capture of the physician's speech.

[0042] S200: Perform real-time speech-to-text processing and natural language understanding processing on the high signal-to-noise ratio physician speech stream to distinguish the physician speech stream into text containing process descriptions and text containing executable instructions, and process the aforementioned text to form one or more data entries to be signed.

[0043] This step is performed by the central processing unit (CPU) core cluster 220 within the surgical interaction and verification processor 130. The CPU core cluster 220 receives a clean audio stream output from the DSP core 210. First, a deep neural network speech-to-text (ASR) engine optimized for medical terminology processes the speech stream. This ASR model, such as an end-to-end model integrating a convolutional neural network (CNN) and a long short-term memory network (LSTM), converts continuous audio waveforms into discrete sequences of text strings. For example, for a 3.5-second speech input, the ASR engine might output the text: “Isolating the trigone of the gallbladder, note the cystic duct and cystic artery.” Subsequently, a natural language understanding (NLU) module performs real-time analysis of the ASR output text. The NLU module employs a hybrid strategy, combining rule-based pattern matching and machine learning-based intent classification.

[0044] For example, for the text "Dissecting the gallbladder triangle...", the NLU module categorizes it as "process description text" based on keywords such as "dissecting" and "separating". For another text converted from speech input, "Give me a non-abrasive gripper", the NLU module first identifies the verb "give" and the noun "gripper", categorizing it as an "instrument request" intent within "executable instruction text". Regardless of the text type, the system processes and formats it. For instance, process description text is recorded as is, while instruction text is parsed into a structured JSON object: {"timestamp": "2025-11-20T15:10:22.450Z", "type": "instrumentrequest", "details": {"name": "atraumaticgrasper", "quantity": 1}}. All these processed texts and JSON objects are temporarily stored as data entries to be signed in a temporary buffer managed by the SGAD-Bus 240 and simultaneously sent to the display unit 140 for real-time preview by the surgical team.

[0045] S300: In response to a submission trigger event, associate the one or more data entries to be signed with the biometric data collected by the biometric sensor, and perform a cryptographic signing operation to generate a timestamp-containing, tamper-proof signed data block.

[0046] This step is completed collaboratively by the CPU core cluster 220, SGAD-Bus 240, biometric sensor 120, and security coprocessor 230. A submission trigger event is a prerequisite for initiating this step. For example, after the surgeon completes a stage of the procedure, they verbally utter the command "Confirm this stage of the procedure." The NLU module on the CPU core cluster 220 recognizes this pre-defined submission trigger voice command, which constitutes the first submission trigger event. In response, the CPU core cluster 220 immediately sends a hardware interrupt signal to the SGAD-Bus 240 controller. The state of the SGAD-Bus 240 instantly switches from "transcription state" to "submission state."

[0047] In this state, the CPU core cluster 220's write access to the temporary cache is hardware locked. The bus controller performs a one-time DMA operation, transferring all unconfirmed data entries in the temporary cache (e.g., an array containing the aforementioned "process description" text and "instrument request" JSON object) as an indivisible snapshot to the internal secure memory of the security coprocessor 230. Almost simultaneously, the biometric sensor 120 (e.g., an iris scanner), receiving instructions from the CPU, is activated. It acquires an image of the physician's iris and sends the image data (e.g., a 300KB JPEG file) to the security coprocessor 230. Within its physically isolated environment, the security coprocessor 230 first compares the received iris image with a physician iris template pre-stored in its secure flash memory. The comparison algorithm calculates a Hamming distance of 0.28, which is less than the preset matching threshold of 0.32, thus authentication is successful. Upon successful verification, the security coprocessor 230 immediately retrieves a high-precision timestamp (e.g., 2025-11-20T15:12:01.888Z) from its internal real-time clock. It then serializes this timestamp, the successfully verified physician's identity identifier (e.g., Dr. Smith_UID), and a snapshot of data received from the bus, forming a contiguous data block. Finally, it uses its built-in hardware SHA-256 engine to calculate a 256-bit hash value for this data block. This hash value, timestamp, identity identifier, and original data snapshot together constitute a signed data block. This data block is then securely transmitted to the hospital's electronic medical record system for archiving via the EMR system interface 150, completing an instant, secure, and non-repudiable digital signature process.

[0048] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0049] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware.

[0050] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A surgical procedure recording and command interaction system based on multimodal biometric verification, characterized in that, include: A multi-channel microphone array is configured to capture multiple raw audio streams within the operating room environment; Biometric sensors are configured to collect biometric data from the surgeon. In addition, a surgical interaction and verification processor is configured as follows: Spatial filtering is performed on the multiple original audio streams to generate a single high signal-to-noise ratio physician speech stream; The high signal-to-noise ratio physician speech stream is subjected to real-time speech-to-text processing and natural language understanding processing to distinguish the physician speech stream into text containing process descriptions and text containing executable instructions, and the aforementioned text is processed to form one or more data entries to be signed. In response to a submission trigger event, the one or more data entries to be signed are associated with the biometric data collected by the biometric sensor, and an cryptographic signing operation is performed to generate a timestamped, tamper-proof signed data block.

2. The system according to claim 1, characterized in that, The surgical interaction and verification processor is a heterogeneous system-on-a-chip, which physically integrates: A digital signal processor core is configured to perform the spatial filtering process; A central processing unit core cluster is configured to perform the real-time speech-to-text processing and the natural language understanding processing; In addition, a security coprocessor is configured to perform the cryptographic signing operation in a physically isolated environment.

3. The system according to claim 2, characterized in that, The heterogeneous system-on-a-chip also integrates a state-gated asynchronous data bus, which is configured to connect the central processing unit core cluster and the security coprocessor.

4. The system according to claim 3, characterized in that, The state-gated asynchronous data bus is configured to operate in two mutually exclusive working states: A transcription state in which the central processing unit core cluster is allowed to write the one or more data entries to be signed into a shared temporary buffer; In addition, in a commit state, write access to the temporary buffer is locked by hardware logic, and a one-time, uninterrupted data block transfer is performed to transmit the entire contents of the temporary buffer as a data snapshot to the security coprocessor.

5. The system according to claim 1, characterized in that, The submission trigger event includes at least one of the following: The surgical interaction and verification processor's natural language understanding process recognizes a preset submission trigger voice command; Alternatively, the biometric sensor may be activated by a specific physical action of the surgeon.

6. The system according to claim 1, characterized in that, The biometric sensor is an iris scanner or a acoustic fingerprint acquisition microphone.

7. The system according to claim 2, characterized in that, The security coprocessor is further configured as follows: Before performing the encryption signature operation, the biometric data collected in real time is compared and verified with a pre-stored biometric template. Furthermore, the cryptographic signature operation will continue only if the comparison and verification are successful.

8. The system according to claim 1, characterized in that, The signed data block includes a data content portion and an encrypted signature portion; wherein the data content portion includes the one or more data entries to be signed, the timestamp, and a verifier identity identifier; the encrypted signature portion is generated by applying an encrypted hash algorithm to the data content portion.

9. The system according to claim 1, characterized in that, The system also includes: A display unit is configured to present, in real time, the text converted from the physician's voice stream and the one or more data entries to be signed, before the cryptographic signing operation is performed, for the attending physician to confirm immediately.

10. A surgical procedure recording and command interaction system based on multimodal biometric verification, characterized in that, include: A spatial filtering module is configured to generate a high signal-to-noise ratio physician speech stream based on multiple raw audio streams acquired from a multi-channel microphone array. A semantic understanding and structuring module is configured to convert the high signal-to-noise ratio physician speech stream into structured data entries, wherein the structured data entries are divided into process descriptive text and executable instruction text; Additionally, a secure signature and solidification module is configured to, in response to a submission trigger event, perform an encrypted signature operation on the structured data entry based on the biometric data of the attending physician collected from a biometric sensor, to generate a timestamped, tamper-proof signed data block.