AI processing-based motion capture recognition pre-judgment system

By integrating visual capture, electromyography, and bioacoustic acquisition modules, the AI ​​processing system solves the problem of real-time perception of operator intent and target object status in complex human-computer interaction scenarios, achieving advanced prediction and adaptive calibration, and improving the system's reliability and efficiency.

CN121370144APending Publication Date: 2026-01-23昆山牙博士口腔门诊部有限公司 +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511505120.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies struggle to anticipate operator intentions and perceive target object states in complex human-computer interaction scenarios, and lack collaborative processing and adaptive calibration of multi-source heterogeneous information, resulting in low system reliability and efficiency.

Method used

An AI-based motion capture and recognition prediction system is adopted, which integrates a visual capture module, an electromyography signal acquisition module, a bioacoustic acquisition module, and an information prompting module. The system uses an AI intelligent processor to fuse the operator's intention with the target object's state, generate composite collaborative instructions, and perform adaptive noise suppression and model calibration.

Benefits of technology

It enables proactive prediction of operator intentions, improves the long-term accuracy and reliability of the system, ensures the accuracy of physiological state monitoring and the reliability of risk warning, and enhances the efficiency and safety of human-machine collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121370144A_ABST
    Figure CN121370144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine collaboration and information processing, discloses an AI processing-based motion capture recognition pre-judgment system, and aims to solve the problems of difficulty in physiological signal monitoring and performance drifting of a myoelectricity intention recognition model in a high-noise environment, and the AI processing-based motion capture recognition pre-judgment system comprises a processing module, a visual capture module, a myoelectricity acquisition module, a bioacoustics acquisition module and an information prompting module, according to the method, myoelectricity and visual information decoding operation intentions are fused, a self-adaptive filter is guided through the intentions, accurate noise suppression is carried out on biological acoustic signals of a target object, and therefore the physiological state risk of the target object is accurately evaluated, meanwhile, an online self-adaptive calibration mechanism is introduced, the actual action of visual recognition serves as a supervision true value, and the accuracy of the physiological state risk of the target object is improved. And the myoelectricity model is continuously calibrated, so that the accuracy of long-term identification is ensured. According to the method, the decoding intention and the physiological risk level are integrated, the composite cooperation instruction is intelligently generated and fed back to the operator, the situation awareness ability of the operator is enhanced, and the safety and efficiency of man-machine cooperation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-machine collaboration and information processing technology, and in particular to a motion capture, recognition and prediction system based on AI processing. Background Technology

[0002] In complex human-computer interaction scenarios such as modern surgery, precision manipulation, or high-risk operations, operator intent recognition and controlled object status monitoring are two core aspects that ensure task safety and efficiency. Existing technologies typically utilize bioelectrical signals such as electromyography (EMG) to decode the operator's intentions and monitor the physiological state of the target object (such as a patient during surgery) through independent sensing systems. However, these technologies face several challenges in practical applications.

[0003] First, in environments such as operating rooms, various medical devices (such as high-frequency electrosurgical units, ultrasonic bone scalpels, and suction devices) generate strong and dynamically changing acoustic noise. This environmental noise can easily drown out the weak bioacoustic signals (such as heart sounds and lung sounds) collected from the surgical subject, resulting in poor effectiveness of traditional signal noise reduction methods. It is difficult to achieve continuous and reliable assessment of the physiological state, making it impossible for operators to obtain timely and accurate feedback on the impact of their operations on the physiological state.

[0004] Secondly, the performance of systems that rely on electromyography (EMG) signals for intention recognition is significantly affected by changes in the operator's physiological state. During long-term, high-intensity tasks, operators inevitably experience muscle fatigue, and factors such as sweating can also alter the conductivity of the skin surface. These factors can cause the EMG signal pattern to drift. For recognition models that are trained offline and deployed statically, this change in signal pattern will cause their recognition accuracy to gradually decrease over time, resulting in reduced system reliability.

[0005] Finally, current human-machine collaborative systems often present the operator's intention information and the target object's status information as two separate channels. The operator needs to mentally connect the next operation with the target object's real-time status, which undoubtedly increases their cognitive load. This separate presentation of information lacks deep integration and intelligent prediction of the inherent correlation between operational intention and status risk, and cannot provide the operator with a composite collaborative instruction that integrates action guidance and risk warning, thus limiting the efficiency and safety of human-machine collaboration. Summary of the Invention

[0006] The purpose of this invention is to provide an AI-based motion capture recognition and prediction system, which solves the problems of existing auxiliary systems in terms of information acquisition and processing, difficulty in simultaneously achieving advance prediction of the operator's intention and real-time perception of the state of the object being acted upon, and lack of collaborative processing and adaptive calibration capabilities for multi-source heterogeneous information.

[0007] To address the aforementioned technical problems, this invention provides an AI-based motion capture, recognition, and prediction system.

[0008] The system includes: a visual capture module, an electromyography (EMG) signal acquisition module, a bioacoustic acquisition module, an information prompting module, and an AI intelligent processor. The AI ​​intelligent processor is communicatively connected to the visual capture module, the EMG signal acquisition module, the bioacoustic acquisition module, and the information prompting module.

[0009] The visual capture module is used to collect the operator's visual motion data in real time; The electromyography (EMG) signal acquisition module is used to acquire EMG signals of the operator's limbs in real time. The bioacoustic acquisition module is used to acquire the heart sounds or respiratory sounds of the surgical subject in real time as bioacoustic signals. The information prompt module is used to output prompt information to the user; The AI ​​intelligent processor is configured to perform the following operations: Based on the visual motion data, visual features and the operator's actual actions are identified. Based on the electromyographic signals, the operator's pre-operational intention is decoded; Based on the bioacoustic signals, the real-time physiological state of the surgical subject was analyzed. By integrating the pre-operation intention with the real-time physiological state, a composite collaborative instruction is generated; The information prompting module is controlled to output the prompting information according to the composite collaborative instruction.

[0010] In one specific embodiment, the composite cooperative instruction includes at least one of the following: Action instructions are used to specify the consumables that need to be prepared or the auxiliary operations that need to be performed. Priority instructions are used to characterize the urgency of the action instructions; Risk warning instructions are used to characterize risk information associated with the real-time physiological state.

[0011] In one specific embodiment, the process by which the AI ​​intelligent processor obtains the operator's pre-operation intention based on the electromyographic signal decoding specifically includes: First, an electromyographic intention recognition model is pre-constructed by collecting electromyographic signal samples S corresponding to the operator performing a preset set of standard movements. emg ; Extract features from the electromyographic signal samples to form a feature vector F emg This process can be represented by the following formula: F emg =f extract.emg (S emg ); Among them, f extract.emg This is a preset electromyography signal feature extraction function; Furthermore, by utilizing the correspondence between the feature vector and the preset standard action, a classifier is trained to obtain the electromyographic intention recognition model.

[0012] Then, during system operation, the electromyographic intention recognition model is used to process the electromyographic signals acquired in real time to output the corresponding pre-operation intention.

[0013] In one specific embodiment, the process by which the AI ​​intelligent processor obtains the real-time physiological state of the surgical subject based on the bioacoustic signal analysis specifically includes: First, a physiological acoustic baseline model is constructed by acquiring the bioacoustic signal S of the surgical subject in a resting state. bio.rest ; The acoustic features of the bioacoustic signals in the resting state are extracted to form a baseline feature vector F. baseline This process can be represented by the following formula: F baseline =f extract.acoustic (S bio.rest ); Among them, f extract.acoustic This is a preset acoustic feature extraction function; And, the baseline feature vector F baseline Stored as the physiological acoustic baseline model.

[0014] Then, during system operation, the current acoustic features F extracted from the real-time acquired bioacoustic signals are used. current.acoustic The physiological acoustic baseline model is compared to assess and determine the real-time physiological state of the surgical subject.

[0015] In a preferred embodiment, the identification and prediction system further includes: an environmental acoustic acquisition module for acquiring environmental noise signals, and the AI ​​intelligent processor is further configured to perform the following operations: According to the decoded pre-operation intent I pre (t), determine the instrument noise model N associated with the pre-operation intention.instr Using the aforementioned instrument noise model, the bioacoustic signal S was analyzed. bio (t) and the environmental noise signal S env The mixed signal (t) is subjected to adaptive filtering noise suppression processing, and the signal S′ obtained after the noise suppression processing is... bio As the bioacoustic signal used for the analysis of the real-time physiological state, the processing procedure can be represented by the following formula: S′ bio (t)=h filter (S bio (t)+S env (t),N instr (I pre (t))); Among them, h filter This is a preset adaptive filtering function.

[0016] In one specific embodiment, the AI ​​intelligent processor is further configured to: The visual feature F identified from the visual motion data vis In (t), extract operation-related features; And utilize the visual features to decode the pre-operation intention I obtained from the electromyographic signals. pre (t) performs fusion correction to update the pre-operational intent, and this update process can be performed by a multimodal fusion model h. fusion express: I final (t)=h fusion (I pre (t),F vis (t)); Among them, I final (t) represents the updated pre-operation intent.

[0017] In a preferred embodiment, the AI ​​intelligent processor is further configured to perform the following operations: Extract the operator's actual action A based on the visual motion data. vis (t); The pre-operation intention I final (t) and the actual action A vis (t) Perform a consistency check; When the conflict measurement between the pre-operation intention and the actual executed action meets the preset calibration conditions, the set D containing conflict data pairs is used. set The parameter θ of the electromyographic intention recognition model is calibrated online, and the calibration process can be represented by the following formula: θ new =hupdate (θ old D set ); Among them, h update For the parameter update function, θ old For the model parameters before calibration, θ new These are the calibrated model parameters.

[0018] Specifically, the visual capture module may include: augmented reality glasses worn on the user's head, the augmented reality glasses being equipped with a first camera, and an environmental camera deployed in the surgical environment.

[0019] Specifically, the electromyography signal acquisition module can be a wearable surface electromyography sensing armband or wristband, and the bioacoustic acquisition module can be a contact bioacoustic sensor.

[0020] Preferably, the information prompting module may include: a semi-transparent display screen disposed on the augmented reality glasses, and an in-ear prompter communicatively connected to the augmented reality glasses.

[0021] In summary, the present invention has at least one of the following beneficial technical effects: 1. This invention, by setting up an electromyography (EMG) signal acquisition module, processes the operator's EMG signals to decode the pre-operational intent, and before the operator actually performs the action, integrates real-time physiological state to intelligently generate composite collaborative instructions, which are output through an information prompt module. This achieves advanced prediction of the operator's intent, enabling assistants to prepare instruments or perform corresponding operations in advance, shortening the waiting and reaction time in the surgical procedure, and improving the smoothness and efficiency of the overall collaborative work.

[0022] 2. This invention, by setting up a visual capture module, is not only used to identify visual features for fusion correction of the pre-operational intent of electromyography decoding, but also to identify the operator's actual execution action. By comparing the consistency between the pre-operational intent and the actual execution action, when the conflict between the two meets the preset conditions, the system can perform online calibration of the electromyographic intent recognition model, which solves the problem of electromyographic signal pattern drift caused by factors such as operator fatigue, and ensures the long-term accuracy and reliability of the system's intent recognition.

[0023] 3. This invention acquires the real-time physiological state of the surgical subject by setting up a bioacoustic acquisition module and acquires environmental noise by using an environmental acoustic acquisition module. By using the decoded pre-operation intention, it predicts the instruments to be used and determines the corresponding instrument noise model. Then, it performs targeted noise suppression processing on the mixed acoustic signals, effectively overcoming the interference of instrument noise in the surgical environment on weak physiological acoustic signals, ensuring the accuracy of physiological state monitoring, and thus improving the reliability of the system for risk warning. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the architecture of the AI-based motion capture, recognition, and prediction system according to an embodiment of the present invention. Figure 2 This is a schematic diagram of an acoustic signal enhancement mechanism guided by intent according to an embodiment of the present invention; Figure 3 Flowchart of the online adaptive calibration mechanism in an embodiment of the present invention. Detailed Implementation

[0025] Please refer to the appendix. Figure 1 , Figure 1 This is a schematic diagram of the architecture of an AI-based motion capture, recognition, and prediction system according to an embodiment of the present invention. The present invention provides an AI-based motion capture, recognition, and prediction system, which may include: an AI intelligent processor 10, a visual capture module 20, an electromyography signal acquisition module 30, a bioacoustic acquisition module 40, and an information prompting module 50.

[0026] In a preferred embodiment, the AI ​​intelligent processor 10 may be an embedded computing platform integrating a high-performance graphics processing unit (GPU), such as NVIDIA's Jetson series, to meet the computing power requirements for real-time processing of multimodal data streams and inference of neural network models. Alternatively, the AI ​​intelligent processor 10 may also be a high-performance personal computer or workstation with wireless communication capabilities.

[0027] The AI ​​intelligent processor 10 is the computing and control unit of this system, and can be implemented by one or more microprocessors, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). The AI ​​intelligent processor 10 contains internal memory for storing program instructions and data during processing.

[0028] The AI ​​intelligent processor 10 establishes data communication connections with the visual capture module 20, the electromyography signal acquisition module 30, the bioacoustic acquisition module 40, and the information prompting module 50 through wired or wireless communication interfaces, respectively, to receive data from each acquisition module and send instructions to the information prompting module.

[0029] The visual capture module 20 is used to collect visual motion data of the operator and the surgical environment. In one specific embodiment, the visual capture module 20 includes a first camera 21 and an environmental camera 22.

[0030] The first camera 21 is integrated into augmented reality glasses worn by a user (e.g., an assistant) to acquire a first-person view video stream. An environmental camera 22 is deployed at a predetermined location in the surgical environment, such as above the ceiling or the side of the operating table, to acquire a global or specific view video stream of the surgical area. Both the first camera 21 and the environmental camera 22 transmit their acquired video data streams to the AI ​​intelligent processor 10 in real time.

[0031] Specifically, the first camera 21 can be a miniature camera integrated into augmented reality (AR) glasses, with a resolution of at least 1080p and a frame rate of at least 60fps to ensure clear capture of the operator's fine hand movements. The environmental camera 22 can be one or more wide-angle or depth cameras used to acquire three-dimensional information of the scene to help determine the spatial relative position of the operator and the surgical subject.

[0032] The electromyography (EMG) signal acquisition module 30 is used to acquire EMG signals from the operator's limbs.

[0033] In a preferred embodiment, the electromyography (EMG) signal acquisition module 30 can be a dry electrode array armband containing at least eight channels, which performs signal acquisition differentially with a sampling frequency set to no less than 1000 Hz to capture sufficiently rich EMG signal details. Using dry electrodes avoids the use of conductive gel, improving the operator's ease of wear and long-term comfort.

[0034] When in use, it is worn on the forearm of the limb where the operator (e.g., the surgeon) performs the main operation, and collects multi-channel surface electromyography (sEMG) signals generated by muscle activity, and transmits the digitized signal data wirelessly to the AI ​​intelligent processor 10 in real time.

[0035] The bioacoustic acquisition module 40 is used to acquire the heart sounds or respiratory sounds of the surgical subject as bioacoustic signals.

[0036] In one specific embodiment, the module is a contact bioacoustic sensor, such as an electronic stethoscope probe, which is fixed to a specific location on the surface of the surgical subject's body with medical adhesive, such as the third or fourth intercostal space on the left sternal border, to collect clear heart sounds and breath sounds, and transmit the collected acoustic signal data to the AI ​​intelligent processor 10.

[0037] In a preferred embodiment, the system further includes an ambient acoustic acquisition module 60. This module consists of one or more microphones deployed in the surgical environment to acquire ambient noise signals, particularly the sounds generated when surgical instruments are in operation. The ambient acoustic acquisition module 60 transmits the acquired acoustic signal data to the AI ​​intelligent processor 10 for subsequent noise suppression processing.

[0038] The information prompt module 50 is used to output prompt information to the user. In one specific embodiment, the information prompt module 50 is integrated with the aforementioned augmented reality glasses.

[0039] The information prompting module 50 includes a semi-transparent display screen 51 mounted on the augmented reality glasses and an in-ear prompter 52 that is communicatively connected to the augmented reality glasses. The semi-transparent display screen 51 is used to overlay text or graphic instructions in the user's field of vision, and the in-ear prompter 52 is used to play voice instructions or prompts to the user. The information prompting module 50 receives instructions from the AI ​​intelligent processor 10 and converts them into corresponding visual or auditory information for output.

[0040] During system operation, the visual capture module 20, electromyography signal acquisition module 30, bioacoustic acquisition module 40, and environmental acoustic acquisition module 60 concurrently and continuously acquire their respective data and send them to the AI ​​intelligent processor 10. The AI ​​intelligent processor 10 performs a series of analyses, decodings, fusions, and decision-making processes on the received multimodal data, and finally generates instructions to drive the information prompt module 50 to provide users with operation guidance and risk warnings.

[0041] Before the system enters real-time working state, the AI ​​intelligent processor 10 performs an initialization phase, which mainly includes building an electromyographic intention recognition model for a specific operator and building a physiological acoustic baseline model for a specific surgical subject.

[0042] To build an electromyography intention recognition model, the system first defines a set of preset standard actions related to the application scenario. For example, the set may include a series of basic operation actions such as holding a scalpel, picking up a cotton ball, and rotating the wrist to suture.

[0043] Subsequently, under the control of the AI ​​intelligent processor 10, instructions are presented to the operator in sequence through an external display device, guiding the operator to repeatedly execute each standard action in the above set.

[0044] During this process, the electromyography signal acquisition module 30 synchronously and continuously acquires multi-channel surface electromyography signals from the operator's forearm, and associates these signal segments with the corresponding standard action labels to form the original training dataset.

[0045] After receiving the original training dataset, the AI ​​intelligent processor 10 first preprocesses the electromyographic signal samples. The preprocessing includes: applying a bandpass filter, such as a Butterworth filter, with a passband range of 20Hz to 500Hz, to filter out motion artifacts and high-frequency noise; and applying a 50Hz notch filter to eliminate power frequency interference.

[0046] After preprocessing, the AI ​​intelligent processor 10 performs windowing processing on the signal data, for example, using a sliding window of 200 milliseconds with a step size of 50 milliseconds. Within each window, features that characterize the muscle activity state are extracted to form a feature vector. These features may include, but are not limited to, time-domain features such as integrated electromyography (iEMG), root mean square (RMS), and zero-crossing point (ZC); and frequency-domain features such as average power frequency (MPF) and median frequency (MF). By combining multiple features, the intrinsic patterns of electromyographic signals can be more comprehensively characterized.

[0047] For an electromyography (EMG) signal acquisition module containing C channels, the feature vector F extracted within the time window t is... emg (t) can be expressed as: F emg (t)=[f 1,1 ,...,f 1,k ,...,f C,1 ,...,f C,k ] T ; Among them, f C,i represents the i-th feature extracted from the C-th channel, and k is the total number of features extracted for each channel.

[0048] Finally, the AI ​​intelligent processor 10 uses the feature vectors extracted from all time windows and their corresponding standard action labels to construct a fully labeled training set. Based on this training set, a classifier algorithm is used for training to obtain an electromyography intention recognition model.

[0049] The classifier can be a support vector machine (SVM) or a variant of a recurrent neural network (RNN), such as a long short-term memory network (LSTM).

[0050] After training, the electromyography intention recognition model is stored in the memory of the AI ​​intelligent processor 10 for subsequent real-time intention decoding.

[0051] To construct a physiological acoustic baseline model, the system acquires a bioacoustic signal S for a duration (e.g., 60 seconds) within a preset stable physiological cycle (e.g., after the surgical subject is under anesthesia and before the surgery officially begins). bio.rest (t), the AI ​​intelligent processor 10 preprocesses the bioacoustic signal segment, including filtering to retain the main frequency components of heart sounds and breath sounds.

[0052] Subsequently, the AI ​​intelligent processor 10 extracts a set of acoustic features from the preprocessed signal that characterize the basic physiological state of the surgical subject. These acoustic features may include: heart rate (HR) and S1-S2 time interval calculated from heart sound signals; respiratory rate (RR) calculated from breath sound signals; and the energy spectrum distribution characteristics of the signal. These features collectively constitute a multidimensional baseline feature vector.

[0053] The AI ​​intelligent processor 10 performs statistical analysis on multiple baseline feature vectors calculated within the stable period to obtain their mean vector μ. baseline The covariance matrix Σ baseline .

[0054] μ baseline =E[F baseline ]; Σ baseline =E[(F baseline -μ baseline (F) baseline -μ baseline ) T ]; Among them, F baseline From signal segment S bio.rest Any baseline feature vector extracted from (t), where E[·] represents the expected computation.

[0055] The mean vector μ baseline The covariance matrix Σ baseline It is stored together as the physiological acoustic baseline model of the surgical subject for subsequent comparison and evaluation of real-time physiological status.

[0056] Please refer to the appendix. Figure 2 , Figure 2 This is a schematic diagram of an acoustic signal enhancement mechanism intended to be guided according to an embodiment of the present invention.

[0057] After the initialization phase is completed, the system enters the real-time collaborative workflow. The AI ​​intelligent processor 10 starts all acquisition modules, including the visual capture module 20, the electromyography signal acquisition module 30, the bioacoustic acquisition module 40, and the environmental acoustic acquisition module 60.

[0058] Each module begins to continuously collect its own data using a preset sampling frequency and a synchronous clock signal, and transmits the data stream to the AI ​​intelligent processor 10 in real time.

[0059] To ensure precise time alignment of multiple data streams, the AI ​​intelligent processor 10 operates an internal master clock and performs time synchronization calibration on all acquisition modules via Network Time Protocol (NTP) or a dedicated hardware synchronization signal. Each acquisition module adds a high-precision timestamp when generating data packets. The AI ​​intelligent processor 10 has an independent buffer queue at the receiving end, which sorts and aligns data packets from different channels according to the timestamps to generate time-synchronized multimodal data frames for subsequent algorithm processing.

[0060] After receiving the multimodal data stream, the AI ​​intelligent processor 10 performs an intent-guided acoustic signal enhancement process.

[0061] Specifically, the AI ​​intelligent processor 10 first performs a preliminary, low-latency decoding of the electromyographic signal stream from the electromyographic signal acquisition module 30 to obtain an immediate pre-operation intention I. pre (t).

[0062] Meanwhile, the AI ​​intelligent processor 10 receives a mixed signal d(t) from the bioacoustic acquisition module 40 and a reference noise signal x(t) from the environmental acoustic acquisition module 60.

[0063] The mixed signal d(t) here can be modeled as the target bioacoustic signal S. bio (t) and the environmental noise signal S after propagation through a specific path env.filtered Linear superposition of (t).

[0064] The AI ​​intelligent processor 10 internally stores a library of instrument noise models, which pre-stores characteristic models of acoustic noise generated by various surgical instruments (e.g., high-frequency electrosurgical units, ultrasonic bone scalpels) during operation.

[0065] The AI ​​intelligent processor 10 utilizes the pre-operation intent I obtained from the aforementioned real-time decoding. pre (t) serves as an index to query and load the instrument noise model N associated with this intent from the database. instr The model N instr Parameters used to configure an adaptive filter.

[0066] The AI ​​intelligent processor 10 employs an adaptive filter to process the mixed signal d(t) to eliminate noise components. In one specific embodiment, this filter is a finite impulse response (FIR) filter using the least mean square (LMS) algorithm. The goal of this filter is to estimate the noise component y(t) in the mixed signal d(t) using a reference noise signal x(t), and to use the difference between the two as an estimate of the target bioacoustic signal. This process is defined by the following formula: y(t)=W(t) TX(t); e(t) = d(t) - y(t); W(t+1)=W(t)+2μe(t)X(t); Where: t is the discrete-time index; X(t) = [x(t), x(t-1), ..., x(t-L+1)] T The input vector of the reference noise signal x(t) at time t is L, where L is the order of the filter; W(t) = [w0(t), w1(t), ..., w L-1 (t)] T y(t) is the filter's weight vector at time t; y(t) is the filter's output at time t, i.e., the estimate of the noise component in the mixed signal; e(t) is the error signal, which theoretically is the target bioacoustic signal S. bio (t) is an estimate of the convergence rate and stability; μ is a step size factor that controls the convergence rate and stability.

[0067] Instrument noise model N loaded from the instrument noise model library instr , is used to set the initial weight vector W(0) and step size factor μ of the adaptive filter, so that the filter can converge quickly for specific types of instrument noise.

[0068] After the above adaptive filtering process, the output error signal e(t) is the noise-suppressed bioacoustic signal S′. bio (t), this signal, along with the video stream from the visual capture module 20 and the electromyographic signal stream from the electromyographic signal acquisition module 30, is sent to the subsequent processing flow.

[0069] In subsequent processing, the AI ​​intelligent processor 10 performs a fusion decision step to generate a more reliable final operational intent.

[0070] Specifically, the AI ​​intelligent processor 10 not only uses the electromyography intention recognition model to obtain the preliminary pre-operation intention, but also extracts operation-related visual features from the video stream of the visual capture module 20, such as the operator's hand posture, the type and orientation of the instrument being held.

[0071] Subsequently, a weighted decision or confidence fusion strategy is adopted, for example, setting a dynamic confidence score for the electromyographic intention recognition model and the visual action recognition model respectively.

[0072] When visual information is clear and can identify a specific action with high confidence (e.g., clearly capturing the action of holding a scalpel), its weight will be increased accordingly, and the recognition result will be used as the final operational intention.

[0073] When visual information is blurry or obscured, but the electromyographic signal pattern is clear and stable, the results of electromyographic recognition are relied upon more heavily.

[0074] If both sides have equal confidence levels, arbitration can be performed using preset rules, or a pending state can be output.

[0075] This fusion step effectively combines the predictive power of electromyography (EMG) signals with the accuracy of visual information to output a more robust final operational intent for subsequent risk assessment and instruction generation.

[0076] While determining the final operational intent, the AI ​​intelligent processor 10 performs a real-time physiological risk assessment on the noise-suppressed bioacoustic signal. The AI ​​intelligent processor 10 extracts the acoustic feature vector x within a sliding time window of the real-time signal stream in the same manner as when building the baseline model. realtime .

[0077] Subsequently, the deviation from the physiological state was quantified by calculating the Mahalanobis distance between the real-time feature vector and the physiological acoustic baseline model. The formula for calculating the Mahalanobis distance is as follows: Where: D M (x realtime ) is the calculated Mahalanobis distance, representing the difference between the current state and the baseline stable state; x realtime It is the acoustic feature vector calculated in real time; μ base and Σ base These are the mean vector and covariance matrix in the pre-constructed physiological acoustic baseline model, respectively.

[0078] The AI ​​intelligent processor 10 will calculate the Mahalanobis distance D. M With a set of preset risk thresholds (e.g., T) alert and T critical T alert <T critical (Compare)

[0079] If D M ≤T alert If so, the risk level is determined to be normal; If T alert <D M ≤T critical If so, it is determined to be a warning; If D M >T critical If so, it is considered a critical situation.

[0080] This enabled a real-time, quantitative risk assessment of physiological states.

[0081] Please refer to the appendix. Figure 3 , Figure 3 This is a flowchart of an online adaptive calibration mechanism according to an embodiment of the present invention.

[0082] To ensure the accuracy of the electromyography (EMG) intention recognition model during long-term use, the AI ​​intelligent processor 10 is also configured to perform an online adaptive calibration process, which uses data provided by the visual capture module 20 to continuously monitor and fine-tune the EMG intention recognition model.

[0083] Specifically, the AI ​​intelligent processor 10 runs a separate visual action recognition model in parallel, which is dedicated to processing the video data stream from the visual capture module 20. This visual action recognition model (e.g., a three-dimensional convolutional neural network 3D-CNN) is pre-trained to recognize the actual actions performed by the operator in the surgical environment.

[0084] When the system is operating in real time, the visual action recognition model analyzes the video data and outputs a recognition result that is considered to be the actual action performed by the operator, denoted as A. vis (t).

[0085] The AI ​​intelligent processor 10 continuously processes the final operational intent I obtained from the aforementioned steps. final (t) and the actual executed action A identified by the visual model vis (t) performs a consistency comparison, when I final (t)≠A vis When (t), it is recorded as a conflict event.

[0086] To quantify the degree of conflict between intention and action, the AI ​​intelligent processor 10 calculates the conflict rate, i.e., the conflict metric C, within a sliding time window of size N. R (t), the formula for calculating this conflict metric is as follows: Where: t is the current time index; N is the length of the sliding window; δ(a,b) is an indicator function, which has a value of 1 when a≠b and a value of 0 when a=b.

[0087] The AI ​​intelligent processor 10 will calculate the conflict metric C. R (t) and a preset calibration threshold T calib When comparing, when C R (t)≥T calib When the preset calibration conditions are met, the system triggers the online calibration procedure.

[0088] After the calibration procedure is triggered, the AI ​​intelligent processor 10 extracts the corresponding electromyographic signal feature vector F from all recorded conflict events within the current sliding window. emg(i) and the visually confirmed actual action A vis (i) These data are combined into a set of corrected data pairs: D set ={(F emg (i),A vis (i))|I final (i)≠A vis (i)}.

[0089] The AI ​​intelligent processor 10 utilizes this correction dataset D set The parameters θ of the existing electromyography intention recognition model are updated online using incremental learning, such as fine-tuning the model parameters by performing one or several steps of gradient descent. This parameter update process can be represented by the following equation: Where: θ old These are the model parameters before calibration; θ new These are the calibrated model parameters; η is a preset learning rate; L(D) set ;θ old ) is in the calibration dataset D set The loss function value calculated above; This represents the gradient operator with respect to the model parameter θ.

[0090] Through this online calibration mechanism, the electromyographic intention recognition model can adaptively adjust to compensate for changes in electromyographic signal patterns caused by factors such as operator muscle fatigue and slight electrode position shifts, thereby maintaining the accuracy of the system in recognizing pre-operation intentions.

[0091] After generating structured composite collaborative instructions, the AI ​​intelligent processor 10 immediately initiates the information output process, converting the instructions into visual and auditory information that the user can perceive, and presenting it through the information prompt module 50.

[0092] A composite collaborative instruction is a data packet containing multiple fields. Its structure can be defined as including three fields: instruction type, instruction content, and priority. The AI ​​intelligent processor 10 first parses the data packet and extracts the values ​​of each field.

[0093] Based on the parsed priority field, the AI ​​intelligent processor 10 determines the way the information is presented. For example, the priority can be divided into three levels: normal, warning and critical. Different levels correspond to a set of preset visual and auditory presentation parameters.

[0094] This structured, composite collaborative instruction is generated by the AI ​​intelligent processor 10 based on an internal decision matrix or rule engine, which correlates and matches the final operational intent with the physiological state risk level. For example: If the final operational intent is to use an electrosurgical cutter for cutting, and the risk level is normal, then the generated instruction type is action confirmation, the content is confirmation of electrosurgical operation, and the priority is normal.

[0095] If the final operational intent is to perform suturing, and the risk level is warning (e.g., a slight abnormality in heart rate is detected), then the generated instruction type is risk warning, the content is "Attention, heart rate fluctuation", and the priority is warning.

[0096] If the ultimate intention of the operation is to apply pressure and the risk level is critical (e.g., a heart murmur or weakening is detected), the generated instruction type is intervention instruction, the content is "Danger! Stop compression immediately and check for bleeding points!", and the priority is critical.

[0097] For the visual presentation through the semi-transparent display screen 51, the information output unit of the AI ​​intelligent processor 10 calls a graphics rendering interface.

[0098] If the instruction priority is normal, the instruction content is rendered as static text in a designated information area of ​​the display screen using a preset non-alert color (e.g., white or green).

[0099] If the priority of the instruction is warning, the instruction content will be rendered as static text with a standardized yellow warning icon.

[0100] If the priority of the instruction is critical, the instruction content will be rendered as flashing text with a red emergency alarm icon to capture more of the user's visual attention.

[0101] For auditory presentation via the in-ear cue 52, the information output unit of the AI ​​intelligent processor 10 calls an audio processing interface.

[0102] For action command type commands, the system calls a text-to-speech (TTS) engine to convert the text string of the command content field into a speech signal, which is then played through the in-ear prompter 52.

[0103] For risk warning instructions, the system determines the output method based on their priority.

[0104] If the priority is warning, a medium-frequency alert tone will be played first, followed by the TTS-converted voice message. If the priority is critical, a high-frequency, repetitive alarm tone will be played first, followed by a TTS-converted voice message to ensure that the user notices the risk information immediately.

[0105] The AI ​​intelligent processor 10 ensures that high-priority instructions are presented overriding or taking precedence over low-priority instructions. For example, when a critical-level risk warning instruction is generated, the system immediately interrupts the currently playing regular-level action instruction voice and replaces it with alarm and emergency voice prompts; simultaneously, the visual information on the display screen is also overridden by the emergency alarm information. This instruction arbitration mechanism ensures that the most critical information is delivered to the user without delay.

Claims

1. A motion capture, recognition, and prediction system based on AI processing, characterized in that, include: The visual capture module is used to collect the operator's visual motion data in real time; An electromyography (EMG) signal acquisition module is used to acquire EMG signals of the operator's limbs in real time. The bioacoustic acquisition module is used to acquire the heart sounds and respiratory sounds of the surgical subject in real time as bioacoustic signals. The information prompt module is used to output prompt messages to the user; The AI ​​intelligent processor is communicatively connected to the visual capture module, the electromyography signal acquisition module, the bioacoustic acquisition module, and the information prompting module. The AI ​​intelligent processor is configured as follows: Based on the visual motion data, visual features and the operator's actual actions are identified. Based on the electromyographic signals, the operator's pre-operational intention is decoded; Based on the bioacoustic signals, the real-time physiological state of the surgical subject was analyzed. By integrating the pre-operation intention with the real-time physiological state, a composite collaborative instruction is generated; The information prompting module is controlled to output the prompting information according to the composite collaborative instruction.

2. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The composite cooperative instruction includes: Action instructions are used to specify the consumables that need to be prepared and the auxiliary operations that need to be performed. Priority instructions are used to characterize the urgency of the action instructions; Risk warning instructions are used to characterize risk information associated with the real-time physiological state.

3. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The process of decoding the operator's pre-operational intent based on the electromyographic signals specifically includes: Collect electromyographic signal samples corresponding to the operator performing a preset set of standard movements; extract features from the electromyographic signal samples to form feature vectors; use the correspondence between the feature vectors and the preset standard movements to train a classifier to construct the electromyographic intention recognition model; The electromyographic intention recognition model is used to process the real-time acquired electromyographic signals and output the corresponding pre-operation intention.

4. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The analysis of the real-time physiological state of the surgical subject based on the bioacoustic signals specifically includes: Bioacoustic signals of the surgical subject in a resting state are collected, acoustic features of the bioacoustic signals in the resting state are extracted, the acoustic features are stored, and the physiological acoustic baseline model is constructed. The real-time acquired bioacoustic signals are compared with the constructed physiological acoustic baseline model to assess and determine the real-time physiological state of the surgical subject.

5. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The identification and prediction system also includes: An environmental acoustic acquisition module is used to acquire environmental noise signals. The AI ​​intelligent processor is also configured to: Based on the decoded pre-operation intention, determine the instrument noise model associated with the pre-operation intention; Using the aforementioned instrument noise model, the mixed signal of the bioacoustic signal and the environmental noise signal is subjected to adaptive filtering noise suppression processing, and the signal after noise suppression processing is used as the bioacoustic signal for the analysis of the real-time physiological state.

6. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The AI ​​intelligent processor is also configured to: Extract the visual features identified based on the visual motion data; The visual features are used to fuse and correct the pre-operational intent obtained from the decoded electromyographic signals, thereby updating the pre-operational intent.

7. The motion capture, recognition, and prediction system based on AI processing according to claim 3, characterized in that, The AI ​​intelligent processor is also configured to: Extract the actual actions performed by the operator based on the visual motion data; Compare the consistency between the pre-operation intention and the actual action performed; When the conflict measurement between the pre-operation intention and the actual executed action meets the preset calibration conditions, the electromyographic intention recognition model is calibrated online using the visual motion data.

8. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The visual capture module includes: Augmented reality glasses worn on the user's head, the augmented reality glasses being equipped with a first camera and an environmental camera deployed in the surgical environment.

9. The motion capture, recognition, and prediction system based on AI processing according to claim 1, characterized in that, The electromyography (EMG) signal acquisition module is a wearable surface EMG sensor armband or wristband, and the bioacoustic acquisition module is a contact bioacoustic sensor.

10. The motion capture, recognition, and prediction system based on AI processing according to claim 8, characterized in that, The information prompt module includes: A semi-transparent display screen is mounted on the augmented reality glasses, and an in-ear prompter is communicatively connected to the augmented reality glasses.

Citation Information

Cited By

  • Laparoscope real-time navigation method based on augmented reality

    CN121796060A

  • A laparoscopic real-time navigation method based on augmented reality

    CN121796060B