Voice-controlled colonoscopy patient lifting system
Through the voice-controlled colonoscopy patient lift system, the neural network model is used to process the doctor's voice signals, and the rapid and accurate adjustment of the electric lift device is achieved, solving the instability and cross-infection of the artificial lift method, and improving the efficiency and accuracy of colonoscopy.
Patent Information
- Application Number
- CN202510608969.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
AI Technical Summary
In existing colonoscopy, it is difficult for artificial lifting to maintain stable lifting status for a long time, which affects the accuracy of the inspection and poses a risk of cross-infection. It is difficult for existing device lifting methods to quickly and accurately adjust the height and angle.
A voice-controlled colonoscopy patient lift system uses a voice-controlled system to collect, process, convert, generate and encode the doctor's voice signal data, use a neural network model to generate control command codes, and control the electric lift device for rapid and accurate adjustments.
It realizes rapid and precise adjustment of lift height and angle without manual lifting, improving the efficiency and accuracy of colonoscopy and reducing the risk of cross-infection.
Smart Images

Figure CN120472897A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical examination systems, and more particularly, relates to a voice-controlled colonoscopy patient lifting system. Background Art
[0002] Colonoscopy is an extremely important medical examination method that can accurately identify a variety of upper gastrointestinal tract lesions. Currently, the main methods for supporting patients during colonoscopy include manual support and device support. The manual support method often has some limitations during colonoscopy.
[0003] The manual lifting method cannot ensure the stability of the patient's head and other parts of the body, and is prone to shaking, which affects the operation of the colonoscopy. For example, it may cause deviations in the angle and position of the colonoscope when inserted, increasing the patient's discomfort and even causing damage to the throat and other parts. Long-term lifting will cause fatigue in the arms and other parts of the medical staff, especially during complex or time-consuming colonoscopy examinations. Fatigue may cause the strength and accuracy of the lifting to decrease, thus affecting the examination results. In addition, the medical staff's hands are in direct contact with the patient. If the hands are not thoroughly cleaned or the patient's skin is damaged, there may be a risk of cross infection. At the same time, the existing device lifting method is difficult to quickly and accurately adjust the height and angle according to the doctor's instructions. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a voice-controlled colonoscopy patient lifting system to solve the problems in the existing technology that manual lifting methods are difficult to maintain a stable lifting state for a long time and possible cross-infection, and to improve the accuracy and efficiency of the device lifting method.
[0005] The purpose and efficacy of the voice-controlled colonoscopy patient support system of the present invention are achieved by the following specific technical means:
[0006] A voice-controlled colonoscopy patient lifting system comprising:
[0007] An acquisition module is used to collect voice signal data of different doctors' control instructions during the colonoscopy examination to obtain a control instruction data set;
[0008] A processing module, wherein the processing module is used to perform data preprocessing on the control instruction data set;
[0009] A conversion module, the conversion module is used to convert the control instruction data in the control instruction data set into a control instruction spectrogram;
[0010] A generation module, the generation module is used to input the control instruction spectrogram into the neural network model to generate the control instruction probability distribution and digital code;
[0011] An encoding module, configured to encode the control instruction probability distribution and the digital code to obtain an intraoperative control instruction code set;
[0012] The control module is used to control the electric lifting device for rapid and precise adjustment according to the intraoperative control instruction code set.
[0013] As a further solution of the present invention, the management system includes the following control steps:
[0014] S1: Obtain the control instruction dataset of the doctor during the operation and perform data preprocessing on the control instruction dataset;
[0015] S2: Performing short-time Fourier transform on the control instruction data in the control instruction data set to convert the control instruction data into a control instruction spectrogram;
[0016] S3: Establish a neural network model, input the control command spectrogram into the neural network model, and obtain the control command probability distribution and digital code;
[0017] S4: Encode the control instruction probability distribution and the digital code to obtain the intraoperative control instruction code set;
[0018] S5: Control and adjust the electric lifting device based on the intraoperative control instruction code set.
[0019] As a further solution of the present invention, the establishment of a neural network model, inputting the control instruction spectrogram into the neural network model, and obtaining the control instruction probability distribution and digital code include:
[0020] The input control instruction spectrogram is divided into two paths, upper and lower, for feature extraction. The control instruction spectrogram of the upper path is subjected to three convolutional layers to extract local time-frequency features, and the context information in the control instruction spectrogram is fused through three deconvolutional layers to generate the first feature information and the second feature information. The first feature information and the second feature information are respectively used to represent the global time-frequency contour and local high-frequency details of the input control instruction spectrogram. The control instruction spectrogram of the lower path is subjected to three convolutional layers with residual connections to suppress equipment noise interference, and then the third feature information and the fourth feature information are generated through three deconvolution layers embedded with a channel attention mechanism. The third feature information and the fourth feature information are respectively used to represent fine parameter features and coarse-grained action features. Feature processing is performed on the first feature information, the second feature information, the third feature information and the fourth feature information to obtain the control instruction probability distribution and digital coding.
[0021] As a further solution of the present invention, the feature processing of the first feature information, the second feature information, the third feature information, and the fourth feature information to obtain the control instruction probability distribution and digital code includes:
[0022] The first feature information and the fourth feature information are spliced along the channel dimension and input into two convolutional layers to extract complementary features to generate a joint feature vector A. The second feature information and the third feature information are added element by element, and the complementary details are enhanced through two convolutional layers to generate a joint feature vector B. The feature vector A is mapped through a fully connected layer to output the probability distribution of 12 types of control instructions. The feature vector B is passed through an independent fully connected layer to output a 12-dimensional digital code.
[0023] As a further solution of the present invention, encoding the control instruction probability distribution and the digital code to obtain the intraoperative control instruction code set includes:
[0024] The highest probability value index in the 12 categories of probability distribution is taken as the main instruction code. If the probability value is greater than 0.5, it is directly adopted. The maximum value index in the 12-dimensional digital code vector is taken as the verification code. If its value is greater than 0.7, it passes the verification. When the main instruction code and the verification code conflict, the judgment is made based on the probability difference. If the absolute value of the probability difference exceeds 0.2, it is processed according to the coding confidence. Based on the system control instruction code, a control instruction code set containing the instruction name, digital code, confidence probability, timestamp and original data is generated.
[0025] As a further solution of the present invention, performing short-time Fourier transform on the control instruction data in the control instruction data set to convert the control instruction data into a control instruction spectrogram includes:
[0026] The control command data is framed and sliced using a Hanning window with a 25ms window length and a 75% overlap rate. A 512-point FFT is performed on each frame, and the frequency resolution is improved by zero padding. The linear spectrum is mapped to the Mel scale that conforms to the human auditory characteristics using a 40-dimensional Mel filter bank. The Mel spectrum is then subjected to natural logarithm compression to compress the dynamic range to 0-60dB, generating a 128×128 pixel grayscale spectrogram. Dynamic time warping is then applied to unify the time axis to 128 steps.
[0027] Spectral subtraction was used to remove the pre-collected operating room noise template, and 20% of the time segments or 30% of the frequency bands were randomly masked to simulate signal loss. Global contrast normalization was then used to eliminate device gain differences. It was tested whether 80% of the energy of the spectrogram was concentrated in the core frequency band of 500Hz-4kHz. 5% of the samples were randomly selected for three-dimensional visualization manual verification to confirm that the high-frequency resonance peaks and the time series energy surge characteristics were consistent with expectations.
[0028] As a further solution of the present invention, the system control instruction code includes:
[0029] The strict correspondence between 12 groups of control instructions and digital codes is predefined to form the core logic basis for machine execution. The 12 groups of control instructions are divided into system control instructions, displacement control instructions and posture control instructions;
[0030] System control instructions are represented by digital code 0 corresponding to the "start operation" in the control instruction, and digital code 11 corresponding to the "end operation" in the control instruction. Digital codes 0 and 11 trigger the start and stop of the power respectively and have the highest execution priority;
[0031] Displacement control instructions are represented by codes 1-4 corresponding to two-dimensional plane movement. Digital code 1 is "move up", driving the Y-axis to move 10cm in the positive direction, digital code 2 is "move down", driving the Y-axis to move 10cm in the negative direction, digital code 3 is "move left", driving the X-axis to move 10cm in the positive direction, and digital code 4 is "move right", driving the X-axis to move 10cm in the negative direction.
[0032] Attitude control instructions are represented by codes 5-10 corresponding to rotation actions around the X / Y / Z axis. Digital code 5 is "rotate left", which executes a 15° clockwise rotation around the X axis. Digital code 6 is "rotate right", which executes a 15° counterclockwise rotation around the X axis. Digital code 7 is "rotate up", which executes a 15° clockwise rotation around the Y axis. Digital code 8 is "rotate down", which executes a 15° counterclockwise rotation around the Y axis. Digital code 9 is "rotate forward", which executes a 15° clockwise rotation around the Z axis. Digital code 10 is "rotate backward", which executes a 15° counterclockwise rotation around the Z axis.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The present invention proposes an electric lifting method based on intelligent voice control, which does not require manual lifting and can quickly and accurately adjust the lifting height and angle according to the doctor's instructions, thereby improving the efficiency and accuracy of colonoscopy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The present invention is a flow chart of the steps of a voice-controlled colonoscopy patient lifting system. DETAILED DESCRIPTION
[0036] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the technical solutions of the present invention, but are not intended to limit the scope of protection of the present invention.
[0037] Example 1:
[0038] As attached Figure 1 As shown:
[0039] The present invention provides a voice-controlled colonoscopy patient lifting system, comprising:
[0040] An acquisition module is used to collect voice signal data of different doctors' control instructions during the colonoscopy examination to obtain a control instruction data set;
[0041] A processing module, wherein the processing module is used to perform data preprocessing on the control instruction data set;
[0042] A conversion module, the conversion module is used to convert the control instruction data in the control instruction data set into a control instruction spectrogram;
[0043] A generation module, the generation module is used to input the control instruction spectrogram into the neural network model to generate the control instruction probability distribution and digital code;
[0044] An encoding module, configured to encode the control instruction probability distribution and the digital code to obtain an intraoperative control instruction code set;
[0045] The control module is used to control the electric lifting device for rapid and precise adjustment according to the intraoperative control instruction code set.
[0046] The system includes the following control steps:
[0047] Step S1: Acquire a control instruction data set of a doctor during surgery and perform data preprocessing on the control instruction data set.
[0048] Specifically, dual-channel high-sensitivity directional microphones are deployed on both sides of the examination bed head, at a 45° angle to the doctor's mouth, and synchronously connected to a professional sound card. The doctor uses a foot switch to trigger control command data collection and obtain the control command data set.
[0049] Furthermore, the control instruction data set was cleaned to remove silence segments longer than 2 seconds and low signal-to-noise ratio segments with a signal-to-noise ratio below 20 dB, and interpolation was performed to repair short-term signal loss.
[0050] Step S2: performing short-time Fourier transform on the control instruction data in the control instruction data set to convert the control instruction data into a control instruction spectrogram.
[0051] Specifically, the control command data is framed and processed, and the continuous control command data is cut using a Hanning window with a window length of 25ms and a 75% overlap rate. A 512-point FFT is performed on each frame, and the frequency resolution is improved by zero padding. The linear spectrum is mapped to the Mel scale that conforms to the human auditory characteristics through a 40-dimensional Mel filter group. The natural logarithm of the Mel spectrum is taken to compress the dynamic range to 0-60dB, and a 128×128 pixel grayscale spectrogram is generated. Dynamic time warping is applied to unify the time axis to 128 steps.
[0052] Furthermore, spectral subtraction was used to remove the pre-collected operating room noise template, randomly masking 20% of the time segments or 30% of the frequency bands to simulate signal loss, and then global contrast normalization was used to eliminate device gain differences. It was detected whether 80% of the energy of the spectrogram was concentrated in the 500Hz-4kHz core frequency band, and 5% of the samples were randomly selected for three-dimensional visualization manual verification to confirm that the high-frequency resonance peaks and the time series energy surge characteristics were in line with expectations.
[0053] Step S3: Establish a neural network model, input the control instruction spectrogram into the neural network model, and obtain the control instruction probability distribution and digital code.
[0054] Specifically, the input control instruction spectrogram is divided into two paths, upper and lower, for feature extraction. The control instruction spectrogram of the upper path is subjected to three convolutional layers to extract local time-frequency features, and the context information in the control instruction spectrogram is fused through three deconvolutional layers to generate first feature information and second feature information. The first feature information and the second feature information are respectively used to represent the global time-frequency contour and local high-frequency details of the input control instruction spectrogram. The control instruction spectrogram of the lower path is subjected to three convolutional layers with residual connections to suppress equipment noise interference, and then the third feature information and the fourth feature information are generated through three deconvolution layers embedded with a channel attention mechanism. The third feature information and the fourth feature information are respectively used to represent fine parameter features and coarse-grained action features. Feature processing is performed on the first feature information, the second feature information, the third feature information and the fourth feature information to obtain the control instruction probability distribution and digital coding.
[0055] Furthermore, the first feature information and the fourth feature information are spliced along the channel dimension and input into two convolutional layers to extract complementary features to generate a joint feature vector A. The second feature information and the third feature information are added element by element, and the complementary details are enhanced through two convolutional layers to generate a joint feature vector B. The feature vector A is mapped by a fully connected layer to output the probability distribution of 12 types of control instructions. The feature vector B is passed through an independent fully connected layer to output a 12-dimensional digital code.
[0056] It can be understood that the strict correspondence between the predefined 12 groups of control instructions and digital codes constitutes the core logical basis of machine execution, and the 12 groups of control instructions are divided into system control instructions, displacement control instructions and posture control instructions.
[0057] System control instructions are represented by digital code 0 corresponding to the "start operation" in the control instruction, and digital code 11 corresponding to the "end operation" in the control instruction. Digital codes 0 and 11 trigger the start and stop of the power respectively, and have the highest execution priority.
[0058] Displacement control instructions are represented by codes 1-4, corresponding to two-dimensional plane movement. Digital code 1 is "move up", driving a 10cm displacement in the positive direction of the Y axis, digital code 2 is "move down", driving a 10cm displacement in the negative direction of the Y axis, digital code 3 is "move left", driving a 10cm displacement in the positive direction of the X axis, and digital code 4 is "move right", driving a 10cm displacement in the negative direction of the X axis.
[0059] Attitude control instructions are represented by codes 5-10 corresponding to rotation actions around the X / Y / Z axis. Digital code 5 is "rotate left", which executes a 15° clockwise rotation around the X axis. Digital code 6 is "rotate right", which executes a 15° counterclockwise rotation around the X axis. Digital code 7 is "rotate up", which executes a 15° clockwise rotation around the Y axis. Digital code 8 is "rotate down", which executes a 15° counterclockwise rotation around the Y axis. Digital code 9 is "rotate forward", which executes a 15° clockwise rotation around the Z axis. Digital code 10 is "rotate backward", which executes a 15° counterclockwise rotation around the Z axis.
[0060] Step S4: Encode the control instruction probability distribution and the digital code to obtain an intraoperative control instruction code set.
[0061] Specifically, the highest probability value index in the 12 categories of probability distribution is taken as the main instruction code. If the probability value is greater than 0.5, it is directly adopted. The maximum value index in the 12-dimensional digital code vector is taken as the verification code. If its value is greater than 0.7, it passes the verification. For example, a control instruction spectrogram is transformed by a neural network model to obtain the following probability distribution P1 = [0.02, 0.08, 0.05, 0.1, 0.6, 0.01, 0.03, 0.02, 0.04, 0.01, 0.03, 0.01] and digital code C1 = [0.1, 0.2, 0.05, 0.3, 0.85, 0.1, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], and the control instruction corresponding to the highest probability in the above probability distribution P1 and digital code C1 is moving right, then the digital code 4 corresponding to the control instruction moving right is output.
[0062] When the main instruction code conflicts with the verification code, a judgment is made based on the probability difference. If the absolute value of the probability difference exceeds 0.2, it is processed according to the coding confidence. For example, a control instruction spectrogram is transformed by the neural network model to obtain the following probability distribution P2 = [0.01, 0.1, 0.05, 0.55, 0.2, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.09] and digital code C2 = [0.0, 0.1, 0.0, 0.2, 0.82, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], where the control instruction corresponding to the highest probability in the probability distribution P2 is to move left. And the highest probability 0.55 is greater than 0.5, the control instruction corresponding to the highest probability in the digital code C2 is to move right and the highest probability 0.82 is greater than 0.7, and the probability distribution P2 conflicts with the control instruction corresponding to the digital code C2. At this time, the absolute value of the difference between the probability distribution P2 and the digital code C2 is 0.27, which is greater than 0.2. Because the confidence of the digital code C2 is higher than the confidence of the probability distribution P2, the control instruction output this time is to move right. The confidence level is expressed as the difference between the maximum probability and the preset threshold. Based on the system control instruction code, a control instruction code set including the instruction name, digital code, confidence probability, timestamp and original data is generated.
[0063] Step S5: Control and adjust the electric lifting device based on the intraoperative control instruction code set.
[0064] The above embodiments can be implemented in whole or in part through software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wireless method (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0065] It should be understood that the term "and / or" as used herein simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent the existence of A alone, the existence of both A and B, or the existence of B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0066] It should be understood that in the embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0067] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A voice-controlled colonoscopy patient lifting system, characterized in that: include: An acquisition module is used to collect voice signal data of different doctors' control instructions during the colonoscopy examination to obtain a control instruction data set; A processing module, wherein the processing module is used to perform data preprocessing on the control instruction data set; A conversion module, the conversion module is used to convert the control instruction data in the control instruction data set into a control instruction spectrogram; A generation module, the generation module is used to input the control instruction spectrogram into the neural network model to generate the control instruction probability distribution and digital code; An encoding module, configured to encode the control instruction probability distribution and the digital code to obtain an intraoperative control instruction code set; The control module is used to control the electric lifting device for rapid and precise adjustment according to the intraoperative control instruction code set.
2. A voice-controlled colonoscopy patient lifting system according to claim 1, characterized in that: The lifting system includes the following control steps: S1: Obtain the control instruction dataset of the doctor during the operation and perform data preprocessing on the control instruction dataset; S2: Performing short-time Fourier transform on the control instruction data in the control instruction data set to convert the control instruction data into a control instruction spectrogram; S3: Establish a neural network model, input the control command spectrogram into the neural network model, and obtain the control command probability distribution and digital code; S4: Encode the control instruction probability distribution and the digital code to obtain the intraoperative control instruction code set; S5: Control and adjust the electric lifting device based on the intraoperative control instruction code set.
3. The voice-controlled colonoscopy patient lifting system according to claim 2, characterized in that: The process of establishing a neural network model, inputting the control instruction spectrogram into the neural network model, and obtaining the control instruction probability distribution and digital coding includes: The input control instruction spectrogram is divided into two paths, upper and lower, for feature extraction. The control instruction spectrogram of the upper path is subjected to three convolutional layers to extract local time-frequency features, and the context information in the control instruction spectrogram is fused through three deconvolutional layers to generate the first feature information and the second feature information. The first feature information and the second feature information are respectively used to represent the global time-frequency contour and local high-frequency details of the input control instruction spectrogram. The control instruction spectrogram of the lower path is subjected to three convolutional layers with residual connections to suppress equipment noise interference, and then the third feature information and the fourth feature information are generated through three deconvolution layers embedded with a channel attention mechanism. The third feature information and the fourth feature information are respectively used to represent fine parameter features and coarse-grained action features. Feature processing is performed on the first feature information, the second feature information, the third feature information and the fourth feature information to obtain the control instruction probability distribution and digital coding.
4. The voice-controlled colonoscopy patient lifting system according to claim 3, characterized in that: The performing feature processing on the first feature information, the second feature information, the third feature information, and the fourth feature information to obtain the control instruction probability distribution and digital code includes: The first feature information and the fourth feature information are spliced along the channel dimension and input into two convolutional layers to extract complementary features to generate a joint feature vector A. The second feature information and the third feature information are added element by element, and the complementary details are enhanced through two convolutional layers to generate a joint feature vector B. The feature vector A is mapped through a fully connected layer to output the probability distribution of 12 types of control instructions. The feature vector B is passed through an independent fully connected layer to output a 12-dimensional digital code.
5. The voice-controlled colonoscopy patient lifting system according to claim 2, characterized in that: The step of encoding the control instruction probability distribution and the digital code to obtain the intraoperative control instruction code set includes: The highest probability value index in the 12 categories of probability distribution is taken as the main instruction code. If the probability value is greater than 0.5, it is directly adopted. The maximum value index in the 12-dimensional digital code vector is taken as the verification code. If its value is greater than 0.7, it passes the verification. When the main instruction code and the verification code conflict, the judgment is made based on the probability difference. If the absolute value of the probability difference exceeds 0.2, it is processed according to the coding confidence. Based on the system control instruction code, a control instruction code set containing the instruction name, digital code, confidence probability, timestamp and original data is generated.
6. The voice-controlled colonoscopy patient lifting system according to claim 2, characterized in that: The performing short-time Fourier transform on the control instruction data in the control instruction data set to convert the control instruction data into a control instruction spectrogram includes: The control command data is framed and sliced using a Hanning window with a 25ms window length and a 75% overlap rate. A 512-point FFT is performed on each frame, and the frequency resolution is improved by zero padding. The linear spectrum is mapped to the Mel scale that conforms to the human auditory characteristics using a 40-dimensional Mel filter bank. The Mel spectrum is then subjected to natural logarithm compression to compress the dynamic range to 0-60dB, generating a 128×128 pixel grayscale spectrogram. Dynamic time warping is then applied to unify the time axis to 128 steps. Spectral subtraction was used to remove the pre-collected operating room noise template, and 20% of the time segments or 30% of the frequency bands were randomly masked to simulate signal loss. Global contrast normalization was then used to eliminate device gain differences. It was tested whether 80% of the energy of the spectrogram was concentrated in the core frequency band of 500Hz-4kHz. 5% of the samples were randomly selected for three-dimensional visualization manual verification to confirm that the high-frequency resonance peaks and the time series energy surge characteristics were consistent with expectations.
7. The voice-controlled colonoscopy patient lifting system according to claim 5, characterized in that: The system control instruction code includes: The strict correspondence between 12 groups of control instructions and digital codes is predefined to form the core logic basis for machine execution. The 12 groups of control instructions are divided into system control instructions, displacement control instructions and posture control instructions; System control instructions are represented by digital code 0 corresponding to the "start operation" in the control instruction, and digital code 11 corresponding to the "end operation" in the control instruction. Digital codes 0 and 11 trigger the start and stop of the power respectively and have the highest execution priority; Displacement control instructions are represented by codes 1-4, corresponding to two-dimensional plane movement. Digital code 1 is "move up", driving the Y-axis to move 10cm in the positive direction. Digital code 2 is "move down", driving the Y-axis to move 10cm in the negative direction. Digital code 3 is "move left", driving the X-axis to move 10cm in the positive direction. Digital code 4 is "move right", driving the X-axis to move 10cm in the negative direction. Attitude control instructions are represented by codes 5-10 corresponding to rotation actions around the X / Y / Z axis. Digital code 5 is "rotate left", which executes a 15° clockwise rotation around the X axis. Digital code 6 is "rotate right", which executes a 15° counterclockwise rotation around the X axis. Digital code 7 is "rotate up", which executes a 15° clockwise rotation around the Y axis. Digital code 8 is "rotate down", which executes a 15° counterclockwise rotation around the Y axis. Digital code 9 is "rotate forward", which executes a 15° clockwise rotation around the Z axis. Digital code 10 is "rotate backward", which executes a 15° counterclockwise rotation around the Z axis.
8. The voice-controlled colonoscopy patient lifting system according to claim 2, characterized in that: The step of obtaining a voice signal dataset of the intraoperative doctor's control instructions and performing data preprocessing on the voice signal dataset includes: Deploy dual-channel high-sensitivity directional microphones on both sides of the examination bed head, at a 45° angle to the doctor's mouth. They are synchronously connected to a professional sound card. The doctor uses a foot switch to trigger control command data collection and obtain the control command data set. The control instruction data set is cleaned to remove silence segments longer than 2 seconds and low signal-to-noise ratio segments with a signal-to-noise ratio below 20 dB, and interpolation is performed to repair short-term signal loss.