Control method and surgical robot based on eye positioning and voice recognition
Through eye positioning and voice recognition technology, the problem that doctors cannot operate the equipment with both hands during the operation is solved, and precise control based on the relative position of the pupil and voice commands is achieved, which improves the auxiliary effect of the surgery.
Patent Information
- Application Number
- CN202210667711.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-06-14
AI Technical Summary
During the operation, when doctors need to operate the surgical robot and screen at the same time, they cannot effectively use their hands to control the equipment, and the existing technology lacks an effective solution.
Through eye positioning and sound recognition technology, real-time eye images and voice data of the target operator are obtained, the relative coordinates of the pupil in the eyepiece are determined and control instructions are identified to achieve accurate control of the target screen.
It realizes that without the help of external devices, doctors can accurately control the equipment through the eyes and sound to assist surgical operations.
Smart Images

Figure CN115089300B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of electronic digital data processing technology, and in particular relates to a control method and a surgical robot based on eye positioning and sound recognition. Background Art
[0002] Currently, when controlling electronic devices with screens, external devices are generally used, for example, an external mouse or keyboard is used to control the device and complete system operations.
[0003] However, in some cases, an operator's hands are needed for other tasks, preventing them from controlling the device using a mouse or other similar device. For example, when a doctor is performing surgery with a surgical robot, they may need to manipulate the images on the screen, such as displaying, zooming in, or reducing them, or overlaying them, to assist with the procedure. However, because the doctor also needs to operate the handle, their hands are currently unable to operate an external device such as a mouse.
[0004] Currently, no effective solution has been proposed to the above-mentioned problem of being unable to effectively control the device. Summary of the Invention
[0005] The purpose of this application is to provide a control method and surgical robot based on eye positioning and sound recognition, which can achieve effective control of the equipment.
[0006] On the one hand, a control method based on eye positioning and sound recognition is provided, comprising:
[0007] Acquire real-time eye images and voice data of the target operator through the eyepiece;
[0008] Determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the corresponding target screen according to the relative coordinate data;
[0009] A control instruction is recognized from the voice data, and the control instruction is executed at a coordinate position in the corresponding target screen.
[0010] In one embodiment, determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining the coordinate position of the corresponding target screen according to the relative coordinate data includes:
[0011] identifying, from the eye image, position coordinates of the pupil in the eyepiece as first coordinate information;
[0012] Retrieving a pre-generated correspondence between the eye image coordinate system and the screen coordinate system as a first correspondence;
[0013] The coordinate position of the pupil in the eye image corresponding to the target screen is determined according to the first coordinate information and the first corresponding relationship.
[0014] In one embodiment, pre-generating a correspondence between the eye image coordinate system and the screen coordinate system includes:
[0015] displaying a plurality of position guide points on the target screen;
[0016] displaying guidance information for guiding the target operator to look at each of the plurality of position guidance points one by one;
[0017] Acquire multiple eye images of the target operator gazing at each position guide point through the eyepiece;
[0018] A correspondence between the eye image coordinate system and the screen coordinate system is formed according to the positions of the pupils in the eyepiece in the plurality of eye images and the position information of the corresponding position guide points on the screen.
[0019] In one embodiment, after obtaining multiple eye images of the target operator gazing at each position guide point one by one through the eyepiece, the method further includes:
[0020] determining whether a shape formed by pupil position points in the plurality of eye images is the same as a shape formed by the plurality of position guide points;
[0021] If they are determined to be different, a trigger is provided to redisplay the plurality of position guide points on the target screen.
[0022] In one embodiment, after forming a correspondence between the eye image coordinate system and the screen coordinate system based on the positions of the pupils in the eyepiece in the plurality of eye images and the position information of the corresponding position guide points on the screen, the method further includes:
[0023] Detecting whether the eyes of the target operator leave and then return to the eyepiece through a pressure sensor provided at the edge of the eyepiece;
[0024] In the case where it is determined that the eyes of the target operator leave and then return to the eyepiece, displaying a center position guide point at the center position of the target screen;
[0025] acquiring an eye image of the target operator gazing at the central position guide point through the eyepiece as a target eye image;
[0026] The correspondence between the eye image coordinate system and the screen coordinate system is calibrated according to the position of the pupil in the eyepiece in the target eye image.
[0027] In one embodiment, identifying the control instruction from the voice data includes:
[0028] performing voiceprint recognition on the voice data to determine voice content in the voice data that matches the voiceprint of the target operator;
[0029] The determined voice content is subjected to text recognition to determine the control instruction.
[0030] In one embodiment, acquiring a real-time eye image of the target operator through an eyepiece includes:
[0031] The eye image of the target operator through the eyepiece is obtained through the installed camera.
[0032] In one embodiment, the eyepiece is an eyepiece on a surgical robot, and the control instruction is an operation instruction for a display screen on the surgical robot during surgery.
[0033] On the other hand, an eye-based positioning method is provided, comprising:
[0034] Acquire a real-time eye image of the target object through the eyepiece;
[0035] Determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the corresponding target screen according to the relative coordinate data;
[0036] The corresponding coordinate position is displayed in the target screen in the form of an icon.
[0037] In one embodiment, the eyepiece and the camera for acquiring eye images are arranged in fixed positions.
[0038] In one embodiment, determining relative coordinate data between the pupil and the eyepiece in the eye image according to the eye image, and determining the coordinate position of the corresponding target screen according to the relative coordinate data includes:
[0039] identifying, from the eye image, the position coordinates of the pupil in the eyepiece as first coordinate information;
[0040] Retrieving a pre-generated correspondence between the eye image coordinate system and the screen coordinate system as a first correspondence;
[0041] The coordinate position of the pupil in the eye image corresponding to the target screen is determined according to the first coordinate information and the first corresponding relationship.
[0042] In another aspect, a surgical robot is provided, comprising:
[0043] A surgical component for a target operator to perform a surgical operation;
[0044] a camera assembly, disposed opposite to the eyepiece, and configured to capture an eye image of the target operator in the eyepiece during the target operator performing a surgical operation;
[0045] a processor connected to the camera assembly, and configured to determine a coordinate position on a screen of a corresponding display screen based on relative coordinate data of the pupil in the eye image in the eyepiece;
[0046] A display screen is in communication with the processor and is used for displaying operations during the surgical procedure.
[0047] In one embodiment, the surgical robot further comprises:
[0048] a sound receiving component, configured to obtain voice data of the target operator during the surgical operation;
[0049] The processor is further configured to identify a control instruction from the voice data and control the execution of the control instruction at a coordinate position of the display screen.
[0050] On the other hand, a surgical robot is provided, comprising a processor and a memory for storing instructions executable by the processor, wherein the processor implements the steps of the above method when executing the instructions.
[0051] On the other hand, a computer-readable storage medium is provided, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0052] This application provides a control method based on eye positioning and voice recognition. This method obtains a real-time eye image of the target operator through the eyepiece to determine the corresponding position on the target screen. It then identifies control commands by acquiring voice data, thereby triggering the execution of control commands at the corresponding coordinate position on the target screen. This solution solves the existing problem of being unable to accurately control the device when the operator's hands are required to perform other operations, achieving the technical effect of accurate control based on the relative position of the pupil in the eyepiece and voice commands. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0054] Figure 1 This is a schematic diagram of the architecture of the surgical robot provided by this application;
[0055] Figure 2 This is a schematic diagram of the architecture of the doctor console provided by this application;
[0056] Figure 3 This is a schematic diagram of the architecture of the imaging trolley provided in this application;
[0057] Figure 4 This is a flow chart of a control method based on eye positioning and sound recognition provided by the present application;
[0058] Figure 5 This is a schematic diagram of determining a mapping relationship based on position guidance points provided by this application;
[0059] Figure 6 This is a calibration diagram for determining a mapping relationship based on position guide points provided by this application;
[0060] Figure 7 This is a temporary calibration diagram for determining a mapping relationship based on position guidance points provided by this application;
[0061] Figure 8 This is a schematic diagram of the mapping relationship between the eyeball position and the display screen coordinate position provided by this application;
[0062] Figure 9 This is a schematic diagram of the camera installation of the system provided by this application;
[0063] Figure 10 This is a schematic diagram of the installation of the radio equipment for the system provided by this application;
[0064] Figure 11 This is a flow chart of the eye-based positioning method provided by this application;
[0065] Figure 12 This is a hardware structure block diagram of an electronic device based on an eye positioning and voice recognition control method provided by the present application;
[0066] Figure 13 This is a schematic diagram of the module structure of an embodiment of a control device based on eye positioning and sound recognition provided by the present application. DETAILED DESCRIPTION
[0067] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0068] In the embodiments of the present application, considering that when operating the device, sometimes the user's hands cannot be freed to operate external devices (such as keyboard, mouse, etc.), in this example, considering that the mouse movement function can be realized through eye positioning, and the operation triggering can be realized through voice commands, so that the user can control the device without using both hands.
[0069] The method is specifically described by taking the application of the control method based on eye positioning and sound recognition to a surgical robot as an example. Figure 1 As shown, the surgical robot may include: an imaging cart 10, a side trolley 11, an operating table trolley 12, a tool cart 13, and a doctor's console 20. Based on this surgical robot, a doctor can remotely operate the robot through the doctor's console 20 to perform surgery on a patient on the operating table trolley 12. The imaging cart 10 is used to provide auxiliary image data to doctors / nurses, etc., and the tool cart 13 is used to provide doctors with tools required for surgery. The side trolley 11 includes at least one imaging arm 110 and a tool arm 112. An image acquisition device is mounted on the imaging arm 110. The image acquisition device is communicatively connected to a display device to acquire image information of the surgical environment and provide it to the display device for display.
[0070] The above-mentioned image acquisition device is used to obtain image information of the surgical environment including human tissues and organs, surgical instruments, blood vessels and body fluids and provide it to the display device. The surgical instrument 113 is mounted on the tool arm 112. The image acquisition device can be as follows: Figure 1 As shown, Figure 1 The endoscope 111 is inserted into the patient's body, and the endoscope 111 and the surgical instrument 113 respectively enter the patient's position through an incision on the patient's body to achieve minimally invasive surgical treatment.
[0071] The doctor console 20 can be Figure 2The system shown includes two manipulator arms 2001 and 2002. The control handles at the ends of these arms detect the operator's hand motion information, which serves as the motion control input for the entire system. A trolley component serves as a base for mounting other components. Foot switches 2003 and 2004 can be mounted on the trolley component to detect on / off control signals from the operator (i.e., the surgical operator). The adjustment component electrically adjusts the position of the manipulator arms, imaging component, operator handrails, and other devices. The imaging component (i.e., display screen) provides the operator with a stereoscopic image captured by the imaging system, providing reliable image information for surgical operations. During surgery, the operator, seated at the doctor's console outside the sterilized area, controls the surgical instruments and laparoscope by manipulating the control handles at the ends of the manipulator arms. The operator observes the intracavitary image transmitted through the eyepieces and uses both hands to control the movements of the patient's surgical platform's robotic arms and instruments, completing various operations to achieve the purpose of performing surgery on the patient. The operator can also control some actions using the foot switches, such as those for electrocautery and electrocoagulation.
[0072] The image trolley can be as follows Figure 3 As shown, it includes: an image host 301, a keyboard 302, a mouse 303, and a display screen 304, wherein the image host 301 is used to provide image information for the endoscope of the surgical robot main control end. The controller can control the image host by controlling the keyboard 302 and the mouse 303. The control result of the image host is displayed through the display screen 304. For example, real-time surgical video and related data during the operation can be displayed.
[0073] This is obviously unreasonable because the surgeon needs surgical instruments to perform surgical operations and needs to control the keyboard and mouse to operate the doctor's console. Therefore, in this case, it is considered that the image components in the doctor's console can be controlled through eye positioning and voice recognition, so that the surgeon can control the image components in the doctor's console without operating the keyboard and mouse.
[0074] Figure 4It is a method flow chart of an embodiment of the control method based on eye positioning and sound recognition provided by the present application. Although the present application provides the method operation steps or device structure as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of the present application and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be connected according to the method or module structure shown in the embodiment or drawings for sequential execution or parallel execution (for example, a parallel processor or a multi-threaded processing environment, or even a distributed processing environment).
[0075] Specifically, such as Figure 4 As shown, the above-mentioned control method based on eye positioning and voice recognition may include the following steps:
[0076] Step 401: Acquire real-time eye image and voice data of the target operator through the eyepiece;
[0077] For example, when the target operator (which may be the above-mentioned operator, i.e., surgical operator, surgeon) is performing surgery, he or she views the image component (i.e., display screen) through the eyepiece, and the distance and setting between the eyepiece and the display screen remain stable. Considering that a camera component can be set up to obtain the eye image of the target operator through the eyepiece, the position of the pupil in the eye image can be matched to the corresponding coordinate position in the display screen.
[0078] Furthermore, a sound receiving device can be provided to obtain the target operator's voice data, identify the target operator's voice content from the voice data, and analyze the control instructions therein, so as to combine the coordinate position and the control instructions to realize the control of the master control system.
[0079] For example, when performing surgery with a surgical robot, the doctor sometimes needs to perform some additional system operations while performing the surgery, such as: retrieving CT image data and ultrasound image data, etc., and superimposing the images to form a 3D image to assist the surgery. However, because the doctor is operating the handle, he cannot free his hands to do other things. Through the analysis of eye images and voice data, the user's control instructions can be identified, for example: displaying CT image data, then the CT image data can be displayed at the position coordinates determined based on the eye image.
[0080] Step 402: determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the pupil in the corresponding target screen according to the relative coordinate data;
[0081] Specifically, to determine the coordinate position of the pupil in the eye image corresponding to the target screen based on the relative coordinate data of the pupil in the eye image, the pupil's position coordinates in the eyepiece can be identified from the eye image as first coordinate information; a pre-generated correspondence between the eye image coordinate system and the screen coordinate system can be retrieved as the first correspondence; and the coordinate position of the pupil in the eye image corresponding to the target screen can be determined based on the first coordinate information and the first correspondence. When actually determining the position and coordinate system, the center point of the pupil in the eyeball can be used as the positioning point. That is, the pupil's position can be first identified from the eyeball, and then the position of the pupil center point in the eye image can be determined, with the position of the pupil center point serving as the basis for determination.
[0082] That is, a correspondence between the coordinates of the eye image and the coordinates of the display screen is established in advance, and the coordinates of the pupil position in the eye image are identified, so that the coordinate position on the display screen can be converted.
[0083] Step 403: Identify a control instruction from the voice data, and execute the control instruction at the coordinate position in the corresponding target screen.
[0084] Considering that the operator's visual range is different, the correspondence between the eye image coordinate system and the screen coordinate system can be generated in advance for each operator. When implementing, the correspondence can be generated in the following way:
[0085] S1: Displaying a plurality of position guide points on the target screen;
[0086] S2: displaying guidance information to guide the target operator to look at each of the plurality of position guidance points one by one;
[0087] S3: Acquire multiple eye images of the target operator gazing at each position guide point one by one through the eyepiece;
[0088] S4: forming a correspondence between an eye image coordinate system and a screen coordinate system according to the positions of the pupils in the eyepiece in the plurality of eye images and the position information of the corresponding position guide points on the screen.
[0089] Considering that the display screen is rectangular, we can set five position guide points to establish the corresponding relationship, such as Figure 5As shown, five bright spots (1, 2, 3, 4, and 5) can be displayed on the display screen in sequence. The operator, according to the guidance, uses the pupil in the eyeball to locate these five guide points in sequence, pausing for 1 second each time, and locates all the bright spots in sequence, thereby determining a rectangular eyeball visual range, thereby achieving the position mapping of the visual range and the screen. That is, the operator's pupil positioning point 1 can form a position point (i.e., the upper left corner), the operator's pupil positioning point 2 can form a position point (i.e., the upper right corner), the operator's pupil positioning point 3 can form a position point (i.e., the lower left corner), the operator's pupil positioning point 4 can form a position point (i.e., the lower right corner), and the operator's pupil positioning point 5 can form a position point (i.e., the center point). In this way, a coordinate system can be formed for the eye image, and the display screen can also form a coordinate system. During implementation, the center point can be used as the (0,0) point of the coordinate system, or the point in the lower left corner can be used as the (0,0) point of the coordinate system. During implementation, the (0,0) point of the coordinate system can be set according to needs to ensure that the coordinate system of the display screen is consistent with the coordinate system formed by the eye image.
[0090] In actual implementation, the positioning method of forming a rectangle and a center point using the above five points can be used, or other positioning methods can be used, for example, three points forming a triangle, five points forming a pentagon, etc. This application does not make any specific limitations on the basic positioning method of the coordinate system.
[0091] Considering that errors or low accuracy may occur in the process of forming the coordinate system correspondence, in this example, a calibration method is provided. After obtaining multiple eye images of the target operator looking at each position guide point through the eyepiece, it is possible to determine whether the shape formed by the pupil position point in the multiple eye images is the same as the shape formed by the multiple position guide points. If it is determined that they are not the same, the multiple position guide points are triggered to be re-displayed on the target screen. That is, Figure 6 As shown, the figure formed by the guide points is a rectangle, but the figure formed by the pupil position is not a rectangle. In this case, it can be determined that the established correspondence is inaccurate. Therefore, a calibration can be triggered. The guide points can be redisplayed to guide the gaze and re-establish the correspondence. Specifically, the eyepiece coordinate system can be established based on the relative position of the pupil in the eye image. For example, four corner guide points and a center point can be selected for guided positioning. The four corner points define the range of eye movement, and the center point assists the four points. Thus, a visual range figure is formed based on the pupil position in the captured image, and the shape is then compared with the figure formed by the guide points to achieve calibration.
[0092] During the operation, the operator sometimes moves the eyepiece away and then moves back. In this case, in order to ensure the accuracy of the coordinate system correspondence, the coordinate system can be calibrated. Considering that the field of view is the same for the same operator, the midpoint of the line of sight often shifts when the operator moves away and then moves back. Therefore, during calibration, only the center point can be calibrated. After calibration, the coordinate system can be translated according to the previously determined coordinate system. Specifically, Figure 7 As shown, when the doctor is operating, sometimes the head will move away from the doctor's eyepiece and then return to the eyepiece. Because the person's face will be offset compared to the last time, the original calibration and center point will also be offset. At this time, a temporary calibration method can be used. During the temporary calibration, only the center point is calibrated, and the visual range is assumed to be unchanged. Therefore, only the center point (5) is displayed in the middle position to remind the operator to only look at the center point. For example, it lasts for about 2 seconds, and then the red dot disappears. In this process, the coordinate system calibration is completed. Specifically, when the operator looks at the center guide point on the screen, the eye image is captured. Assuming the width and height of the visual range of the first calibration, the position after the center point calibration can be translated, that is, to ensure that the center point of the visual range is the same as the center point of the display screen.
[0093] That is, after forming a correspondence between the eye image coordinate system and the screen coordinate system based on the positions of the pupils in the multiple eye images and the position information of the position guide points in the corresponding screen, the pressure sensor set at the edge of the eyepiece can be used to detect whether the eyes of the target operator leave and return to the eyepiece; when it is determined that the eyes of the target operator leave and return to the eyepiece, a center position guide point is displayed at the center position of the target screen; the eye image of the target operator looking at the center position guide point through the eyepiece is obtained as the target eye image; and the correspondence between the eye image coordinate system and the screen coordinate system is calibrated according to the position of the pupil in the eyepiece in the target eye image.
[0094] When the coordinate position on the screen is determined based on the real-time eye image data and the predetermined correspondence, the coordinate position can be determined as follows: Figure 8As shown, the camera captures an image of the operator's eye as they gaze at a specific point. The pupil position in the eyepiece coordinate system is p(x,y). First, calculate the relative distance of point p to the upper-left corner of the eye's range of motion, o. Based on the mapping between the visual range and the display screen, o corresponds to the coordinates of point (0,0) on the display screen. The relative distance is actually the coordinates (x0, y0) on the display screen. During the calculation, subtract the x-coordinate of point o from the x-coordinate of point p, which is (x - x'), and subtract the y-coordinate of point o from the y-coordinate of point p, which is (y - y'). This yields the coordinates of point p relative to o. Multiply the obtained x-coordinate by the screen height / visual range height H, and the y-coordinate by the screen width / visual range width W. The coordinate position on the display screen can be inferred from the relative coordinates in the eyepiece coordinate system.
[0095] That is, during implementation, the pupil coordinates of the image can be obtained first, and then the relative coordinates of the pupil and the pupil range can be calculated. The relative coordinates are then multiplied by the ratio of the corresponding visual range width and height to the display screen width and height to calculate the coordinate position of the pupil on the display screen.
[0096] In this example, eye tracking and voice recognition are combined to achieve ultimate control. Voice recognition can identify control commands from the operator's voice data. Considering the sometimes noisy surgical environment, to ensure accurate control command recognition, pre-recorded voice information can be used for each operator to form a voiceprint corresponding to each operator. During voice recognition, the voiceprint can be used to identify the voice data belonging to the current operator from the acquired voice data. The voice data belonging to the current operator is then enhanced, while voice data not belonging to the current operator is removed or weakened. Voice recognition is then performed on the voice data belonging to the current operator to identify the current operator's voice control commands. Specifically, identifying control commands from the voice data can include: performing voiceprint recognition on the voice data to determine the voice content in the voice data that matches the target operator's voiceprint; and performing text recognition on the determined voice content to identify the control command. Pre-registering voiceprints can effectively improve recognition accuracy and avoid misoperation caused by incorrectly identifying control commands not belonging to the current operator.
[0097] Acquiring a real-time eye image of the target operator may include acquiring an eye image of the target operator through an eyepiece using a pre-installed camera. Specifically, a camera may be pre-installed and fixedly positioned facing the eyepiece, such that the camera can acquire the image in the eyepiece in real time. The eyepiece is a device used by the target operator to view images, and the target operator's eyes are facing the eyepiece.
[0098] In this example, only one eye can be used as the basis for the eye image during implementation. That is, only one of the eyepieces can be used to acquire the target image. For example, the right eye or the left eye can be used as the basis for image acquisition. The specific eye selected can be determined based on actual needs and circumstances, and this application does not impose specific restrictions on this. Using a single eye instead of two eyes can effectively avoid the problem of biased results and cumbersome calculations caused by selecting two eyes.
[0099] Specifically, the eyepiece may be that of a surgical robot, and the control instructions may be operating instructions for the display screen of the surgical robot during surgery. These operating instructions may allow the surgeon to call up CT images, 3D modeling images, or ultrasound images, and fuse them with existing endoscopic images to assist the surgeon in performing the surgery. Specifically, these instructions may include, but are not limited to, operations such as zooming in, zooming out, rotating, hiding, and displaying the fused image.
[0100] In this example, a surgical robot is also provided, which may include:
[0101] 1) surgical components, used by the target operator to perform surgical operations;
[0102] 2) a camera assembly, disposed opposite the eyepiece, for capturing an eye image of the target operator in the eyepiece while the target operator is performing a surgical operation;
[0103] 3) a processor connected to the camera assembly, configured to determine a coordinate position on a corresponding display screen based on the relative coordinate data of the pupil in the eye image;
[0104] 4) a display screen, in communication with the processor, for displaying operations during the surgical procedure;
[0105] 5) a sound receiving component for acquiring voice data of the target operator during the surgical operation;
[0106] The processor may also be configured to identify a control instruction from the voice data and control the execution of the control instruction at the coordinate position of the display screen.
[0107] That is, the control system combining the above-mentioned eye positioning and voiceprint control can be embedded into the software and hardware system of the above-mentioned doctor's console. During the operation, the camera device can obtain real-time image information of the doctor's eye area, and then convert it into position coordinates on the system display screen. The system can then be operated by the voice commands issued by the doctor. By combining the voice commands with the position coordinates on the display screen, precise control can be achieved. Specifically, voiceprint recognition can be added according to the accuracy requirements to improve the accuracy of voice recognition, and the correspondence between control commands and business execution models can be set, so as to achieve efficient control. That is, interactive control based on eye positioning and voiceprint control is carried out without the help of external devices. This eliminates the need to add an additional keyboard and mouse for the doctor, and can assist the doctor in performing more accurate operations.
[0108] Specifically, the surgical robot-based eye positioning and voiceprint control system provided in this example can embed voice and eye recognition hardware into the hardware structure of the doctor's console, which can include cameras and audio equipment. The software system module can be embedded in the main control end's operating system and run independently. The software system can include: pupil position calibration system module, eye positioning system module, speech semantic recognition system module, command parsing system module, and command execution module. The command execution module can be embedded in any laparoscope host and other extended hardware systems, thereby achieving richer and more comprehensive system control, not just image-based control.
[0109] The camera device can be installed at the corresponding position of the eyepiece in the robot main control terminal to capture the eye image. The image can generate a virtual coordinate system of the eyepiece. The movement of the pupil in this coordinate system can be converted into the coordinate movement of the real display screen. For example, it can be equivalent to the movement of the mouse on the display screen.
[0110] The above-mentioned sound receiving device can obtain voice information in the environment and recognize it as voice commands;
[0111] The eye calibration system module can pre-generate a position mapping between the eye's range of motion on the eyepiece (the eyepiece's virtual coordinate system) and the real display screen. Based on this mapping, the eye's movement can then be converted into movement on the real display screen.
[0112] The eye tracking system module dynamically identifies eyes and pupils based on images captured by the camera device, converts the pupil's position in the virtual coordinate system of the eyepiece, and then maps it to the position on the real display screen. This position simulates the movement of the mouse and, combined with voice control technology and specific command execution calls, can control the system.
[0113] Speech and semantic recognition system module, used to recognize sounds and analyze the semantics expressed by the sounds;
[0114] The command parsing system module analyzes the corresponding voiceprint pattern based on the user's voice semantic information and calls the corresponding command execution module based on the voiceprint pattern;
[0115] The command execution module can be configured based on specific business needs. Voiceprints and command controls can be assigned on a one-to-one basis. For example, different voiceprints can be configured for different operators, and possible control operations can be configured for different operators, thus forming a one-to-one voiceprint pattern. For example, operator A's control operation 1 corresponds to voiceprint pattern A-1, operator A's control operation 2 corresponds to voiceprint pattern A-2, operator B's control operation 1 corresponds to voiceprint pattern B-1, and operator B's control operation 2 corresponds to voiceprint pattern B-2. This one-to-one model effectively improves the accuracy of voice recognition and command execution. Commands can include, but are not limited to, at least one of the following: displaying a screen, moving a window, zooming in, zooming out, rotating, hiding, etc., thereby enabling the integrated display of medical data. In other words, multiple voiceprint patterns can be configured. After analyzing the voice information from the audio receiving device to determine the corresponding voiceprint pattern, the corresponding execution command in the local command execution module is invoked. If remote control is required, the corresponding execution command may be remotely invoked via TCP / IP.
[0116] Specifically, the relative coordinate information of the pupil can be calculated based on the captured image of the eye to calculate the coordinate position of the corresponding display screen, and the semantics can be identified and parsed based on the voiceprint information. The corresponding voiceprint pattern can be analyzed based on the parsed semantics to execute the corresponding execution instruction.
[0117] For example, if the pupil moves to position P on the screen and the voice command "Show CT image" is given, the CT image will be displayed on P. If the voice command "Show CT image in the upper right corner" is given, or "CT image width 200, height 180" is given, P will be moved to the upper right corner of the screen and the display will be modified to 200 pixels wide and 100 pixels high. In other words, display control is achieved through eye recognition and voice recognition.
[0118] Specifically, through eye tracking and voice recognition, during robotic surgery, the surgeon can manipulate the fused image and system resources in the image host using only their eyes and voice, without the need for a mouse, keyboard, or hands. For example, they can view and manipulate medical images, thereby improving surgical success rates. Specifically, during surgery, the surgeon can control the fusion display of medical images and the patient's laparoscopic surgical images, allowing them to gain a more comprehensive understanding of the surgical site and avoid risks, thereby improving surgical success rates.
[0119] like Figure 9 The figure shows the camera installation diagram of the system, which includes: camera 1 and eyepiece 2. The camera is used to capture the eye image of the operator's single eye, and the eyepiece is used to fix the operator's eye position. The operator views the display screen through the eyepiece. d represents the distance between the camera device and the eyepiece. Once fixed, it will not change, that is, the camera and the eyepiece are fixed. Figure 10 The figure shows a schematic diagram of the installation of the system's sound receiving device. The sound receiving device (i.e., microphone) can be installed directly below the eyepiece of the doctor's main control terminal to collect voice information in the environment.
[0120] pass Figure 9 The camera in the Figure 10 The audio receiving device in the device can recognize the voice information in the environment, pass the image to the eye positioning module for image analysis and convert it into position coordinates on the display screen, pass the voice information to the voiceprint recognition module for voice analysis to obtain execution instructions and be transmitted to the instruction execution module for instruction execution. Finally, the medical image is fused and displayed through the rendering module, and the video is output to the doctor's display screen (i.e., the micro display located at the doctor's end) through the laparoscope software for display control.
[0121] The coordinate system calibration operation is provided in this example in at least two stages. Stage 1: For different operators, eye alignment calibration is initiated upon the operator's first use, which can be considered the initial eye alignment calibration. Specifically, five bright spots (1, 2, 3, 4, and 5) can be displayed sequentially on the display screen. The operator follows the guidance and uses the pupil of the eye to locate these five guide points, pausing for 1 second each time. After locating all the bright spots in sequence, a rectangular eye visual range is determined, thereby achieving the mapping of the visual range and the screen position. Stage 2: Each time the operator's eyes leave and return to the eyepiece, considering that the facial position will be slightly offset compared to the initial calibration, a temporary eye alignment calibration can be performed. The temporary eye alignment calibration can only calibrate the center point. The visual range is assumed to remain unchanged. Therefore, only the center position guide point can be displayed to remind the operator to focus only on the center point. For example, this lasts for approximately 2 seconds, after which the red dot disappears. In this process, the coordinate system calibration is completed.
[0122] Specifically, whether the target operator's eyes leave and return to the eyepiece can be detected by using a pressure sensor provided at the edge of the eyepiece, or by performing pupil recognition on the image in the eyepiece. When it is determined that the target operator's eyes leave and return to the eyepiece, a center position guide point is displayed at the center of the target screen. An eye image of the target operator gazing at the center position guide point is obtained as a target eye image. Based on the position of the pupil in the target eye image, the correspondence between the eye image coordinate system and the screen coordinate system is calibrated, that is, temporary eye positioning calibration is performed.
[0123] After completing the above-mentioned initial eye positioning calibration and temporary eye positioning calibration, eye-based positioning operations can be performed. The movement of the eye on the eyepiece is mapped to the system coordinates on the display screen, and the voice information is parsed into execution instructions. The system combines the position coordinates and execution instructions to execute the corresponding execution instructions at the position coordinates. Specifically, the main control end can perform the parsing to obtain the position coordinates and execution instructions, and then pass the parsed position coordinates and execution instructions to the laparoscope host via TCP\IP for instruction execution, thereby performing specific operations. At the same time, the execution instructions can be transmitted to the laparoscope host through the network for image fusion operations, and can also perform remote system operations on other extended systems to achieve distributed control at the LAN level.
[0124] In this example, an eye-based positioning method is also provided. Figure 11 As shown, the following steps may be included:
[0125] Step 1101: Acquire a real-time eye image of a target object through an eyepiece;
[0126] Step 1102: determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the pupil in the corresponding target screen according to the relative coordinate data;
[0127] Step 1103: Display the corresponding coordinate position in the target screen in the form of an icon.
[0128] Among them, the eyepiece and the camera for obtaining eye images are set in fixed positions, that is, in a scenario where the head position is fixed relative to the camera, the position of the exit pupil can be determined through the real-time eye image of the target object, and then the position of the target object's line of sight in the display screen can be determined based on a pre-established mapping relationship, so that the target object can lock the target position on the screen without the help of hands, only by relying on the rotation of the eyes, that is, the movement control of the position in the display screen can be achieved by rotating the eyes.
[0129] Specifically, determining relative coordinate data between the pupil in the eye image and the eyepiece based on the eye image, and determining the coordinate position of the corresponding target screen based on the relative coordinate data, may include: identifying the position coordinates of the pupil in the eye image as first coordinate information; retrieving a pre-generated correspondence between the eye image coordinate system and the screen coordinate system as the first correspondence; and determining the coordinate position of the pupil in the eye image corresponding to the target screen based on the first coordinate information and the first correspondence. Determining the first correspondence can be achieved using the method described above, and this application will not elaborate further.
[0130] However, it is worth noting that the above is an explanation of the control method based on eye positioning and sound recognition, and the eye-based positioning method, using the application in a surgical robot as an example. In actual implementation, the above-mentioned control method based on eye positioning and sound recognition, and the eye-based positioning method can also be applied to other scenarios, such as microbial observation, fine parts processing, virtual reality, and other scenarios that require observation and operation with an eyepiece.
[0131] The method embodiments provided in the above embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on an electronic device as an example, Figure 12 This is a hardware structure diagram of an electronic device based on an eye positioning and voice recognition control method provided by this application. Figure 12 As shown, the electronic device 10 may include one or more (only one is shown in the figure) processors 02 (the processor 02 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 04 for storing data, and a transmission module 06 for communication functions. It will be understood by those skilled in the art that Figure 12 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 12 More or fewer components than shown, or with Figure 12 Different configurations shown.
[0132] The memory 04 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the control method based on eye positioning and voice recognition in the embodiment of the present application. The processor 02 executes various functional applications and data processing by running the software programs and modules stored in the memory 04, that is, implementing the control method based on eye positioning and voice recognition of the above-mentioned application. The memory 04 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 04 may further include a memory remotely located relative to the processor 02, and these remote memories may be connected to the electronic device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0133] The transmission module 06 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of the electronic device 10. In one embodiment, the transmission module 06 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 06 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0134] At the software level, control devices based on eye positioning and voice recognition can be Figure 13 As shown, this may include:
[0135] An acquisition module 1301 is used to acquire real-time eye images and voice data of a target operator through an eyepiece;
[0136] a determination module 1302 for determining relative coordinate data of the pupil in the eye image in the eyepiece based on the eye image, and determining a coordinate position of the corresponding target screen based on the relative coordinate data;
[0137] The execution module 1303 is configured to identify a control instruction from the voice data and execute the control instruction at a coordinate position in the corresponding target screen.
[0138] In one embodiment, the above-mentioned determination module 1302 can be specifically used to identify the position coordinates of the pupil in the eyepiece from the eye image as the first coordinate information; retrieve the correspondence between the pre-generated eye image coordinate system and the screen coordinate system as the first correspondence; and determine the coordinate position of the pupil in the eye image corresponding to the target screen based on the first coordinate information and the first correspondence.
[0139] In one embodiment, pre-generating a correspondence between the eye image coordinate system and the screen coordinate system may include: displaying a plurality of position guide points on the target screen; displaying guidance information for guiding the target operator to look at each of the plurality of position guide points one by one; obtaining a plurality of eye images through the eyepiece as the target operator looks at each position guide point one by one; and forming a correspondence between the eye image coordinate system and the screen coordinate system based on the positions of the pupils in the plurality of eye images and the position information of the corresponding position guide points on the screen.
[0140] In one embodiment, after obtaining multiple eye images of the target operator looking at each position guide point one by one through the eyepiece, it is also possible to determine whether the shape formed by the pupil position point in the multiple eye images is the same as the shape formed by the multiple position guide points; if it is determined that they are not the same, the multiple position guide points are triggered to be re-displayed on the target screen.
[0141] In one embodiment, after forming a correspondence between the eye image coordinate system and the screen coordinate system based on the positions of the pupils in the eyepiece in the multiple eye images and the position information of the corresponding position guide points in the screen, it is also possible to detect whether the eyes of the target operator leave and return to the eyepiece through a pressure sensor provided at the edge of the eyepiece; when it is determined that the eyes of the target operator leave and return to the eyepiece, a center position guide point is displayed at the center position of the target screen; an eye image of the target operator looking at the center position guide point through the eyepiece is obtained as the target eye image; and the correspondence between the eye image coordinate system and the screen coordinate system is calibrated according to the position of the pupil in the eyepiece in the target eye image.
[0142] In one embodiment, identifying the control instruction from the voice data may include: performing voiceprint recognition on the voice data to determine the voice content in the voice data that matches the voiceprint of the target operator; and performing text recognition on the determined voice content to determine the control instruction.
[0143] In one embodiment, the acquisition module 1301 may be specifically configured to acquire an eye image of the target operator through an eyepiece via an installed camera.
[0144] In one embodiment, the eyepiece is an eyepiece on a surgical robot, and the control instruction is an operation instruction for a display screen on the surgical robot during surgery.
[0145] In this example, an eyeball-based positioning device is also provided, which can be used to obtain a real-time eye image of a target object through an eyepiece; based on the eye image, the relative coordinate data of the pupil in the eye image in the eyepiece is determined, and the coordinate position in the corresponding target screen is determined based on the relative coordinate data; the corresponding coordinate position is displayed in the target screen in the form of an icon.
[0146] In one embodiment, the eyepiece and the camera for acquiring eye images are arranged in fixed positions.
[0147] In one embodiment, determining the relative coordinate data between the pupil and the eyepiece in the eye image based on the eye image, and determining the coordinate position in the corresponding target screen based on the relative coordinate data may include: identifying the position coordinates of the pupil in the eye image as first coordinate information; retrieving the correspondence between the pre-generated eye image coordinate system and the screen coordinate system as the first correspondence; and determining the coordinate position of the pupil in the eye image corresponding to the target screen based on the first coordinate information and the first correspondence.
[0148] The embodiments of the present application also provide a specific implementation of an electronic device capable of implementing all steps of the control method based on eye positioning and voice recognition in the above embodiment. The electronic device specifically includes the following contents: a processor, a memory, a communication interface, and a bus; wherein the processor, the memory, and the communication interface communicate with each other via the bus; the processor is used to call a computer program in the memory, and when the processor executes the computer program, all steps of the control method based on eye positioning and voice recognition in the above embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0149] Step 1: Acquire real-time eye image and voice data of the target operator through the eyepiece;
[0150] Step 2: determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining the coordinate position of the corresponding target screen according to the relative coordinate data;
[0151] Step 3: Identify a control instruction from the voice data, and execute the control instruction at the coordinate position in the corresponding target screen.
[0152] As can be seen from the above description, the embodiments of the present application obtain a real-time eye image of the target operator to determine the corresponding position on the target screen, and then identify the control command by acquiring voice data, thereby triggering the execution of the control command at the corresponding coordinate position on the target screen. This solution solves the existing problem of being unable to accurately control the device when the operator's hands are required to perform other operations, achieving the technical effect of accurate control based on eye position and voice commands.
[0153] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the control method based on eye positioning and voice recognition in the above-mentioned embodiment. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, all steps of the control method based on eye positioning and voice recognition in the above-mentioned embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0154] Step 1: Acquire real-time eye image and voice data of the target operator through the eyepiece;
[0155] Step 2: determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining the coordinate position of the corresponding target screen according to the relative coordinate data;
[0156] Step 3: Identify a control instruction from the voice data, and execute the control instruction at the coordinate position in the corresponding target screen.
[0157] As can be seen from the above description, the embodiments of the present application obtain a real-time eye image of the target operator to determine the corresponding position on the target screen, and then identify the control command by acquiring voice data, thereby triggering the execution of the control command at the corresponding coordinate position on the target screen. This solution solves the existing problem of being unable to accurately control the device when the operator's hands are required to perform other operations, achieving the technical effect of accurate control based on eye position and voice commands.
[0158] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the hardware + program embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0159] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0160] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment).
[0161] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0162] Although the present specification embodiment provides the method operation steps as described in the embodiment or flow chart, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiment is only one way in the order of execution of many steps and does not represent a unique execution order. When the device or terminal product in practice is executed, it can be performed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (such as a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements not only include those elements, but also include other elements not clearly listed, or also include elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements.
[0163] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing the embodiments of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules that implement the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0164] Those skilled in the art will also appreciate that, in addition to implementing the controller in pure computer-readable program code, it is entirely possible to implement the same functionality by logically programming the method steps in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered structures within the hardware component. Alternatively, the devices for implementing various functions can be considered both software modules implementing the method and structures within the hardware component.
[0165] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0166] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0168] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0169] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0170] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0171] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0172] Embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.
[0173] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the embodiments in this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.
[0174] The above description is merely an example of the embodiments of this specification and is not intended to limit the embodiments of this specification. For those skilled in the art, various modifications and variations of the embodiments of this specification are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of this specification shall be included within the scope of the claims of the embodiments of this specification.
Claims
1. A control method based on eye positioning and voice recognition, characterized in that: include: Acquire real-time eye images and voice data of the target operator through an eyepiece, wherein the eyepiece and the camera for acquiring the eye image are set in a fixed position; Determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the corresponding target screen according to the relative coordinate data; Recognizing a control instruction from the voice data, and executing the control instruction at a coordinate position in the corresponding target screen; Among them, based on the eye image, the relative coordinate data of the pupil in the eye image in the eyepiece is determined, and the coordinate position in the corresponding target screen is determined based on the relative coordinate data, including: identifying the position coordinates of the pupil in the eyepiece from the eye image as first coordinate information; retrieving the correspondence between the pre-generated eye image coordinate system and the screen coordinate system as the first correspondence; and determining the coordinate position of the pupil in the eye image corresponding to the target screen based on the first coordinate information and the first correspondence.
2. The method according to claim 1, characterized in that The correspondence between the eye image coordinate system and the screen coordinate system is pre-generated, including: displaying a plurality of position guide points on the target screen; displaying guidance information for guiding the target operator to look at each of the plurality of position guidance points one by one; Acquire multiple eye images of the target operator gazing at each position guide point through the eyepiece; A correspondence between the eye image coordinate system and the screen coordinate system is formed according to the positions of the pupils in the eyepiece in the plurality of eye images and the position information of the corresponding position guide points on the screen.
3. The method according to claim 2, characterized in that After acquiring a plurality of eye images of the target operator gazing at each position guide point one by one through the eyepiece, the method further includes: determining whether a shape formed by pupil position points in the plurality of eye images is the same as a shape formed by the plurality of position guide points; If they are determined to be different, a trigger is provided to redisplay the plurality of position guide points on the target screen.
4. The method according to claim 2, characterized in that After forming a correspondence between the eye image coordinate system and the screen coordinate system based on the positions of the pupils in the eyepiece in the plurality of eye images and the position information of the corresponding position guide points on the screen, the method further includes: Detecting whether the eyes of the target operator leave and then return to the eyepiece through a pressure sensor provided at the edge of the eyepiece; In the case where it is determined that the eyes of the target operator leave and then return to the eyepiece, displaying a center position guide point at the center position of the target screen; acquiring an eye image of the target operator gazing at the central position guide point through the eyepiece as a target eye image; The correspondence between the eye image coordinate system and the screen coordinate system is calibrated according to the position of the pupil in the eyepiece in the target eye image.
5. The method according to claim 1, wherein Recognizing the control instruction from the voice data includes: performing voiceprint recognition on the voice data to determine voice content in the voice data that matches the voiceprint of the target operator; The determined voice content is subjected to text recognition to determine the control instruction.
6. The method according to any one of claims 1 to 5, characterized in that The step of acquiring a real-time eye image of the target operator through the eyepiece includes: The eye image of the target operator through the eyepiece is obtained through the installed camera.
7. The method according to claim 6, characterized in that The eyepiece is an eyepiece on a surgical robot, and the control instruction is an operation instruction for a display screen on the surgical robot during surgery.
8. A positioning method based on eyeballs, characterized in that: include: Acquire a real-time eye image of the target object through the eyepiece; Determining relative coordinate data of the pupil in the eye image in the eyepiece according to the eye image, and determining a coordinate position of the corresponding target screen according to the relative coordinate data; Displaying the corresponding coordinate position in the target screen in the form of an icon; The eyepiece and the camera for acquiring eye images are set in fixed positions; Among them, according to the eye image, the relative coordinate data between the pupil in the eye image and the eyepiece is determined, and the coordinate position in the corresponding target screen is determined according to the relative coordinate data, including: identifying the position coordinates of the pupil in the eyepiece from the eye image as first coordinate information; retrieving the correspondence between the pre-generated eye image coordinate system and the screen coordinate system as the first correspondence; and determining the coordinate position of the pupil in the eye image corresponding to the target screen according to the first coordinate information and the first correspondence.
9. A surgical robot, characterized in that: include: A surgical component for a target operator to perform a surgical operation; a camera assembly, disposed opposite to the eyepiece, and configured to capture an eye image of the target operator in the eyepiece during the target operator performing a surgical operation; a processor connected to the camera assembly, and configured to determine a coordinate position on a screen of a corresponding display screen based on relative coordinate data of the pupil in the eye image in the eyepiece; a display screen, in communication with the processor, for displaying operations during surgery; The eyepiece and the camera for acquiring eye images are set in fixed positions; Among them, according to the eye image, the relative coordinate data between the pupil in the eye image and the eyepiece is determined, and the coordinate position in the corresponding target screen is determined according to the relative coordinate data, including: identifying the position coordinates of the pupil in the eyepiece from the eye image as first coordinate information; retrieving the correspondence between the pre-generated eye image coordinate system and the screen coordinate system as the first correspondence; and determining the coordinate position of the pupil in the eye image corresponding to the target screen according to the first coordinate information and the first correspondence.
10. The surgical robot according to claim 9, characterized in that: Also includes: a sound receiving component, configured to obtain voice data of the target operator during the surgical operation; The processor is further configured to identify a control instruction from the voice data and control the execution of the control instruction at a coordinate position of the display screen.
11. A surgical robot comprising a processor and a memory for storing instructions executable by the processor, characterized in that: When the processor executes the instructions, the steps of the method according to any one of claims 1 to 7 are implemented.
12. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Medical devices, systems and methods using eye gaze tracking for stereo viewer
CN110200702A
Medical devices, systems, and methods using eye gaze tracking
US20170172675A1