Human-computer interaction method and system integrating AI voice and intelligent feedback
Through the human-computer interaction method that integrates AI voice and intelligent feedback, and using AI voiceprint recognition and visual detection models, command control is realized by moving forward and backward without the need for fingers to leave the mouse, which solves the problems of inconvenient gesture operation and inability to intelligent feedback in the existing technology, and improves the user experience.
Patent Information
- Application Number
- CN202510586543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In the prior art, when performing gesture operations, hands need to switch back and forth between the mouse and the camera, affecting the user experience, and cannot achieve intelligent feedback for different users.
Through the human-computer interaction method that integrates AI voice and intelligent feedback, the AI voiceprint recognition model is used to identify and obtain interaction rate parameters, and the intermediate camera and left and right reflectors collect finger sliding information in real time, determine the fingertip contour through the visual detection model and judge the distance from the mouse surface, so as to realize gesture detection and control of the two fingers separately.
It realizes command control by moving the fingers on the left and right sides without leaving the mouse without leaving the mouse, improving the user experience, and providing intelligent feedback to different users.
Smart Images

Figure CN120103983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction software technology, and in particular to a human-computer interaction method and system integrating AI voice and intelligent feedback. Background Art
[0002] As a traditional human-computer interaction device, the mouse can realize basic operations such as pointer movement and left and right clicks. On the basis of these basic operations, shortcut keys can also be set on the basis of the mouse. By installing mouse management software on the operating system, when the shortcut key is pressed, the keyboard combination key can be simulated to be pressed, thereby realizing more and more complex shortcut key functions. At present, gesture recognition and interaction can also be realized based on the camera collection gestures of the operating system, but this kind of reliance on the camera of the operating system, and the hand has to switch back and forth between the mouse and the camera during gesture operation, which affects the user experience. It is also impossible to achieve adaptation for different users and intelligent feedback for different users. Summary of the invention
[0003] To this end, it is necessary to provide a human-computer interaction method and system that integrates AI voice and intelligent feedback to solve the existing problems of having to switch back and forth between the mouse and the camera when performing gesture operations and being unable to provide intelligent feedback for different users.
[0004] To achieve the above object, the present invention provides a human-computer interaction method integrating AI voice and intelligent feedback, comprising the following steps: Acquire the operator's voice information and identify the operator through the AI voiceprint recognition model; After identification, the interaction rate parameters corresponding to the operator are obtained; The middle camera on the mouse and the left and right reflectors are used to obtain photos containing the left and right finger sliding on the mouse in real time; The photo is divided into left and right sides, and the visual detection model is used to detect the fingertip contours of the left and right fingers of each side of the segmented photo; After obtaining the bottom position of the fingertip contour and the mouse surface less than a preset distance, the scrolling distance and direction on the left and right sides are calculated in real time according to the interaction rate parameter, the finger sliding distance and direction, and control instructions containing the scrolling distance and direction on the left and right sides are sent to the first area and the second area of the operating system interface respectively.
[0005] Furthermore, the method further comprises the steps of: After obtaining the bottom positions of the left and right fingertip contours and contacting the mouse surface at the same time; After the bottom positions of the left and right fingertip contours are greater than the preset distance from the mouse surface at the same time within the preset first interval; After the bottom positions of the left and right fingertip contours touch the mouse surface again at the same time within the preset second interval; After the bottom positions of the left and right fingertip contours are greater than the preset distance from the mouse surface again within the preset third interval; Trigger the preset two-finger control command to the operating system.
[0006] Furthermore, the interaction rate parameter initialization step is also included: Get the initialization instruction, obtain the operator's voice information and the entered interaction rate parameters, extract the voiceprint feature information through the AI voiceprint recognition model and save the corresponding relationship with the entered interaction rate parameters.
[0007] Furthermore, after obtaining that the bottom position of the fingertip contour is greater than a preset distance from the mouse surface, the bending angle changes of the left and right fingers are obtained, and corresponding control instructions are sent to the operating system according to the bending angle changes of the fingers.
[0008] Furthermore, the method further comprises the steps of: After obtaining the voice recognition command triggered by the mouse button, the mouse obtains the operator's voice information and performs voice recognition through the AI voice recognition model. After the voice recognition is the preset control command, the control command is sent to the operating system.
[0009] Furthermore, the first area is a left area of the mouse pointer, and the second area is a right area of the mouse pointer.
[0010] Furthermore, the first area is the left area in the middle of the screen, and the second area is the right area in the middle of the screen.
[0011] Further, the first area is a first screen area, the second area is a second screen area, and the first screen and the second screen are different screens.
[0012] Furthermore, the AI voiceprint recognition model is a local voiceprint recognition model or an online API voiceprint recognition model.
[0013] The present invention provides a human-computer interaction system integrating AI voice and intelligent feedback, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of any method described in the present invention are implemented.
[0014] Different from the existing technology, the above technical solution uses the AI voiceprint recognition model to identify the user's voiceprint and determine the interaction rate parameters, realizing intelligent feedback for different users, and then the middle camera can respectively collect the side images of the left and right fingers placed on the human-computer interaction device (mouse), and determine the fingertip contour according to the visual detection model and then judge the distance from the mouse surface, so as to realize gesture detection of the two fingers respectively, and correspondingly control different areas of the operating system. At this time, the hand does not need to move away from the mouse, and the command control can be realized by moving the left and right fingers back and forth, which improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A method flow chart of an embodiment of the present invention; Figure 2 A physical diagram of the human-computer interaction device of the present invention; Figure 3 is a cross-sectional view of the human-computer interaction device of the present invention; Figure 4 The present invention is a method flow chart of another embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0017] See also Figures 1 to 4 The present invention provides a human-computer interaction method integrating AI voice and intelligent feedback, which can be applied in Figure 2-Figure 3 On a human-computer interaction device (such as a mouse), the human-computer interaction device includes a device body 1, and the surface of the device body 1 protruding from the upper middle position includes an image acquisition module 2. The image acquisition module 2 includes left and right reflectors 20 at the top and a bottom camera 21. The bottom camera 21 is located in the middle position directly below the left and right reflectors. In order to capture the images of the left and right fingers into the bottom camera 21, a convex lens may also be included in front of the reflector. The bottom of the device body 1 may also include a photoelectric sensor 3 for collecting displacement and a main control chip 4 at the bottom. The main control chip is used to collect camera and sensor information and convert it into corresponding control instructions, and then send it to the operating system via wired (USB cable) or wireless (Bluetooth or 2.4G wireless) methods.
[0018] See also Figure 1As shown, the method of the present invention includes the following steps: Step S101 obtains the voice information of the operator and performs identity recognition through the AI voiceprint recognition model. Here, the sound can be collected through a camera connected to the operating system or a microphone can be set on the human-computer interaction device to collect the sound, and then the AI voiceprint recognition model obtains the sound and recognizes it, and the recognized voiceprint feature information is compared with the pre-stored voiceprint feature information to determine whether there is the same voiceprint information. It should be noted that the AI voiceprint recognition model is a local voiceprint recognition model or an online API voiceprint recognition model. The local voiceprint model can be deployed on the operating system or on the main control chip 4 of the human-computer interaction device. The local voiceprint model can be a single-state hidden Markov model (HMM) or a Gaussian mixture Markov model (GMM), and the online API voiceprint recognition model such as Yunzhisheng or Alibaba Cloud voiceprint recognition API interface can realize voiceprint recognition by sending the sound to the API interface. The sound information here can be the sound information of the operating system guiding the user to say a preset text, so that rapid recognition can be achieved.
[0019] After voiceprint recognition, the process proceeds to step S102 to obtain the interaction rate parameters corresponding to the operator after identity recognition. Each operator can store its corresponding interaction rate parameters. Step S103 uses the middle camera on the mouse and the left and right mirrors to obtain photos containing the left and right finger sliding on the mouse in real time. Figure 3As shown, the left finger (index finger) operating the left button of the mouse will be reflected into the left half of the lens of the lower camera through the left reflector, and the right finger (middle finger) operating the right button of the mouse will be reflected into the right half of the lens of the lower camera through the right reflector, realizing the bilateral acquisition of a single lens, where the image of the side of the finger, that is, the image between the finger and the finger, is acquired. Then enter step S104 to divide the photo into left and right sides, and use the visual detection model to detect the fingertip contours of the fingers on the left and right sides of the divided photo on each side. Preferably, it is necessary to mirror-flip the image of one side so that the fingers on both sides face the same direction, which is convenient for the rapid processing of the visual detection model. The visual detection model can use the YOLOv8 model, and the trained YOLOv8 model can be deployed in the main control chip, and the main control chip can use Rockchip's RK3588 chip. After deployment, the collected pictures are segmented and mirrored and input into the YOLOv8 model, and then the fingertip contours are obtained. The training process of the YOLOv8 model can be performed as follows: first, the camera is used to collect side images of fingers at different positions, different bends, and different users, and then each image is opened using labelImg, and then manually annotated, the outline of the last section of the finger is marked with a box, and saved as an annotation file corresponding to the image. Then, most (such as 90%) images are placed in the training folder of the YOLOv8 model, and a small part of the images are placed in the verification folder of the YOLOv8 model. The saved annotation files and images are placed in the marked training folder and verification folder, and then the YOLOv8 model training is started. After training, the trained model file can be saved, and the model file and the YOLOv8 model are placed in the memory of the main control chip, and then the model is called to realize the fingertip contour position.
[0020] Then, in step S105, after the bottom position of the fingertip contour is obtained and the distance from the mouse surface is less than a preset distance (such as less than 3 pixels), it means that the finger is in contact with the mouse surface, and the command triggering process begins. The scrolling distance and direction of the left and right sides are calculated in real time according to the interaction rate parameter, the finger sliding distance and direction, and the control instructions containing the scrolling distance and direction of the left and right sides are sent to the first area and the second area of the operating system interface respectively. If the left finger slides down by 5 pixels and the interaction rate parameter is 10, then 50 pixels can be obtained by multiplication, and the first area of the operating system can be scrolled down by 50 pixels. If the right finger slides up by 7 pixels and the interaction rate parameter is 10, then 70 pixels can be obtained by multiplication, and the second area of the operating system can be scrolled up by 70 pixels. Thus, operations in different areas are realized. For different operators, if the interaction rate parameter is 12, the left finger slides down by 5 pixels, and 60 pixels can be obtained by multiplication, and the first area of the operating system can be scrolled down by 60 pixels. Intelligent feedback for different users is realized.
[0021] The above solution uses the AI voiceprint recognition model to identify the user's voiceprint and determine the interaction rate parameters to achieve intelligent feedback for different users. Then, the middle camera can collect the side images of the left and right fingers placed on the human-computer interaction device (mouse) respectively, and determine the fingertip contour based on the visual detection model and then judge the distance from the mouse surface, thereby achieving gesture detection of the two fingers respectively, and correspondingly controlling different areas of the operating system. At this time, the hand does not need to move away from the mouse, and command control can be achieved by moving the left and right fingers back and forth, improving the user experience.
[0022] In some embodiments, gesture operation can also be achieved by clicking with left and right fingers, such as Figure 4As shown, it also includes the steps: Step S401 obtains the bottom positions of the left and right fingertip contours and contacts the mouse surface at the same time. The "simultaneous" here does not mean exactly the same moment, but it can be less than a certain interval, such as less than 100 milliseconds, which can be considered as simultaneous. Then step S402 is after the bottom positions of the left and right fingertip contours and the mouse surface are simultaneously greater than the preset distance within the preset first interval time (such as 200 milliseconds); that is, the fingertips leave the mouse surface, but the hand does not leave the mouse. Then step S403 is after the bottom positions of the left and right fingertip contours and the mouse surface are contacted again at the same time within the preset second interval time (such as 200 milliseconds); and step S404 is after the bottom positions of the left and right fingertip contours and the mouse surface are again greater than the preset distance within the preset third interval time (such as 200 milliseconds); enter step S405 to trigger the preset two-finger control instruction to the operating system. The two-finger control instruction can be preset as a shortcut key, a combination key or other control instructions, such as page turning, returning to the desktop and other instructions. In some embodiments, this two-finger control instruction can be used as Figure 1 In the embodiment, the trigger instruction of step 105 is performed after the two fingers perform such an operation, and the scroll control instruction judgment and sending operation of step 105 is performed, otherwise the scroll control instruction judgment is not performed to avoid erroneous operation. Then, after performing steps S401 to S405 again, the trigger instruction can be cancelled, and the scroll control instruction judgment of step S105 is not performed. Figure 1 The trigger instruction of step 105 of the embodiment may also be implemented by a button set on the human-computer interaction device. The trigger is performed after the button is pressed, and the trigger is canceled after the button is pressed again.
[0023] In order to have corresponding interaction rate parameters after identity recognition, the present invention also includes an interaction rate parameter initialization step: obtaining an initialization instruction, such as a user clicking an initialization button on the software, obtaining the operator's voice information and the entered interaction rate parameters, extracting voiceprint feature information through an AI voiceprint recognition model, and saving the corresponding relationship with the entered interaction rate parameters. In this way, the voiceprint feature information and the corresponding interaction rate parameters can be saved. Of course, if no corresponding voiceprint feature information is detected, the default interaction rate parameter (such as 10) can also be used.
[0024] Furthermore, the present invention can also realize more gesture instructions. After obtaining the bottom position of the fingertip contour and the mouse surface greater than the preset distance, that is, after the two fingers are lifted, the bending angle change of the left and right fingers is obtained, and the corresponding control instructions are sent to the operating system according to the bending angle change of the fingers. The bending of the fingertips can be detected by a visual detection model, such as the YOLOv8 model. First, the box corresponding to the outline of the finger is obtained, and the box will frame the entire finger. Since the aspect ratio (height divided by width) of the box will be close to larger when the finger is bent, and the box will become flat when the finger is straightened, the aspect ratio will become smaller, and the aspect ratio change of the box of the entire finger can be obtained by judging and detecting, and the detection of finger bending can be realized. If the finger changes from straight to bent, the next page operation can be triggered. The specific scheme is to detect that within a preset time, the change in the aspect ratio of the box from small to large is greater than a certain value, then the next page operation is triggered, and the corresponding left finger sends the next page operation to the first area, and the corresponding right finger sends the next page operation to the second area. If it is detected that the change in the aspect ratio of the box from large to small is greater than a certain value within a preset time, the previous page operation is triggered, that is, the finger is bent and then straightened for the previous page, and vice versa for the next page, thus realizing the finger's page turning operation.
[0025] In order to further realize voice operation, the step is also included: after obtaining the voice recognition instruction triggered by the mouse button, such as pressing the voice recognition button. The mouse obtains the voice information of the operator and performs voice recognition through the AI voice recognition model (different from the voiceprint model). After the voice recognition is a preset control instruction, the control instruction is sent to the operating system. Voice recognition is now relatively mature. The focus of the present invention is that triggering and voice recognition can be realized through the buttons and microphone on the mouse, such as controlling the sound of the operating system to become louder or smaller.
[0026] It should be noted that the present invention does not limit the positions of the first area and the second area, which can be set according to the user. In some embodiments, the first area is the left area of the mouse pointer, and the second area is the right area of the mouse pointer. In this way, after moving the mouse, the mouse pointer moves, and the first area and the second area can be changed at any time, which is convenient for user operation.
[0027] In some embodiments, the first area is the left area in the middle of the screen, and the second area is the right area in the middle of the screen. In this way, the operating system can quickly place different documents or web pages in the left half screen area and the right half screen area of the screen, and the left finger can control the document in the left half screen, and the right finger can control the document in the right half screen, which is convenient for operation.
[0028] In some embodiments, the first area is a first screen area, the second area is a second screen area, and the first screen and the second screen are different screens. In this way, multiple connected screens can be controlled separately, which is convenient for users to control. It should be noted that the scrolling of the present invention can also be a picture carousel scrolling, so that the first area can play multiple pictures, and the second area can also play multiple pictures. The picture can be quickly scrolled to the one you want to view by using the left and right fingers.
[0029] The present invention provides a human-computer interaction system integrating AI voice and intelligent feedback, including a memory and a processor, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the method described in any one of the present invention are implemented. The memory of this embodiment may be a memory set in an electronic device, and the electronic device may read the contents of the memory and implement the effects of the present invention. The memory may also be a separate memory, and the memory is connected to the electronic device, so that the electronic device can read the contents of the memory and implement the method steps of the present invention. The system of the present invention recognizes the user's voiceprint and determines the interaction rate parameters through an AI voiceprint recognition model, thereby realizing intelligent feedback for different users, and then the side images of the left and right fingers placed on the human-computer interaction device (mouse) can be collected respectively through the middle camera, and the distance from the mouse surface is judged after the fingertip contour is determined according to the visual detection model, thereby realizing gesture detection of the two fingers respectively, and correspondingly different areas of the operating system can be controlled. At this time, the hand does not need to move the mouse away, and command control can be realized by moving the left and right fingers forward and backward, thereby improving the user experience.
[0030] It should be noted that, although the above embodiments have been described in this article, the patent protection scope of the present invention is not limited thereby. Therefore, based on the innovative concept of the present invention, changes and modifications made to the embodiments described herein, or equivalent structures or equivalent process changes made using the contents of the present invention specification and drawings, directly or indirectly applying the above technical solutions to other related technical fields, are all included in the patent protection scope of the present invention.
Claims
1. A human-computer interaction method integrating AI voice and intelligent feedback, characterized in that: The steps include: Acquire the operator's voice information and identify the operator through the AI voiceprint recognition model; After identification, the interaction rate parameters corresponding to the operator are obtained; The middle camera on the mouse and the left and right reflectors are used to obtain photos containing the left and right finger sliding on the mouse in real time; The photo is divided into left and right sides, and the visual detection model is used to detect the fingertip contours of the left and right fingers of each side of the segmented photo; After obtaining the bottom position of the fingertip contour and the mouse surface less than a preset distance, the scrolling distance and direction on the left and right sides are calculated in real time according to the interaction rate parameter, the finger sliding distance and direction, and control instructions containing the scrolling distance and direction on the left and right sides are sent to the first area and the second area of the operating system interface respectively.
2. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: Also includes the steps: After obtaining the bottom positions of the left and right fingertip contours and contacting the mouse surface at the same time; After the bottom positions of the left and right fingertip contours are greater than the preset distance from the mouse surface at the same time within the preset first interval; After the bottom positions of the left and right fingertip contours touch the mouse surface again at the same time within the preset second interval; After the bottom positions of the left and right fingertip contours are greater than the preset distance from the mouse surface again within the preset third interval; Trigger the preset two-finger control command to the operating system.
3. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: It also includes the interaction rate parameter initialization step: Get the initialization instruction, obtain the operator's voice information and the entered interaction rate parameters, extract the voiceprint feature information through the AI voiceprint recognition model and save the corresponding relationship with the entered interaction rate parameters.
4. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: After obtaining that the bottom position of the fingertip contour is greater than a preset distance from the mouse surface, the bending angle changes of the left and right fingers are obtained, and corresponding control instructions are sent to the operating system according to the bending angle changes of the fingers.
5. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: Also includes the steps: After obtaining the voice recognition command triggered by the mouse button, the mouse obtains the operator's voice information and performs voice recognition through the AI voice recognition model. After the voice recognition is the preset control command, the control command is sent to the operating system.
6. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: The first area is the left area of the mouse pointer, and the second area is the right area of the mouse pointer.
7. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: The first area is the left area in the middle of the screen, and the second area is the right area in the middle of the screen.
8. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: The first area is a first screen area, the second area is a second screen area, and the first screen and the second screen are different screens.
9. The human-computer interaction method integrating AI voice and intelligent feedback according to claim 1, characterized in that: The AI voiceprint recognition model is a local voiceprint recognition model or an online API voiceprint recognition model.
10. Human-computer interaction system integrating AI voice and intelligent feedback, characterized by: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Human-machine interaction method and device based on sight tracing and gesture discriminating
CN101344816A
Television man-machine interaction method based on handwriting input and fingertip mouse
CN102184021A
Method for realizing disembodied virtual mouse based on monocular vision
CN104699243A
A method for recognizing ponter control commands based on finger motions on the mobile device and a mobile device which controls ponter based on finger motions
KR101189633B1
Devices for use with computers
US20150193023A1