Air handwriting human-computer interaction technology based on monocular camera

By using aerial handwriting human-computer interaction technology based on a monocular camera, the problems of inconvenient device carrying and writing long texts in the air have been solved, improving comfort and universality, avoiding text superposition and character distortion, and enhancing the practicality of the system.

CN116311520BActive Publication Date: 2026-02-13ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310278218.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2026-02-13
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

Existing intelligent human-computer interaction technologies suffer from problems such as inconvenient device portability, significant limitations, low comfort, insufficient universality, and text overlap and character distortion when writing long texts in the air.

Method used

The system employs an aerial handwriting human-computer interaction technology based on a monocular camera. By acquiring real-time two-dimensional video images from the camera, it detects the hand movement area, obtains a complete binarized image of the hand movement area, combines the geometric features of the hand to obtain the hand contour, segments the palm and fingertip contours, and uses the fingertip contour features to complete the matching and generate aerial handwritten text.

Benefits of technology

It improves the comfort and universality of human-computer interaction systems, reduces device size and space limitations, avoids text overlap and character distortion when writing long texts in the air, and enhances the practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311520B_ABST
    Figure CN116311520B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent man-machine interaction, in particular to an air handwriting man-machine interaction technology based on a monocular camera, which comprises the following steps: acquiring a real-time two-dimensional video image based on the monocular camera and taking frames, detecting and acquiring a complete binary image of a hand movement area in the two-dimensional video image, acquiring a hand contour for hand segmentation based on the complete binary image of the hand movement area, acquiring a palm contour, segmenting the palm contour to acquire a fingertip contour, and determining handwriting start or end according to the number of the fingertip contour, completing fingertip matching to acquire fingertip coordinates based on fingertip contour features, and generating air writing text in response to fingertip movement and based on a coordinate system virtual sliding technology. The application improves the comfort and universality of a man-machine interaction system, reduces the limitations of the man-machine interaction system, avoids text superposition and character distortion problems when a user writes a long text in the air through the coordinate system virtual sliding technology, and thus enhances the practicability of the man-machine interaction system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent human-computer interaction, and in particular to an air handwriting human-computer interaction technology based on a monocular camera. BACKGROUND

[0002] With the rapid development of artificial intelligence, the development trend of human-computer interaction is more intelligent and friendly to people, that is, the less the human-computer interaction restricts people, the better. As one of the main carriers of information, text occupies a very important position in the field of human-computer interaction.

[0003] Most of the current common human-computer interaction methods are to use keyboard, mouse, touch screen or handwriting board and other contact type devices for information interaction, but due to the limitations of the number of keys, screen size and the like, it is not conducive to long text input, and it is not convenient to carry, and in virtual reality and smart large screen scenarios, the limitations are great.

[0004] Another human-computer interaction method is air handwriting interaction technology, as a new type of human-computer interaction method, because its writing method is more natural and humanized, it can allow users to perform non-contact human-computer interaction, has less restrictions, and provides users with a more comfortable and free experience. At present, the air handwriting interaction technology is mainly based on three types of systems, namely wearing / handheld, 3D sensors with depth information (such as Leap Motion, Kinect, etc.), and WIFI signals. However, the user wearing / handheld device for human-computer interaction will be affected by the connection line and the size of the handheld device, which will cause a certain sense of restraint to the user and cannot provide good comfort to the user; when air handwriting human-computer interaction is assisted by WIFI signal, the requirement for WIFI environment is relatively strict, and the limitations are great; and although the 3D sensor with depth information such as Leap Motion and Kinect can track and locate the fingertip position well, such 3D sensor is large in size and expensive in price, and has insufficient universality. And most of the current air handwriting technologies are for researching single characters or numbers, because when writing long text, the text will be overlapped, the character structure will be distorted, and the recognition difficulty will be increased. However, in actual application, users often need to write text continuously, and the practicality of the air handwriting interaction system is improved. SUMMARY

[0005] In order to solve the technical problems of inconvenient device carrying, great limitations, low comfort, insufficient universality and text overlapping and character distortion when air writing long text in the existing intelligent human-computer interaction technology, the present application provides an air handwriting human-computer interaction technology based on a monocular camera, improves the comfort and universality of the human-computer interaction system, reduces the limitations of the human-computer interaction system, avoids the problems of text overlapping and character distortion when the user writes long text in the air, and thus enhances the practicality of the human-computer interaction system.

[0006] To achieve the above object, the present application provides the following technical solutions.

[0007] An aspect of the present application is to provide a monocular camera-based air handwriting human-computer interaction technology, which comprises the following steps:

[0008] Obtaining a real-time two-dimensional video image of a camera, taking frames of the two-dimensional video image to obtain two-dimensional video image frames, and obtaining a two-dimensional video image sequence based on the continuous two-dimensional video image frames; wherein the camera is a monocular camera provided by a PC device or an external monocular camera.

[0009] Obtaining the two-dimensional video image in the two-dimensional video image sequence, detecting a hand motion region in the two-dimensional video image, and obtaining a hand motion region complete binary image based on hand motion information and color features in the hand motion region.

[0010] Obtaining a hand contour based on a hand geometric structure feature in the hand motion region complete binary image, performing hand segmentation on the hand contour, and obtaining a palm contour.

[0011] Further segmenting the palm contour to obtain a fingertip contour, judging the number of the fingertip contour, and determining the start or end of the writing according to the judgment result.

[0012] Completing fingertip matching based on the fingertip contour feature and obtaining a fingertip coordinate.

[0013] In response to the movement of the fingertip, obtaining a fingertip coordinate sequence passed in the movement of the fingertip based on a coordinate system virtual sliding technology, and generating air writing text according to the fingertip coordinate sequence.

[0014] Further, the obtaining of the hand motion region complete binary image based on the hand motion information and color features in the hand motion region comprises the following steps:

[0015] Processing the two-dimensional video image based on an average background difference method to obtain a hand motion region preliminary binary image.

[0016] Performing shadow detection on the two-dimensional video image based on an HSV color space to obtain a shadow detection result, performing shadow elimination on the hand motion region preliminary binary image based on the shadow detection result, and obtaining a hand motion region complete binary image.

[0017] Further, the processing of the two-dimensional video image based on the average background difference method to obtain the hand motion region preliminary binary image comprises the following steps:

[0018] The initial background is obtained by taking the average pixel value of the two-dimensional video image from several frames before the hand enters the detection range of the camera. Based on the initial background Update weight α and the current frame image. The background image is updated in real time during the background subtraction process to obtain the current background image. k≥1, where k represents the number of samples;

[0019] Based on the background image and the current frame image Obtaining difference images using background subtraction The difference image is obtained by setting a threshold th. Perform binarization processing to obtain a preliminary binarized image of the currently sampled hand movement region. .

[0020] Furthermore, the step of performing shadow detection on the two-dimensional video image based on the HSV color space to obtain shadow detection results, and then removing shadows from the preliminary binarized image of the hand movement region based on the shadow detection results to obtain a complete binarized image of the hand movement region, includes the following steps:

[0021] The background image and the current frame image The background images are obtained by converting the RGB color space to the HSV color space. and the current frame image The HSV color space has three components: chromaticity (H), saturation (S), and brightness (V).

[0022] Based on the background image and the current frame image Components in the HSV color space , , , , , Shadow detection is performed on the two-dimensional video image to obtain the shadow detection results. Based on the shadow detection results, a preliminary binarized image of the hand movement area is generated. Perform deshading to obtain a complete binarized image of the hand movement area.

[0023] Furthermore, the step of obtaining the hand contour based on the geometric structure features of the hand in the complete binary image of the hand movement region, and then segmenting the hand contour to obtain the palm contour includes the following steps:

[0024] The hand contour is obtained based on the hand motion region complete binary image, four points with maximum and minimum horizontal and vertical coordinates of the hand contour are connected in sequence to form a minimum circumscribed rectangle of the hand contour, the maximum distance from each point inside the hand contour to the hand contour is calculated, and the maximum value of all the minimum distances is taken as the maximum inscribed circle radius r of the hand contour. At this time, the center of the maximum inscribed circle is point O.

[0025] Two points on the minimum circumscribed rectangle of the hand contour , Coordinates of the center of the maximum inscribed circle and the maximum inscribed circle radius r, the approximate length d of the palm is obtained, the values of pixel points outside the palm are set to 0, and the palm contour is obtained.

[0026] Further, the palm contour is further segmented to obtain a fingertip contour, the number of the fingertip contour is judged, and the handwriting start or end is determined according to the judgment result, including the following steps:

[0027] The maximum inscribed circle radius r is expanded to 1.5 times to draw a circle on the palm contour, the center of the circle coincides with the center O of the maximum inscribed circle, the palm contour intersects with the circle, and the values of pixel points in the circle are set to 0 to obtain a fingertip contour;

[0028] The number N of the fingertip contour is counted, and the handwriting start or end is determined according to the number N of the fingertip contour.

[0029] Further, the fingertip matching is completed based on the fingertip contour feature, and the fingertip coordinates are obtained, including the following steps:

[0030] The fingertip contour is matched with a plurality of fingertip contour templates with different angles prepared in advance, and the fingertip contour template with the minimum matching value is selected. The position of the fingertip contour template in the two-dimensional video image is recorded, and the center of the fingertip contour template is defined as the fingertip coordinates.

[0031] Further, the coordinate system virtual sliding technology is:

[0032] After detecting each frame of the two-dimensional video image, the horizontal coordinates of all the fingertip coordinates recorded before the current frame are shifted to the left by a plurality of pixel units, the vertical coordinates are kept unchanged, the sliding character track is formed by connecting the shifted fingertip coordinates, and the air handwriting text is Q.

[0033] ;

[0034] ;​

[0035] ;

[0036] wherein, is the time interval between two adjacent points, and v is the lateral velocity of the fingertip coordinate.

[0037] A second aspect of the present application is to also provide an electronic device comprising:

[0038] at least one processor;

[0039] a memory in communication with the at least one processor;

[0040] wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0041] A third aspect of the present application is to also provide a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method as described above.

[0042] Compared with the prior art, the present application has the beneficial effects that:

[0043] From the above technical solution, the present application provides an air handwriting human-computer interaction technology based on a monocular camera, two-dimensional video images are collected based on a PC device self-provided monocular camera or an external monocular camera, a complete binary image of a hand motion region is acquired based on hand motion information and color features, a hand contour is acquired in combination with a hand geometric structure feature, a palm contour is obtained by segmenting the hand contour, a fingertip contour is obtained by further segmenting the palm contour, the start or end of text writing is determined by judging the number of the fingertip contour, fingertip matching is completed and fingertip coordinates are acquired based on the fingertip contour feature, fingertip coordinate sequences are acquired based on a coordinate system virtual sliding technology, air handwriting text is generated, a monocular camera which is cheap, small in size, and good in versatility is used, and the air handwriting human-computer interaction technology is not limited by the size of a device and the size of a space, the comfort and the universality of the human-computer interaction system are improved, the limitation of the human-computer interaction system is reduced, and through the coordinate system virtual sliding technology, the problem of text superposition and character distortion does not occur when a user writes a long text in the air, so that the practicality of the human-computer interaction system is enhanced.

[0044] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter.

[0045] The foregoing and other aspects, embodiments and features of the present teachings are more fully described below, in connection with the attached drawings. Other aspects, embodiments and features of the present teachings will become apparent from the following description, including the descriptions of the examples, and from the claims. BRIEF DESCRIPTION OF DRAWINGS

[0046] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical, or nearly identical, component that is illustrated in various figures is represented with a like numeral. For purposes of clarity, not every component is called out in every drawing. There is now being described, by way of example, embodiments of various aspects of the present teachings, with reference to the accompanying drawings, in which:

[0047] Figure 1 Flow chart of a monocular camera based air handwriting human-computer interaction technology embodiment of the present application;

[0048] Figure 2 Two-dimensional video image effect diagram in an embodiment of the present application;

[0049] Figure 3 Hand motion region preliminary binarization image effect diagram in an embodiment of the present application;

[0050] Figure 4 Shadow detection result effect diagram in an embodiment of the present application;

[0051] Figure 5 Hand motion region complete binarization image effect diagram in an embodiment of the present application;

[0052] Figure 6 Hand contour topology structure schematic diagram in an embodiment of the present application;

[0053] Figure 7 Fingertip contour number N effect diagram in an embodiment of the present application;

[0054] Figure 8 Fingertip positioning effect diagram in an embodiment of the present application;

[0055] Figure 9 Air handwriting text trajectory effect diagram in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in connection with the embodiments of the present application. Obviously, the described embodiments are a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the described embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without any creative effort belong to the protection scope of the present application.

[0057] Those skilled in the art of the technology can understand that all the terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs, unless otherwise defined. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with that in the context of the prior art, and should not be interpreted with an idealized or overly formal meaning unless defined as such.

[0058] At present, the human-computer interaction mode mainly includes two kinds, the first kind is based on keyboard, mouse, touch screen or handwriting board and other contact type device to carry out human-computer interaction, the second kind is based on wearing / handheld device, WIFI signal and 3D sensor with depth information air handwriting interaction technology, but the mode based on keyboard, mouse, touch screen or handwriting board and other contact type device to carry out human-computer interaction is limited by the number of keys, screen size, is not conducive to long text input, is inconvenient to carry, has great limitation, the air handwriting interaction technology based on wearing / handheld device, WIFI signal and 3D sensor with depth information has great limitation, low comfort, insufficient universality and the problem that long text air writing will cause text superposition and character distortion. Based on this, in the field of human-computer interaction technology, the embodiment of the application acquires two-dimensional video images based on the monocular camera of PC device or external monocular camera, obtains complete binary image of hand motion area based on hand motion information and color features, obtains hand contour in combination with hand geometry feature, carries out segmentation on hand contour to obtain palm contour, carries out further segmentation on palm contour to obtain fingertip contour, determines the start or end of text writing by judging the number of fingertip contours, completes fingertip matching and acquires fingertip coordinates based on fingertip contour features, acquires fingertip coordinate sequence based on coordinate system virtual sliding technology, generates air handwriting text, uses monocular camera which is cheap, small in size and good in universality, is not limited by device size and space size, improves the comfort and universality of human-computer interaction system, reduces the limitation of human-computer interaction system, and through coordinate system virtual sliding technology, the problem that long text air writing by the user does not appear text superposition and character distortion, thereby enhancing the practicability of human-computer interaction system.

[0059] Please refer to Figure 1 The embodiment of the application provides an air handwriting human-computer interaction technology based on monocular camera, which comprises the following steps:

[0060] Step S10: acquiring real-time two-dimensional video images of the camera, and the effect of the two-dimensional video images is as shown in Figure 2 The two-dimensional video images are taken to acquire two-dimensional video image frames, and a two-dimensional video image sequence is acquired based on continuous two-dimensional video image frames; wherein the camera is a monocular camera of PC device or an external monocular camera.

[0061] Step S20: obtaining the two-dimensional video image in the sequence of two-dimensional video images, detecting a hand motion region in the two-dimensional video image, and obtaining a hand motion region complete binary image based on hand motion information in the hand motion region and color features.

[0062] In some embodiments, the step of obtaining the hand motion region complete binary image based on the hand motion information in the hand motion region and the color features in step S20 comprises the following steps:

[0063] Step S201: processing the two-dimensional video image based on an average background difference method to obtain a hand motion region preliminary binary image.

[0064] In some embodiments, the step of processing the two-dimensional video image based on the average background difference method to obtain the hand motion region preliminary binary image comprises the following steps:

[0065] Step S2011: obtaining an initial background by taking a pixel mean value of the two-dimensional video images of the first 15 frames before the hand enters a camera detection range. updating the background image in real time in a background difference process based on the initial background , an update weight α, and a current frame image to obtain a current background image , where k≥1, and the k represents a sampling number, and the specific process can be represented by the following formula:

[0066]

[0067] wherein, represents a coordinate of a pixel point, , represents a background image of the kth and (k-1)th sampling.

[0068] Step S2012: obtaining a difference image by using background difference based on the background image and the current frame image , performing binary processing on the difference image by setting a threshold value th to obtain a hand motion region preliminary binary image of the kth sampling , and the hand motion region preliminary binary image has an effect as Figure 3 shown, and the specific process can be represented by the following formula:

[0069]

[0070]

[0071] wherein, represents a motion region, represents a background region.

[0072] Step S202: based on the HSV color space, shadow detection is performed on the two-dimensional video image to obtain a shadow detection result, and based on the shadow detection result, shadow elimination is performed on the hand motion region preliminary binarization image to obtain a hand motion region complete binarization image.

[0073] In some embodiments, the shadow detection based on the HSV color space, the two-dimensional video image, the shadow detection result, the shadow elimination based on the shadow detection result, the hand motion region preliminary binarization image, and the hand motion region complete binarization image include the following steps:

[0074] Step S2021: convert the background image and the current frame image from the RGB color space to the HSV color space to obtain the background image and the current frame image in the HSV color space, respectively.

[0075] Step S2022: based on the components of the background image and the current frame image in the HSV color space, , , , , , by selecting a suitable threshold to judge the shadow region in the two-dimensional video image, a shadow detection result is obtained. The effect is shown in Figure 4 , the shadow is removed from the hand motion region preliminary binarization image to obtain a hand motion region complete binarization image , and the hand motion region complete binarization image The effect is shown in Figure 5 , and the specific process can be represented by the following formula:

[0076] ;

[0077] ;

[0078] wherein, the shadow detection result represents a shadow region, Indicates the area of ​​motion. as well as The threshold value for brightness V. as well as These are the thresholds for saturation (S) and chroma (H), respectively.

[0079] Step S30: Based on the geometric features of the hand in the complete binary image of the hand movement area, obtain the hand contour, and perform hand segmentation on the hand contour to obtain the palm contour.

[0080] In some embodiments, obtaining the hand contour based on the geometric structural features of the hand in the complete binary image of the hand movement region, and segmenting the hand contour to obtain the palm contour includes the following steps:

[0081] Step S301: As Figure 6 As shown, the hand contour is obtained based on the topological structure of the complete binary image of the hand movement region. The origin of the coordinate system is placed at the upper left corner of the two-dimensional video image. The four points with the largest and smallest horizontal and vertical coordinates of the hand contour are connected in sequence to form the minimum bounding rectangle AEFB of the hand contour. The minimum distance from each point inside the hand contour to the hand contour is calculated. The maximum value among all distances is taken as the maximum inscribed circle radius r of the hand contour. At this time, the center of the maximum inscribed circle is point O.

[0082] Step S302: Based on the two points of the minimum bounding rectangle AEFB of the hand contour , Coordinates, center of the largest inscribed circle Using the coordinates and the maximum inscribed circle radius r, the approximate length d of the hand is obtained. The values ​​of pixels outside the hand are set to 0 to obtain the hand outline. The approximate length d of the hand is calculated using the following formula:

[0083]

[0084]

[0085] Where h is the perpendicular distance from the center O of the largest inscribed circle to line segment AB.

[0086] Step S40: Further segment the palm outline to obtain the fingertip outline, determine the number of fingertip outlines, and determine the start or end of the text writing based on the result of the determination.

[0087] In some embodiments, further segmenting the palm outline to obtain fingertip outlines, determining the number of fingertip outlines, and determining the start or end of the text writing based on the determination result includes the following steps:

[0088] Step S401: expand the maximum inscribed circle radius r to 1.5 times, draw a circle on the palm profile, the center of the circle coincides with the center of the maximum inscribed circle O, make the palm profile intersect with the circle and set the value of the pixel points in the circle to 0, and obtain the fingertip profile.

[0089] Step S402: count the number of fingertip profiles N, and determine the start or end of text writing according to the number of fingertip profiles N, as shown in the following table: Figure 7 When N=1, the text writing starts, and when N=5, the text writing ends.

[0090] Step S50: complete fingertip matching and obtain fingertip coordinates based on fingertip profile features.

[0091] In some embodiments, the fingertip matching and fingertip coordinates are obtained based on the fingertip profile features, including the following steps:

[0092] Step S501: make a plurality of fingertip profile templates with different angles in advance .

[0093] Step S502: match the fingertip profile with a plurality of fingertip profile templates with different angles, select the smallest matching value of the fingertip profile template , record the position of the fingertip profile template in the two-dimensional video image, and define the center of the fingertip profile template as the fingertip coordinates to realize fingertip positioning, the fingertip positioning effect is shown in the following table: Figure 8 The matching value is calculated as follows:

[0094]

[0095] Wherein, is the pixel point of the fingertip profile in the template matching process, and when the matching value is the smallest, the matching degree of the fingertip profile and the fingertip profile template is the highest.

[0096] Step S60: in response to the movement of the fingertip, obtain the fingertip coordinate sequence passed during the movement of the fingertip based on the coordinate system virtual sliding technology, and generate an air writing text according to the fingertip coordinate sequence, the trajectory effect of the air writing text is shown in the following table: Figure 9 .

[0097] Further, the coordinate system virtual sliding technology is specifically:

[0098] When each frame of the two-dimensional video image is detected, the horizontal coordinate of all the recorded fingertip coordinates before the current frame is translated left by 1 pixel unit, the vertical coordinate is kept unchanged, the sliding character track is formed by connecting the translated fingertip coordinates, and the air handwriting text is Q;

[0099] t=1,2,...T;

[0100] ;

[0101] ;

[0102] wherein, is the time interval between two adjacent points, and v is the horizontal translation speed of the fingertip coordinate, which can be set.

[0103] For example, the fingertip coordinate of the first frame of the two-dimensional video image is (x1, y1). When the second frame of the two-dimensional video image is detected, the fingertip coordinate of the first frame of the two-dimensional video image is translated left by 1 unit. At this time, , the fingertip coordinate of the second frame of the two-dimensional video image is (x2, y2). When the third frame of the two-dimensional video image is detected, the fingertip coordinate of the first frame of the two-dimensional video image is translated left by 1 unit. , At this time, , , the fingertip coordinate of the third frame of the two-dimensional video image is (x3, y3). .

[0104] Correspondingly, the embodiment of the application also provides an electronic device, comprising:

[0105] at least one processor;

[0106] a memory connected in communication with the at least one processor;

[0107] wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any embodiment of the application.

[0108] These computer programs can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the flow Figure 1 one flow or multiple flows and / or blocks Figure 1The steps of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM, flash memory, ROM, EEPROM, or any other non-transitory computer-readable medium, such as those listed above. An exemplary non-transitory computer-readable medium includes a floppy disk, CD-ROM, DVD, Blu-ray disc, hard disk drive, or any other non-transitory medium that can be used to store the desired information. The desired program code can be downloaded into the main memory 1200 from an external source via a computer system bus or network connection. Although a single software module is illustrated in FIG. 10, in other embodiments more than one software module can be used, or specific modules can be divided in sub-modules, combinations of these or other software modules can be used, or these can be stored on more than one computer system.

[0109] The processes described herein can be performed by software programs executable by a computer system or processor. Further, a computer-readable medium can include a storage medium having stored thereon computer-executable instructions. A computer-readable medium can include one or more memory devices and / or carrier waves that store data, which is

[0110] Accordingly, an embodiment of the application also provides a non-transitory computer readable storage medium having stored thereon computer instructions that, when executed by a computer, cause the computer to perform the method described in any of the embodiments of the application.

[0111] While this application has been described in conjunction with the preferred aspects thereof, it is understood that many modifications, changes and variations of the application can be made and will be obvious to one skilled in the art without departing from the spirit and scope of the application as defined in the following claims.

Claims

1. An aerial handwriting human-computer interaction technology based on a monocular camera, characterized in that, The technology includes the following steps: The system acquires real-time two-dimensional video images from a camera, extracts frames from the two-dimensional video images to obtain two-dimensional video image frames, and obtains a two-dimensional video image sequence based on consecutive two-dimensional video image frames; wherein, the camera is a built-in monocular camera of a PC device or an external monocular camera. The two-dimensional video images in the two-dimensional video image sequence are acquired, the hand movement region in the two-dimensional video images is detected, and a complete binarized image of the hand movement region is acquired based on the hand movement information and color features in the hand movement region. Based on the geometric features of the hand in the complete binary image of the hand movement region, the hand contour is obtained, and the hand contour is segmented to obtain the palm contour. This includes the following steps: The hand contour is obtained based on the topological structure of the hand contour in the complete binary image of the hand movement region; the four points with the largest and smallest x and y coordinates of the hand contour are connected sequentially to form the minimum bounding rectangle of the hand contour; the minimum distance from each point inside the hand contour to the hand contour itself is calculated; the maximum value among all the minimum distances is taken as the maximum inscribed circle radius r of the hand contour. At this point, the center of the maximum inscribed circle is point O; and the minimum bounding rectangle of the hand contour is used as the radius r. , Coordinates, center of the largest inscribed circle Using the coordinates and the maximum inscribed circle radius r, we obtain the approximate length d of the palm, set the values ​​of pixels outside the palm to 0, and get the palm outline. The palm outline is further segmented to obtain the fingertip outline, the number of fingertip outlines is judged, and the start or end of the writing is determined based on the judgment result; Fingertip matching and fingertip coordinates are obtained based on fingertip contour features; In response to fingertip movement and based on coordinate system virtual sliding technology, the fingertip coordinate sequence during the fingertip movement is obtained, and text written in the air is generated according to the fingertip coordinate sequence.

2. The aerial handwriting human-computer interaction technology based on a monocular camera according to claim 1, characterized in that, The process of obtaining a complete binarized image of the hand movement region based on hand movement information and color features in the hand movement region includes the following steps: The two-dimensional video image is processed using the average background subtraction method to obtain a preliminary binarized image of the hand movement area; Based on the HSV color space, shadow detection is performed on the two-dimensional video image to obtain the shadow detection result. Based on the shadow detection result, the shadow is eliminated in the preliminary binarized image of the hand movement area to obtain the complete binarized image of the hand movement area.

3. The aerial handwriting human-computer interaction technology based on a monocular camera according to claim 2, characterized in that, The process of processing the two-dimensional video image based on the average background subtraction method to obtain a preliminary binarized image of the hand movement region includes the following steps: The initial background is obtained by taking the average pixel value of the two-dimensional video image from several frames before the hand enters the detection range of the camera. Based on the initial background Update weight α and the current frame image. The background image is updated in real time during the background subtraction process to obtain the current background image. k≥1, where k represents the number of samples; Based on the background image and the current frame image Obtaining difference images using background subtraction The difference image is obtained by setting a threshold th. Perform binarization processing to obtain a preliminary binarized image of the currently sampled hand movement region. .

4. The aerial handwriting human-computer interaction technology based on a monocular camera according to claim 3, characterized in that, The process of performing shadow detection on the two-dimensional video image based on the HSV color space, obtaining shadow detection results, and then removing shadows from the preliminary binarized image of the hand movement region based on the shadow detection results to obtain a complete binarized image of the hand movement region includes the following steps: The background image and the current frame image The background images are obtained by converting the RGB color space to the HSV color space. and the current frame image The HSV color space has three components: chromaticity (H), saturation (S), and brightness (V). Based on the background image and the current frame image Components in the HSV color space , , , , , Shadow detection is performed on the two-dimensional video image to obtain the shadow detection results. Based on the shadow detection results, a preliminary binarized image of the hand movement area is generated. Perform deshading to obtain a complete binarized image of the hand movement area.

5. The spatial handwriting human-computer interaction technology based on a monocular camera according to claim 1, characterized in that, The process of further segmenting the palm outline to obtain fingertip outlines, determining the number of fingertip outlines, and determining the start or end of the writing based on the determination result includes the following steps: Enlarge the radius r of the maximum inscribed circle to 1.5 times and draw a circle on the palm outline. The center of the circle coincides with the center O of the maximum inscribed circle. Make the palm outline intersect with the circle and set the value of the pixel in the circle to 0 to obtain the fingertip outline. The number of fingertip outlines N is counted, and the start or end of text writing is determined by the number of fingertip outlines N.

6. The aerial handwriting human-computer interaction technology based on a monocular camera according to claim 1, characterized in that, The process of matching fingertips and obtaining their coordinates based on fingertip contour features includes the following steps: The fingertip contour is matched with several pre-made fingertip contour templates at different angles, and a matching value is selected. The smallest fingertip contour template is used to record the position of the fingertip contour template in the two-dimensional video image, and the center of the fingertip contour template is defined as the fingertip coordinates.

7. The aerial handwriting human-computer interaction technology based on a monocular camera according to claim 1, characterized in that, The virtual sliding coordinate system technique is as follows: After detecting each frame of the two-dimensional video image, the horizontal coordinates of all the fingertip coordinates recorded before the current frame are shifted to the left by a certain number of pixels while keeping the vertical coordinates unchanged. By connecting the shifted fingertip coordinates, a sliding character trajectory is formed, and the handwritten text in the air is Q. ; ; ; in, The time interval between two adjacent points is given by , and v is the lateral movement speed of the fingertip coordinates.

8. An electronic device, comprising: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1 to 7.