Interactive intelligent tablet control method and device and storage medium
By identifying and analyzing the coordinates of the human body key points in the image, determining the interactive intention and interaction points of the interactive smart tablet, the problem of unnatural and convenient traditional operation methods is solved, and a more natural and convenient interactive smart tablet operation is achieved.
Patent Information
- Application Number
- CN202311573345.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
The operation of interactive smart tablets in traditional classrooms relies on touch screens or mouse, which makes the operation not natural and convenient enough.
By detecting the human body-wide area in the current image, identifying the coordinates of the target object's key point, and determining the interaction intention and interaction points of the target object to the interactive smart tablet based on these coordinates.
It realizes the natural, intuitive and convenient operation of interactive intelligent tablets through body movements, improving the naturalness and convenience of operation.
Smart Images

Figure CN120029442A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an interactive intelligent tablet control method, device and storage medium. Background Art
[0002] The interaction method of interactive smart tablets in traditional classrooms mainly relies on external devices such as touch screens or mice. Teachers need to directly touch the interactive smart tablets or use mice to operate them, which is not natural and convenient enough. Summary of the invention
[0003] The present application proposes an interactive smart tablet control method, device and storage medium, which can alleviate the problem in the related art that the operation of using a touch screen or a mouse to interact with an interactive smart tablet is not natural and convenient enough.
[0004] The first embodiment of the present application provides an interactive smart tablet control method, including:
[0005] Detect the human body area in the current image;
[0006] Identify current coordinates of key points of a target object within the human body range area;
[0007] In a case where it is confirmed based on the current coordinates that the target object has an intention to interact with the interactive smart tablet, based on the current coordinates, calculating a position coordinate for representing a position of the target object relative to the interactive smart tablet;
[0008] Based on the position coordinates, an interaction point of the target object with respect to the interactive smart tablet is determined.
[0009] In some embodiments, confirming that the target object has an intention to interact with the interactive smart tablet based on the current coordinates includes:
[0010] Based on the current coordinates, predict a first interaction intention parameter and a second interaction intention parameter, wherein the first interaction intention parameter is used to characterize whether the gesture of the target object has an interaction intention with the interactive smart tablet, and the second interaction intention parameter is used to characterize whether the human body posture of the target object has an interaction intention with the interactive smart tablet;
[0011] When the first interaction intention parameter and the second interaction intention parameter satisfy a preset interaction intention trigger condition, it is confirmed that the target object has an interaction intention with the interactive smart tablet.
[0012] In some embodiments, predicting the first interaction intention parameter based on the current coordinates includes:
[0013] Based on the current coordinates, determining a hand range area of the target human body in the current image;
[0014] Recognize the image within the hand range to obtain a gesture classification result;
[0015] The gesture classification result is matched with gestures in a preset static human gesture library to obtain the first interaction intention parameter.
[0016] In some embodiments, based on the current coordinates, determining the hand range area of the target human body in the current image includes:
[0017] Based on the hand range area in the historical image and the hand motion parameters obtained based on the historical image, predict the first hand range area in the current image; based on the current coordinates and the preset hand center distance, detect the second hand range area in the current image;
[0018] Position fusion is performed on the first hand range area and the second hand range area to obtain the hand range area.
[0019] In some embodiments, based on the current coordinates and a preset hand center distance, detecting a second hand range area in the current image includes:
[0020] Based on the current coordinates, determining the extension line of the elbow and the arm;
[0021] Based on the extension line and the preset hand center distance, a diameter circle is determined; the intersection of the diameter circle and the extension line is the position of the elbow, the radius of the diameter circle is the preset hand center distance, and the center of the diameter circle is located on the extension line and at the hand of the current image;
[0022] The area enclosed by the circumscribed rectangle of the diameter circle is used as the second hand range area.
[0023] In some embodiments, performing position fusion on the first hand range area and the second hand range area to obtain the hand range area includes:
[0024] Performing position fusion on the first hand range area and the second hand range area to obtain a fused area;
[0025] The fusion area is expanded outward by a preset ratio to obtain the hand range area.
[0026] In some embodiments, performing position fusion on the first hand range area and the second hand range area to obtain a fused area includes:
[0027] When the intersection-and-union ratio of the first hand range area and the second hand range area is greater than the intersection-and-union ratio threshold, obtaining an overlapping area of the first hand range area and the second hand range area, and using the overlapping area as the fusion area;
[0028] When the IoU is less than or equal to the IoU threshold, the credibility of the current coordinates is obtained; when the credibility is greater than the credibility threshold, the first hand range area is used as the fusion area; when the credibility is less than or equal to the credibility threshold, the second hand range area is used as the fusion area.
[0029] In some embodiments, predicting a second interaction intention parameter based on the current coordinates includes:
[0030] Based on the current coordinates, determining a classification result of the arm posture of the target object;
[0031] The arm posture classification result is matched with the postures in a preset static human posture library to obtain the second interaction intention parameter.
[0032] In some embodiments, based on the current coordinates, calculating the position coordinates used to characterize the position of the target object relative to the interactive smart tablet includes:
[0033] When the posture of the preset human body parameter model is the i-th posture, projecting the key points in the human body parameter model into the device coordinate system to obtain the coordinates of the projected key points; the device corresponding to the device coordinate system is used to collect the current image;
[0034] For each key point of the target object, calculating the coordinate difference between the current coordinate of each key point and the projected key point coordinate of each key point;
[0035] Based on the coordinate differences, calculating a coordinate difference mean;
[0036] When the error is greater than or equal to the error threshold, i=i+1 is updated until the error is less than the error threshold, and the coordinates of the projected key point when the error is less than the error threshold are used as the position coordinates.
[0037] In some embodiments, determining the interaction point of the target object with the interactive smart tablet based on the position coordinates includes:
[0038] Calculating the position coordinates of the starting point of the virtual ray based on the position coordinates and the preset wrist limit distance;
[0039] The position coordinates of the interaction point are calculated based on the position coordinates of the starting point of the virtual ray and the position coordinates of the hand in the current coordinates.
[0040] In some embodiments, based on the position coordinates and a preset wrist limit distance, calculating the position coordinates of the starting point of the virtual ray includes:
[0041] Calculating a first quotient of the limit distance of the interactive smart tablet and the width of the interactive smart tablet, and calculating the quotient of the z-axis coordinate in the position coordinate and the first quotient to obtain the z-axis coordinate of the starting point of the virtual ray;
[0042] Calculating a second quotient of the x-axis coordinate and the z-axis coordinate in the position coordinate, and calculating the product of the second quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the x-axis coordinate of the starting point of the virtual ray;
[0043] Calculating a third quotient of the y-axis coordinate and the z-axis coordinate in the position coordinate, and calculating the product of the third quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the y-axis coordinate of the starting point of the virtual ray;
[0044] The three-dimensional coordinates composed of the x-axis coordinate, y-axis coordinate and z-axis coordinate of the starting point of the virtual ray are used as the position coordinates of the starting point of the virtual ray.
[0045] In some embodiments, identifying current coordinates of key points of a target object within the human body range region includes:
[0046] Cropping the current image to obtain a sub-image corresponding to the human body range area;
[0047] Using a pre-trained human key point model to identify the sub-image, and predicting the coordinates of the predicted key points of the target object;
[0048] Based on the coordinates of the human body range area, the coordinates of the predicted key points are translated and transformed to obtain the actual coordinates of the predicted key points;
[0049] The actual coordinates are screened based on the human body range area to obtain the current coordinates.
[0050] The second aspect of the present application provides an interactive smart tablet control device, including:
[0051] A detection module is used to detect the human body area in the current image;
[0052] An identification module, used for identifying the current coordinates of key points of a target object within the human body range;
[0053] A calculation module, configured to calculate, based on the current coordinates, position coordinates for representing the position of the target object relative to the interactive smart tablet when it is confirmed that the target object has an intention to interact with the interactive smart tablet based on the current coordinates;
[0054] A determination module is used to determine the interaction point of the target object with respect to the interactive smart tablet based on the position coordinates.
[0055] An embodiment of the third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0056] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect above.
[0057] The technical solution provided in the embodiments of the present application has at least the following technical effects or advantages:
[0058] In an embodiment of the present application, the current coordinates of the key points of the target object in the current image are identified, and based on the current coordinates, it is determined whether the target object has an intention to interact with the interactive smart tablet. If there is an intention to interact, the position coordinates of the target object are calculated based on the current coordinates, and finally the interaction point of the target object relative to the interactive smart tablet is determined based on the position coordinates. The scheme of the embodiment of the present application realizes the recognition of the interaction intention of the target object through the recognition of the key points of the human body, and finally realizes the interaction with the interactive smart tablet. Since the coordinates of the key points of the human body reflect the body movements of the target object, for the target object, it is realized to operate the focus of the interactive smart tablet through body movements, and the operation is natural, intuitive and convenient.
[0059] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] By reading the detailed description of the preferred embodiment below, various other advantages and benefits will become clear to those of ordinary skill in the art. The accompanying drawings are only used for the purpose of illustrating the preferred embodiment and are not considered to be limitations of the present application. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings.
[0061] In the attached picture:
[0062] Figure 1 A flow chart of an interactive smart tablet control method provided by an embodiment of the present application is shown;
[0063] Figure 2 A schematic diagram showing a human body range area provided by an embodiment of the present application is shown;
[0064] Figure 3 A schematic diagram showing key points provided by an embodiment of the present application;
[0065] Figure 4 A schematic diagram showing key points including a hand diameter circle provided by an embodiment of the present application is shown;
[0066] Figure 5 A schematic diagram showing the principle of calculating the starting point position coordinates of a virtual ray provided in an embodiment of the present application is shown;
[0067] Figure 6 A schematic diagram showing the principle of calculating the position coordinates of an interaction point provided by an embodiment of the present application is shown;
[0068] Figure 7 A structural diagram of an interactive smart tablet control device provided by an embodiment of the present application is shown;
[0069] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application is shown;
[0070] Fig. 9 A schematic diagram of a storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0071] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0072] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by technicians in the field to which this application belongs.
[0073] Smart classroom implementation technology refers to how to apply emerging intelligent technologies such as the Internet of Things, big data, cloud computing, and "Internet +" to the construction of smart classrooms from the perspective of educational technology, so as to achieve the transformation of traditional classrooms from physical space environment to intelligence. With the development and introduction of new technologies, the selection of smart classrooms from design to equipment has gradually diversified, focusing more on user experience, people-oriented, and using intelligence to promote interactive communication between teachers and students.
[0074] Interactive smart tablets are an indispensable part of smart classrooms. In current daily teaching, teachers use interactive smart tablets to play slides for teaching. In related technologies, teachers use input devices such as keyboards, mice, or laser pens to interact with interactive devices. When teachers walk around the classroom, they are limited by the effective range of the laser pen and the fixed input position of the input device, making it difficult for teachers to interact with the interactive smart tablet in a timely and effective manner. That is, the use of input devices to interact with interactive smart tablets in related technologies has the problem of inconvenient interaction.
[0075] In order to alleviate the technical problems existing in the related art, the embodiments of the present application propose a control method, device and storage medium for an interactive smart tablet. The method detects the current coordinates of the key points of the target object in the current image, uses the current coordinates to determine the target object's interaction intention with the interactive smart tablet, and when it is confirmed that the target object has the interaction intention with the interactive smart tablet, determines the target object's interaction point with the interactive smart tablet according to the position coordinates that characterize the position of the target object relative to the interactive smart tablet. The scheme of the embodiments of the present application realizes the recognition of the target object's interaction intention through the recognition of the key points of the human body, and finally realizes the interaction with the interactive smart tablet. Since the coordinates of the key points of the human body reflect the body movements of the target object, for the target object, it is realized to operate the focus of the interactive smart tablet through body movements, and the operation is natural, intuitive and convenient.
[0076] A method for controlling an interactive smart tablet according to an embodiment of the present application is described below in conjunction with the accompanying drawings. The method can be applied to an electronic device, and the electronic device includes a processing module and an interactive smart tablet. The processing module can control the interactive smart tablet.
[0077] like Figure 1 As shown, the method specifically comprises the following steps:
[0078] Step 101, detecting the human body range area in the current image;
[0079] Step 102, identifying the current coordinates of the key points of the target object within the human body range;
[0080] Step 103: When it is confirmed based on the current coordinates that the target object has an intention to interact with the interactive smart tablet, a position coordinate for representing the position of the target object relative to the interactive smart tablet is calculated based on the current coordinates;
[0081] Step 104: Determine the interaction point of the target object with the interactive smart tablet based on the location coordinates.
[0082] In the application, a human body detection model can be used to detect the human body range area in the current image, and the human body detection model includes but is not limited to the YOLO model, the SSD model, the Faster R-CNN model, etc. The human body detection model detects the human body range area in the current image, which may include data preparation, data preprocessing, model selection, model training, model evaluation, parameter tuning, model deployment, real-time detection, post-processing, and visualization of detection results.
[0083] In the data preparation stage, multiple images or video data containing a human body are collected. It should be understood that the collected multiple images can come from images collected in real time by a camera, a pre-set image library, etc.
[0084] In the data preprocessing stage, the images in the multiple images or video data are preprocessed, including but not limited to resizing the images, normalizing the images, and / or denoising the images. Data preprocessing of the images helps to improve the performance and accuracy of the human detection model.
[0085] In the model selection stage, a model is selected from the YOLO model, SSD model, and Faster R-CNN model according to the requirements and resource constraints.
[0086] In the model training phase, the human body detection model selected in the model selection phase is trained using the images after data preprocessing. During the training process, the human body detection model learns to recognize the features and bounding boxes of the human body.
[0087] In the model evaluation phase, the trained human detection model is evaluated using the test set. Evaluation indicators may include precision, recall, and accuracy.
[0088] In the parameter tuning stage, according to the evaluation results of the model evaluation stage, the parameters of the human detection model are tuned to improve the accuracy and model performance.
[0089] In the model deployment phase, the trained human body detection model is deployed to actual applications. The human body detection model can be embedded in mobile devices, servers, or cloud platforms.
[0090] In the real-time detection stage, the deployed model is used to detect the human body area in the current image. The human body detection model outputs the detected human body position and bounding box.
[0091] In the post-processing stage, the detection results of the human detection model are post-processed, such as using non-maximum suppression (NMS) to eliminate overlapping bounding boxes and filter out detection results with higher confidence.
[0092] In the visualization stage, the detection results are visualized, and the human body range area, labeling category and other information are drawn on the current image to facilitate user observation and analysis.
[0093] It should be noted that human detection is a complex task involving many details and technologies. In actual applications, the process may need to be adjusted and optimized according to specific needs. At the same time, different human detection models and algorithms may have different processes and steps. Therefore, the specific human detection process may vary depending on the application scenario and the selected model.
[0094] In an example, taking the use of the YOLO model to detect the human body range area as an example, it can be implemented by following the steps below:
[0095] Download the YOLO model: Download the YOLO model file for human detection from the YOLO official website or GitHub. YOLO has multiple versions, such as YOLOv3, YOLOv4, etc. You can choose the appropriate version according to your needs.
[0096] Install dependent libraries: Before using YOLO, you need to install some necessary dependent libraries, such as OpenCV, NumPy, etc. These libraries can be installed through the pip command.
[0097] Load the model: Use the corresponding library (such as Darknet) to load the downloaded YOLO model file. This will create a model object that can be used for subsequent human detection tasks.
[0098] Image preprocessing: Before human detection, the current image to be detected needs to be preprocessed. This includes adjusting the image size, normalizing, converting to the format required by the model, etc.
[0099] Perform human detection: Use the loaded model object to perform human detection on the preprocessed image. Input the image into the model, and the model will return the detected human bounding box and its corresponding confidence.
[0100] Post-processing: The detection results can be post-processed as needed, such as filtering bounding boxes with low confidence and performing non-maximum suppression.
[0101] Visualization results: The detection results can be visualized by drawing bounding boxes, annotating categories, and other information on the image to more intuitively display the results of human body detection.
[0102] It should be noted that the use of the YOLO model may involve some details and parameter adjustments, such as the model's threshold, input image size, etc. Depending on specific needs and application scenarios, these parameters may need to be adjusted to achieve better detection results.
[0103] In this embodiment, the human key point model can be used to identify the coordinates of the key points of the target object in the human body range. In the application, the human key point model includes but is not limited to the OpenPose model, the HRNet model, etc. These models can predict the coordinates of the key points of the human body in the image, such as the coordinates of the key points of the head, arms, legs, etc.
[0104] In an optional embodiment, identifying the current coordinates of the key points of the target object within the human body range may include:
[0105] Crop the current image to obtain a sub-image corresponding to the human body area;
[0106] Use the pre-trained human key point model to recognize the sub-image and predict the coordinates of the predicted key points of the target object;
[0107] Based on the coordinates of the human body range area, the coordinates of the predicted key points are translated and transformed to obtain the actual coordinates of the predicted key points;
[0108] The actual coordinates are filtered based on the human body range area to obtain the current coordinates.
[0109] It should be understood that when the human key point model performs recognition prediction on a sub-image, the number of predicted key points may be greater than the number of key points actually present in the sub-image. For example, if the sub-image only includes the area above the shoulders of the target object, when the human key point model performs prediction on the sub-image, the human key point model can predict key points such as the left waist and the right waist based on the key points such as the left shoulder and the right shoulder included in the sub-image. Therefore, after obtaining the actual coordinates of the predicted key points, the actual coordinates need to be screened based on the human body range area to obtain the current coordinates of the key points located in the human body range area.
[0110] It should be understood that since the cropped sub-image is relative to the bounding box of the human body region, the coordinates of the predicted key points need to be converted back to the original image coordinate system. When performing the translation conversion, the coordinates of the predicted key points can be added to the starting coordinates of the human body region.
[0111] It should be understood that when there are multiple objects in the current image, each object corresponds to a human body range area. For an example, see Figure 2 , Figure 2 The area enclosed by the box marked with "person" is the human body area.
[0112] In applications, as an example, Figure 3As shown in FIG, the key points of the human body are usually represented by 17 joints, namely, nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles. After the human body range area is detected and obtained, the current image can be cropped based on the human body range area to obtain a sub-image that only includes the image encircled by the human body range area, and then the sub-image is recognized to identify the key points in the sub-image.
[0113] In an optional embodiment, confirming that the target object has an intention to interact with the interactive smart tablet based on the current coordinates may include the following steps:
[0114] Based on the current coordinates, predict a first interaction intention parameter and a second interaction intention parameter, wherein the first interaction intention parameter is used to characterize whether the gesture of the target object has an interaction intention with the interactive smart tablet, and the second interaction intention parameter is used to characterize whether the human body posture of the target object has an interaction intention with the interactive smart tablet;
[0115] When the first interaction intention parameter and the second interaction intention parameter satisfy the preset interaction intention triggering condition, it is confirmed that the target object has the interaction intention with the interactive smart tablet.
[0116] It should be understood that when judging whether the gesture of the target object has the intention to interact with the interactive smart tablet, if the gestures of both the left and right hands of the target object are detected at the same time, then as long as the gesture of one hand represents the intention to interact with the interactive smart tablet, it can be determined that the gesture of the target object has the intention to interact with the interactive smart tablet.
[0117] It should be understood that there may be some misjudgment when using gestures alone to judge the interaction intention, because the shapes and movements of gestures may be similar, making it difficult to accurately distinguish different interaction intentions. Combining the intention of human body posture can provide more information, thereby improving the accuracy of judgment. Using human body posture alone to determine the interaction intention may be affected by environmental factors, such as lighting conditions, occlusion, etc., resulting in inaccurate posture estimation. Combining the intention of gestures can make up for the inaccuracy of posture estimation and improve the robustness of the system. In addition, combining the intention of gestures and human body posture can provide richer interaction methods. Gestures can provide an intuitive and natural way of interaction, while human body posture can provide more complex interactive actions and postures, thereby increasing the flexibility of interaction between users and the system. Finally, by combining the intention of gestures and human body posture, the user's interaction intention can be understood more accurately, thereby providing a more accurate and personalized interaction experience. This can increase user satisfaction with the system and improve user experience.
[0118] In this embodiment, the interaction intention trigger condition is used to determine whether the target object in the current image has an interaction intention with the interactive smart tablet. The first interaction intention parameter and the second interaction intention parameter satisfying the interaction intention trigger condition may be that the first interaction intention parameter and the second interaction intention parameter respectively satisfy the interaction intention trigger condition, that is, when the first interaction intention parameter and the second interaction intention parameter respectively satisfy the interaction intention trigger condition, the first interaction intention parameter and the second interaction intention parameter satisfy the preset interaction intention trigger condition. As an example, when the first interaction intention parameter represents that the gesture of the target object has an interaction intention with the interactive smart tablet, it is determined that the first interaction intention parameter meets the interaction intention trigger condition. When the second interaction intention parameter represents that the human body posture of the target object has an interaction intention with the interactive smart tablet, it is determined that the second interaction intention parameter meets the interaction intention trigger condition.
[0119] Of course, the weighted result of the first interaction intention parameter and the second interaction intention parameter may also satisfy the interaction intention trigger condition. In this case, the first interaction intention parameter and the second interaction intention parameter may be weighted to obtain a weighted calculation result, and when the weighted calculation result satisfies the interaction intention trigger condition, it is determined that the first interaction intention parameter and the second interaction intention parameter satisfy the preset interaction intention trigger condition.
[0120] As an example, the first interaction intention parameter and the second interaction intention parameter include but are not limited to expressing the target object's intended interaction degree with the interactive smart tablet in the form of a score. Among them, the higher the score, the more obvious the target object's interaction intention. Accordingly, when judging whether the first interaction intention parameter and the second interaction intention parameter meet the preset interaction intention trigger condition, the first interaction intention parameter and the second interaction intention parameter can be weighted according to the preset weight, and then the weighted result can be compared with the preset threshold. When the weighted result is greater than the preset threshold, it is determined that the interaction intention trigger condition is met, otherwise it is determined that the interaction intention trigger condition is not met. It should be understood that the weight parameter for weighting can be set artificially according to needs, and this embodiment does not limit this.
[0121] In an optional embodiment, predicting the first interaction intention parameter based on the current coordinates may include:
[0122] Based on the current coordinates, determine the hand range area of the target human body in the current image;
[0123] Recognize the image in the hand range area to obtain the gesture classification result;
[0124] The gesture classification result is matched with the gestures in a preset static human gesture library to obtain a first interaction intention parameter.
[0125] It should be understood that after the hand range area of the target person is identified, the current image is screenshotted based on the hand range area to obtain a sub-image that only includes the hand of the target person, and then the sub-image is identified to obtain a gesture classification result.
[0126] In this embodiment, the gesture classification results include but are not limited to five fingers open, thumbs up, ok, scissors hand, etc.
[0127] In this embodiment, the preset static human gesture library also includes multiple gestures. The degree of the target object's interaction intention with the interactive smart tablet can be determined based on the degree of matching between the gesture classification result and the gesture in the static human gesture library. For example, when there is a gesture in the static human gesture library that is the same as the gesture classification result, determining the first interaction intention parameter indicates that the target object has an interaction intention with the interactive smart tablet. When there is no gesture in the static human gesture library that is exactly the same as the gesture classification result, the degree of the interaction intention represented by the first interaction intention parameter can be determined based on the similarity between the gesture classification result and the gesture in the static human gesture library. It should be understood that the higher the degree of similarity between the gesture classification result and the gesture in the static human gesture library, the more obvious the interaction intention represented by the first interaction intention parameter.
[0128] In this embodiment, when the current video frame is the first frame image captured by the image acquisition device on the interactive smart tablet, the second hand range area in the current image is detected based on the current coordinates and the preset hand center distance, and the second hand range area is used as the hand range area of the target human body in the current image. In specific implementation, after obtaining the current coordinates of the key point, the extension line of "left elbow-left arm" and / or "right elbow-right arm" is determined; taking the extension line of "left elbow-left arm" as an example, the left wrist on the extension line is used as the intersection of the diameter circle and the extension line, and the preset hand center distance is used as the radius of the diameter circle. The diameter circle is determined based on the intersection and the radius. The center of the diameter circle is located on the extension line and on the hand of the current image, and the circumscribed rectangle of the diameter circle is used as the second predicted hand range box. For an example, please refer to Figure 4 , Figure 4 The circle in the middle is the hand diameter circle.
[0129] In this embodiment, in order to improve the accuracy of determining the hand range area of the target human body, when the current image has a historical image, the hand range area of the target human body can also be determined in combination with the historical image.
[0130] In an optional embodiment, based on the current coordinates, determining the hand range area of the target human body in the current image includes:
[0131] Based on the hand range area in the historical image and the hand motion parameters obtained based on the historical image, predict the first hand range area in the current image; based on the current coordinates and the preset hand center distance, detect the second hand range area in the current image;
[0132] The first hand range area and the second hand range area are positionally fused to obtain the hand range area. It should be understood that the historical image is an image captured a period of time before the current image is captured.
[0133] The hand motion parameters include but are not limited to the hand motion speed, motion trajectory and other parameters.
[0134] In this embodiment, the hand range area in the historical image can be identified and tracked by a hand detection algorithm or model, and the hand range area in each frame of the image can be recorded. Based on the data of the hand range area in the historical image, the movement speed and trajectory of the hand can be calculated. The movement speed is obtained by calculating the position change of the center point of the hand range area, and the movement trajectory is obtained by calculating the change of the boundary point of the hand range area.
[0135] In some embodiments, based on the current coordinates and a preset hand center distance, detecting a second hand range area in the current image may include the following steps:
[0136] Based on the current coordinates, determine the extension lines of the elbow and arm;
[0137] Based on the extension line and the preset hand center distance, a diameter circle is determined; the intersection of the diameter circle and the extension line is the position of the elbow, the radius of the diameter circle is the preset hand center distance, and the center of the diameter circle is located on the extension line and at the hand of the current image;
[0138] The area enclosed by the circumscribed rectangle of the diameter circle is taken as the second hand range area.
[0139] It should be understood that the human body range area of the current image may include multiple second hand range areas. For example, when the human body range area only includes the left elbow and left arm of the target object, then the human body range area includes one second hand range area, that is, the second hand range area obtained based on the left elbow and the left arm. For another example, when the human body range area simultaneously includes the left elbow, left arm, right elbow and right arm of the target object, then the human body range area includes two second hand range areas, that is, the second hand range area obtained based on the left elbow and the left arm, and the second hand range area obtained based on the right elbow and the right arm.
[0140] In some embodiments, performing position fusion on the first hand range area and the second hand range area to obtain the hand range area may include:
[0141] Performing position fusion on the first hand range area and the second hand range area to obtain a fusion area;
[0142] The fusion area is expanded by a preset ratio to obtain the hand range area.
[0143] It should be understood that expanding the fusion area by a preset ratio can improve boundary coverage, shape adaptability, stability, and robustness.
[0144] Regarding improving the boundary coverage rate, there may be a certain error between the first hand range area and the second hand range area, resulting in the boundary not completely covering the edge of the hand. By performing a preset ratio expansion, the boundary of the hand range area can be expanded, thereby better capturing the boundary information of the hand and improving the boundary coverage rate.
[0145] Regarding improving shape adaptability, the shape of the hand may change to a certain extent, especially in different gestures and movements. The first hand range area and the second hand range area may not be able to fully adapt to the change in the shape of the hand. By performing a preset ratio expansion, the change in the hand shape can be adapted to a certain extent, thereby improving shape adaptability.
[0146] Regarding improving stability, hand movements and posture changes may cause jitter and instability in the hand range area. By expanding the preset ratio, the jitter of the hand range box can be reduced and the stability can be improved. The expansion of the fusion area can smooth the boundaries, reduce unnecessary changes, and make the hand range area more stable.
[0147] Regarding increasing robustness, the first prediction of the hand range area may have a certain error, especially in complex scenes or under poor lighting conditions. By performing a preset expansion, the size of the hand range area can be increased, thereby increasing the robustness of hand detection and tracking and reducing the possibility of misjudgment.
[0148] In this embodiment, when the fusion area is expanded outward by a preset ratio, the centroid (i.e., the center point) of the fusion area is found, and then the fusion area is expanded proportionally to both sides. Specifically, the width and height after expansion can be calculated according to the preset ratio, and the expanded hand range area can be reconstructed with the centroid as the center.
[0149] In some embodiments, performing position fusion on the first hand range area and the second hand range area to obtain a fused area may include the following steps:
[0150] When the intersection-and-union ratio of the first hand range area and the second hand range area is greater than the intersection-and-union ratio threshold, obtaining an overlapping area of the first hand range area and the second hand range area, and using the overlapping area as a fusion area;
[0151] When the IoU is less than or equal to the IoU threshold, the credibility of the current coordinates is obtained. When the credibility is greater than the credibility threshold, the first hand range area is used as the fusion area. When the credibility is less than or equal to the credibility threshold, the second hand range area is used as the fusion area.
[0152] It should be understood that when the coordinates of key points are identified through the human body key point model, the human body key point model not only outputs the coordinates of the key points, but also synchronously outputs the credibility (confidence) of the coordinates. The credibility is used to characterize the reliability and accuracy of the coordinates of the key points output by the human body key point model.
[0153] In an optional embodiment, predicting the second interaction intention parameter based on the current coordinates includes:
[0154] Based on the current coordinates, determine the arm posture classification result of the target object;
[0155] The arm posture classification result is matched with the posture in the preset static human posture library to obtain the second interaction intention parameter.
[0156] The arm posture classification results include but are not limited to using a neural network model to determine the arm posture classification results. The arm posture classification results include but are not limited to arm raised, arm relaxed or arm hanging down.
[0157] Among them, the postures in the preset static human posture library include but are not limited to raising the right hand, raising both hands, raising the left hand, etc.
[0158] In this embodiment, the degree of obviousness of the interaction intention expressed by the second interaction intention parameter can be determined based on the degree of matching between the arm posture classification result and the posture in the preset static human posture library. It should be understood that the higher the degree of matching between the posture existing in the preset static human posture library and the arm posture classification result, the more obvious the interaction intention expressed by the second interaction intention parameter.
[0159] In an optional embodiment, based on the current coordinates, calculating the position coordinates used to characterize the position of the target object relative to the interactive smart tablet includes:
[0160] When the posture of the preset human body parameter model is the i-th posture, the key points in the human body parameter model are projected into the device coordinate system to obtain the coordinates of the projected key points; the device corresponding to the device coordinate system is used to collect the current image;
[0161] For each key point of the target object, calculate the coordinate difference between the current coordinate of each key point and the projected key point coordinate of each key point;
[0162] Based on the coordinate differences, the mean of the coordinate differences is calculated;
[0163] When the error is greater than or equal to the error threshold, i=i+1 is updated until the error is less than the error threshold, and the coordinates of the projected key point when the error is less than the error threshold are used as the position coordinates.
[0164] In this embodiment, the preset human body parameter model can be a standard human body parameter model for Asians. The standard human body parameter model for Asians here can be a human key point skeleton model with a height of 168 cm and a shoulder width of 39 cm. The key points on the human key point skeleton model have movable hinges, so when the movable hinges are rotated, the posture of the human key point skeleton model will change.
[0165] In applications, the error threshold may be preset manually based on experience or according to actual needs, and this embodiment does not specifically limit this.
[0166] In this embodiment, when projecting the human body parameter model to the device coordinate system, the human body parameter model can be first projected to the three-dimensional coordinate system of the device, and then projected from the three-dimensional coordinate system to the two-dimensional coordinate system of the device to obtain the coordinates of the projection key points.
[0167] In some embodiments, determining the interaction point of the target object with the interactive smart tablet based on the location coordinates may include the following steps:
[0168] Based on the position coordinates and the preset wrist limit distance, the position coordinates of the starting point of the virtual ray are calculated;
[0169] The position coordinates of the interaction point are calculated based on the position coordinates of the starting point of the virtual ray and the hand position coordinates in the current coordinates.
[0170] The wrist limit distance is the maximum distance when the user moves one hand horizontally from the leftmost to the rightmost. This distance refers to the distance that the wrist joint (usually the center point of the wrist or a specific point of the wrist) moves from the leftmost position to the rightmost position in the horizontal direction. In the application, the wrist limit distance = A * the distance of the arms being spread out, A < 1, for example, A can be set to 0.7.
[0171] In some embodiments, calculating the position coordinates of the starting point of the virtual ray based on the position coordinates and the preset wrist limit distance may include:
[0172] Calculate a first quotient result of the limit distance of the interactive smart tablet and the width of the interactive smart tablet, and calculate the quotient of the z-axis coordinate in the position coordinate and the first quotient result to obtain the z-axis coordinate of the starting point of the virtual ray;
[0173] Calculate the second quotient of the x-axis coordinate and the z-axis coordinate in the position coordinate, and calculate the product of the second quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the x-axis coordinate of the starting point of the virtual ray;
[0174] Calculate the third quotient of the y-axis coordinate and the z-axis coordinate in the position coordinate, and calculate the product of the third quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the y-axis coordinate of the starting point of the virtual ray;
[0175] The three-dimensional coordinates composed of the x-axis coordinate, y-axis coordinate and z-axis coordinate of the starting point of the virtual ray are used as the position coordinates of the starting point of the virtual ray.
[0176] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the principle of calculating the starting point position coordinates of a virtual ray according to an embodiment of the present application. Figure 5 In the figure, lx1 is the wrist limit distance, ray_start is the starting point of the virtual ray, and w is the width of the interactive smart tablet. According to the definition of similar triangles, we have:
[0177] lx1 / w=z person / z ray_start ;
[0178] x ray_start / z ray_start= x person / z person ;
[0179] y ray_start / z ray_start= y person / z person ;
[0180] Among them, x person is the x-axis coordinate in the position coordinate, y person is the y-axis coordinate in the position coordinate, z person is the z-axis coordinate in the position coordinate. ray_start is the x-axis coordinate of the starting point of the virtual ray, y ray_start is the y-axis coordinate of the starting point of the virtual ray, z ray_start It is the z-axis coordinate of the starting point of the virtual ray.
[0181] It should be understood that when calculating the virtual ray starting point position coordinates according to the principle of similar triangles, the wrist limit distance and the width of the interactive smart tablet are used, so the virtual ray starting point position coordinates calculated by the similarity principle of triangles can follow the user's mobile control interaction point to reach any position on the interactive smart tablet on the width of the interactive smart tablet. Since the width of the interactive smart tablet is usually greater than the height of the interactive smart tablet, and the user's wrist activity range is usually a circle, the virtual ray starting point position coordinates calculated with reference to the width of the interactive smart tablet can also follow the user's mobile control interaction point to reach any position on the interactive smart tablet on the height of the interactive smart tablet.
[0182] In the related art, there is a scheme for controlling the movement of the interaction point based on the change in the position of the human hand relative to the human body, that is, when the position of the human hand relative to the human body changes, the interaction point on the interactive smart tablet will move based on the position change. The disadvantage of this scheme is that when the relative position of the person and the human hand does not change, but the person moves relative to the interactive smart tablet, the interaction point on the interactive smart tablet will not move, but for the user, the user expects the interaction point to change position as the user moves. Therefore, the user will think that the interactive control of the interactive smart tablet is not sensitive, which affects the user's product experience. In this embodiment, a virtual ray starting point is set, and the virtual ray starting point will change with the change of the user's position. Correspondingly, the change in the position of the virtual ray starting point drives the change in the position of the interaction point, so that the change in the user's position affects the change in the position of the interaction point, and improves the user's sensitivity to the interactive control of the interactive smart tablet.
[0183] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the principle of calculating the position coordinates of the interaction point shown in the embodiment of the present application. Figure 6 Medium hand is the hand position, P cast is the location of the interaction point.
[0184] z cast =z ray_start +k*(z hand -z ray_start )=0;
[0185] x cast =x ray_start +k*(x hand -x ray_start );
[0186] y cast =y ray_start +k*(y hand -y ray_start );
[0187] Among them, x hand ,y hand and z hand are the x-axis coordinates, y-axis coordinates, and z-axis coordinates of the hand position coordinates; similarly, x cast ,y cast and z cast are the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the location coordinate of the interaction point.
[0188] It should be understood that the hand position used in this embodiment is the coordinates of the hand whose gesture shows the intention to interact with the interactive smart tablet. If both hands show the intention to interact with the interactive smart tablet, the coordinates of the right hand are used by default as the hand position in this embodiment.
[0189] In the scheme of this embodiment, the current coordinates of the key points of the target object in the current image are identified, and based on the current coordinates, it is determined whether the target object has an intention to interact with the interactive smart tablet. If there is an intention to interact, the position coordinates of the target object are calculated based on the current coordinates, and finally the interaction point of the target object relative to the interactive smart tablet is determined based on the position coordinates. The scheme of the embodiment of the present application realizes the recognition of the interaction intention of the target object through the recognition of the key points of the human body, and finally realizes the interaction with the interactive smart tablet. Since the coordinates of the key points of the human body reflect the body movements of the target object, for the target object, it is realized to operate the focus of the interactive smart tablet through body movements, and the operation is natural, intuitive and convenient.
[0190] The present application also provides an interactive smart tablet control device, which is used to execute the interactive smart tablet control method provided in any of the above embodiments. Figure 7 As shown, the device comprises:
[0191] A detection module 71 is used to detect a human body region in a current image;
[0192] An identification module 72, used to identify the current coordinates of key points of the target object within the human body range;
[0193] A calculation module 73 is used to calculate, based on the current coordinates, a position coordinate representing a position of the target object relative to the interactive smart tablet when it is confirmed that the target object has an intention to interact with the interactive smart tablet based on the current coordinates;
[0194] The determination module 74 is used to determine the interaction point of the target object with respect to the interactive smart tablet based on the position coordinates.
[0195] The calculation module 73 is used for:
[0196] Based on the current coordinates, predict a first interaction intention parameter and a second interaction intention parameter, wherein the first interaction intention parameter is used to characterize whether the gesture of the target object has an interaction intention with the interactive smart tablet, and the second interaction intention parameter is used to characterize whether the human body posture of the target object has an interaction intention with the interactive smart tablet;
[0197] When the first interaction intention parameter and the second interaction intention parameter satisfy the preset interaction intention triggering condition, it is confirmed that the target object has the interaction intention with the interactive smart tablet.
[0198] The calculation module 73 is used for:
[0199] Based on the current coordinates, determine the hand range area of the target human body in the current image;
[0200] Recognize the image in the hand range area to obtain the gesture classification result;
[0201] The gesture classification result is matched with the gestures in a preset static human gesture library to obtain a first interaction intention parameter.
[0202] The calculation module 73 is used for:
[0203] Based on the hand range area in the historical image and the hand motion parameters obtained based on the historical image, predict the first hand range area in the current image; based on the current coordinates and the preset hand center distance, detect the second hand range area in the current image;
[0204] The first hand range area and the second hand range area are positionally fused to obtain a hand range area.
[0205] The calculation module 73 is used for:
[0206] Based on the current coordinates, determine the extension lines of the elbow and arm;
[0207] Based on the extension line and the preset hand center distance, a diameter circle is determined; the intersection of the diameter circle and the extension line is the position of the elbow, the radius of the diameter circle is the preset hand center distance, and the center of the diameter circle is located on the extension line and at the hand of the current image;
[0208] The area enclosed by the circumscribed rectangle of the diameter circle is taken as the second hand range area.
[0209] The calculation module 73 is used for:
[0210] Performing position fusion on the first hand range area and the second hand range area to obtain a fusion area;
[0211] The fusion area is expanded by a preset ratio to obtain the hand range area.
[0212] The calculation module 73 is used for:
[0213] When the intersection-and-union ratio of the first hand range area and the second hand range area is greater than the intersection-and-union ratio threshold, obtaining an overlapping area of the first hand range area and the second hand range area, and using the overlapping area as a fusion area;
[0214] When the IoU is less than or equal to the IoU threshold, the credibility of the current coordinates is obtained. When the credibility is greater than the credibility threshold, the first hand range area is used as the fusion area. When the credibility is less than or equal to the credibility threshold, the second hand range area is used as the fusion area.
[0215] The calculation module 73 is used for:
[0216] Based on the current coordinates, determine the arm posture classification result of the target object;
[0217] The arm posture classification result is matched with the posture in the preset static human posture library to obtain the second interaction intention parameter.
[0218] The calculation module 73 is used for:
[0219] When the posture of the preset human body parameter model is the i-th posture, the key points in the human body parameter model are projected into the device coordinate system to obtain the coordinates of the projected key points; the device corresponding to the device coordinate system is used to collect the current image;
[0220] For each key point of the target object, calculate the coordinate difference between the current coordinate of each key point and the projected key point coordinate of each key point;
[0221] Based on the coordinate differences, the mean of the coordinate differences is calculated;
[0222] When the error is greater than or equal to the error threshold, i=i+1 is updated until the error is less than the error threshold, and the coordinates of the projected key point when the error is less than the error threshold are used as the position coordinates.
[0223] The determination module 74 is used to:
[0224] Based on the position coordinates and the preset wrist limit distance, the position coordinates of the starting point of the virtual ray are calculated;
[0225] The position coordinates of the interaction point are calculated based on the position coordinates of the starting point of the virtual ray and the hand position coordinates in the current coordinates.
[0226] The determination module 74 is used to:
[0227] Calculate a first quotient result of the limit distance of the interactive smart tablet and the width of the interactive smart tablet, and calculate the quotient of the z-axis coordinate in the position coordinate and the first quotient result to obtain the z-axis coordinate of the starting point of the virtual ray;
[0228] Calculate the second quotient of the x-axis coordinate and the z-axis coordinate in the position coordinate, and calculate the product of the second quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the x-axis coordinate of the starting point of the virtual ray;
[0229] Calculate the third quotient of the y-axis coordinate and the z-axis coordinate in the position coordinate, and calculate the product of the third quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the y-axis coordinate of the starting point of the virtual ray;
[0230] The three-dimensional coordinates composed of the x-axis coordinate, y-axis coordinate and z-axis coordinate of the starting point of the virtual ray are used as the position coordinates of the starting point of the virtual ray.
[0231] The identification module 72 is used to:
[0232] Crop the current image to obtain a sub-image corresponding to the human body area;
[0233] Use the pre-trained human key point model to recognize the sub-image and predict the coordinates of the predicted key points of the target object;
[0234] Based on the coordinates of the human body range area, the coordinates of the predicted key points are translated and transformed to obtain the actual coordinates of the predicted key points;
[0235] The actual coordinates are filtered based on the human body range area to obtain the current coordinates.
[0236] The interactive smart tablet control device provided in the embodiment of the present application and the interactive smart tablet control method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.
[0237] The present application also provides an electronic device to execute the above interactive smart tablet control method. Figure 8 It shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 8 As shown, the electronic device 8 includes: a processor 800, a memory 801, a bus 802 and a communication interface 803, and the processor 800, the communication interface 803 and the memory 801 are connected via the bus 802; the memory 801 stores a computer program that can be run on the processor 800, and when the processor 800 runs the computer program, it executes the interactive smart tablet control method provided in any of the aforementioned embodiments of the present application.
[0238] The memory 801 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 803 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0239] The bus 802 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 801 is used to store programs, and the processor 800 executes the programs after receiving the execution instructions. The interactive smart tablet control method disclosed in any implementation of the embodiment of the present application may be applied to the processor 800, or implemented by the processor 800.
[0240] The processor 800 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 800. The above processor 800 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 801, and the processor 800 reads the information in the memory 801 and completes the steps of the above method in combination with its hardware.
[0241] The electronic device provided in the embodiment of the present application and the interactive smart tablet control method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.
[0242] The present application also provides a computer-readable storage medium corresponding to the interactive smart tablet control method provided in the above embodiment. Fig. 9 The computer-readable storage medium shown is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the interactive smart tablet control method provided in any of the aforementioned embodiments will be executed.
[0243] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0244] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the interactive smart tablet control method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0245] It should be noted that:
[0246] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0247] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be interpreted as reflecting the following schematic diagram: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the claims below, the inventive aspects are less than all the features of the single embodiment disclosed above. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.
[0248] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.
[0249] The above is only a preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. An interactive smart tablet control method, It is characterized in that include: Detect the human body area in the current image; Identify current coordinates of key points of a target object within the human body range area; In a case where it is confirmed based on the current coordinates that the target object has an intention to interact with the interactive smart tablet, based on the current coordinates, calculating a position coordinate for representing a position of the target object relative to the interactive smart tablet; Based on the position coordinates, an interaction point of the target object with respect to the interactive smart tablet is determined.
2. The method according to claim 1, It is characterized in that Confirming that the target object has an intention to interact with the interactive smart tablet based on the current coordinates includes: Based on the current coordinates, predict a first interaction intention parameter and a second interaction intention parameter, wherein the first interaction intention parameter is used to characterize whether the gesture of the target object has an interaction intention with the interactive smart tablet, and the second interaction intention parameter is used to characterize whether the human body posture of the target object has an interaction intention with the interactive smart tablet; When the first interaction intention parameter and the second interaction intention parameter satisfy a preset interaction intention trigger condition, it is confirmed that the target object has an interaction intention with the interactive smart tablet.
3. The method according to claim 2, It is characterized in that Predicting a first interaction intention parameter based on the current coordinates includes: Based on the current coordinates, determining a hand range area of the target human body in the current image; Recognize the image within the hand range to obtain a gesture classification result; The gesture classification result is matched with gestures in a preset static human gesture library to obtain the first interaction intention parameter.
4. The method according to claim 3, It is characterized in that Determining a hand range area of the target human body in the current image based on the current coordinates includes: Based on the hand range area in the historical image and the hand motion parameters obtained based on the historical image, predict the first hand range area in the current image; based on the current coordinates and the preset hand center distance, detect the second hand range area in the current image; Position fusion is performed on the first hand range area and the second hand range area to obtain the hand range area.
5. The method according to claim 4, It is characterized in that Detecting a second hand range area in the current image based on the current coordinates and a preset hand center distance includes: Based on the current coordinates, determining the extension line of the elbow and the arm; Based on the extension line and the preset hand center distance, a diameter circle is determined; the intersection of the diameter circle and the extension line is the position of the elbow, the radius of the diameter circle is the preset hand center distance, and the center of the diameter circle is located on the extension line and at the hand of the current image; The area enclosed by the circumscribed rectangle of the diameter circle is used as the second hand range area.
6. The method according to claim 4, It is characterized in that Performing position fusion on the first hand range area and the second hand range area to obtain the hand range area includes: Performing position fusion on the first hand range area and the second hand range area to obtain a fused area; The fusion area is expanded outward by a preset ratio to obtain the hand range area.
7. The method according to claim 6, It is characterized in that Performing position fusion on the first hand range area and the second hand range area to obtain a fusion area includes: When the intersection-and-union ratio of the first hand range area and the second hand range area is greater than the intersection-and-union ratio threshold, obtaining an overlapping area of the first hand range area and the second hand range area, and using the overlapping area as the fusion area; When the IoU is less than or equal to the IoU threshold, the credibility of the current coordinates is obtained; when the credibility is greater than the credibility threshold, the first hand range area is used as the fusion area; when the credibility is less than or equal to the credibility threshold, the second hand range area is used as the fusion area.
8. The method according to claim 2, It is characterized in that Predicting a second interaction intention parameter based on the current coordinates includes: Based on the current coordinates, determining a classification result of the arm posture of the target object; The arm posture classification result is matched with the postures in a preset static human posture library to obtain the second interaction intention parameter.
9. The method according to claim 1, It is characterized in that Based on the current coordinates, calculating the position coordinates used to characterize the position of the target object relative to the interactive smart tablet includes: When the posture of the preset human body parameter model is the i-th posture, projecting the key points in the human body parameter model into the device coordinate system to obtain the coordinates of the projected key points; the device corresponding to the device coordinate system is used to collect the current image; For each key point of the target object, calculating the coordinate difference between the current coordinate of each key point and the projected key point coordinate of each key point; Based on the coordinate differences, calculating a coordinate difference mean; When the error is greater than or equal to the error threshold, i=i+1 is updated until the error is less than the error threshold, and the coordinates of the projected key point when the error is less than the error threshold are used as the position coordinates.
10. The method according to claim 1, It is characterized in that Determining the interaction point of the target object with respect to the interactive smart tablet based on the position coordinates includes: Calculating the position coordinates of the starting point of the virtual ray based on the position coordinates and the preset wrist limit distance; The position coordinates of the interaction point are calculated based on the position coordinates of the starting point of the virtual ray and the position coordinates of the hand in the current coordinates.
11. The method according to claim 10, It is characterized in that Based on the position coordinates and the preset wrist limit distance, the position coordinates of the starting point of the virtual ray are calculated, including: Calculating a first quotient of the limit distance of the interactive smart tablet and the width of the interactive smart tablet, and calculating the quotient of the z-axis coordinate in the position coordinate and the first quotient to obtain the z-axis coordinate of the starting point of the virtual ray; Calculating a second quotient of the x-axis coordinate and the z-axis coordinate in the position coordinate, and calculating the product of the second quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the x-axis coordinate of the starting point of the virtual ray; Calculating a third quotient of the y-axis coordinate and the z-axis coordinate in the position coordinate, and calculating the product of the third quotient and the z-axis coordinate of the starting point of the virtual ray to obtain the y-axis coordinate of the starting point of the virtual ray; The three-dimensional coordinates composed of the x-axis coordinate, y-axis coordinate and z-axis coordinate of the starting point of the virtual ray are used as the position coordinates of the starting point of the virtual ray.
12. The method according to claim 1, It is characterized in that Identifying the current coordinates of key points of the target object within the human body range, including: Cropping the current image to obtain a sub-image corresponding to the human body range area; Using a pre-trained human key point model to identify the sub-image, and predicting the coordinates of the predicted key points of the target object; Based on the coordinates of the human body range area, the coordinates of the predicted key points are translated and transformed to obtain the actual coordinates of the predicted key points; The actual coordinates are screened based on the human body range area to obtain the current coordinates.
13. An interactive intelligent tablet control device, It is characterized in that include: A detection module is used to detect the human body area in the current image; An identification module, used for identifying the current coordinates of key points of a target object within the human body range; A calculation module, configured to calculate, based on the current coordinates, position coordinates for representing the position of the target object relative to the interactive smart tablet when it is confirmed that the target object has an intention to interact with the interactive smart tablet based on the current coordinates; A determination module is used to determine the interaction point of the target object with respect to the interactive smart tablet based on the position coordinates.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that The processor runs the computer program to implement the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, It is characterized in that The program is executed by a processor to implement the method according to any one of claims 1 to 12.