Interactive question answering method, device, storage medium and learning machine

By integrating posture recognition and clustering network technology on the learning machine and combining the body-sensing answerer, the intelligent mapping of user posture to answering options is realized, solving the problem of low intelligence in the existing learning machine and improving user interactivity and learning fun.

CN114721505BActive Publication Date: 2025-06-24BEIJING YUDA ORIENTAL SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210178866.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-06-24
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing learning machines are less intelligent when users answer questions, and lack intelligent interactive answering technology based on artificial intelligence, resulting in insufficient user interaction and learning fun.

Method used

The target image is collected through the camera, the attitude recognition model is used to identify the user's bone key points, and these key points are analyzed through the clustering network to determine the answer options represented by the user's attitude. Optionally, use the somatosensor to correct the coordinate error in posture recognition.

Benefits of technology

It improves the intelligence of the learning machine when answering questions, increases the interaction between users and the learning machine and the fun of answering questions, and provides an alternative way to answer questions when users are not convenient to click on the screen.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114721505B_ABST
    Figure CN114721505B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an interactive answering method, apparatus, storage medium and learning machine. The method includes: obtaining a target image collected by a camera; performing pose recognition on the target image through a pose recognition model to obtain a plurality of skeletal key points; performing clustering analysis on the plurality of skeletal key points through a clustering network to determine the answering option represented by the pose of the user in the target image in the image. Through this technical solution, when the user uses the learning machine to answer questions, the learning machine can automatically recognize the user's current pose and determine the answering option selected by the user according to the user's current pose, thereby improving the intelligence level of the learning machine in answering questions, increasing the interactivity between the user and the learning machine and the fun of answering questions. At the same time, it can also provide another optional answering method for the user when the user is not convenient to click on the answering option on the learning machine screen, such as when the hands are stained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of learning machines, and in particular, to an interactive answering method, device, storage medium, and learning machine. Background Art

[0002] A learning machine is an electronic teaching product that can assist students in learning. Currently, the learning machines on the market usually only have simple learning functions, such as playing courseware and playing the teacher's narration of the courseware. The product form is single. Especially when users need to use the learning machine to answer questions, they can only have simple multiple-choice question interactions with the learning machine through the buttons on the screen. Learning machines for children need more artificial intelligence-based intelligent interactive answering technologies to increase the overall learning fun and interactive effects, but the current learning machines have a low degree of intelligence. Summary of the Invention

[0003] The purpose of the present disclosure is to provide an interactive answering method, device, storage medium, and learning machine to improve the intelligence level of the learning machine for user answering questions.

[0004] In a first aspect, an embodiment of the present disclosure provides an interactive answering method, including:

[0005] Obtain a target image collected by a camera;

[0006] Perform pose recognition on the target image through a pose recognition model to obtain multiple skeletal key points;

[0007] Perform clustering analysis on the multiple skeletal key points through a clustering network to determine the answering option represented by the pose of the user in the target image in the image.

[0008] Optionally, the performing clustering analysis on the multiple skeletal key points through a clustering network to determine the answering option represented by the pose of the user in the target image in the image includes:

[0009] Select one key point from the multiple skeletal key points as a reference point;

[0010] Calculate the deviation value between each skeletal key point and the reference point respectively, and obtain a deviation vector according to all the deviation values;

[0011] Input the deviation vector into the clustering network to obtain the cluster corresponding to the deviation vector;

[0012] Determine the corresponding answering option according to the label of the cluster.

[0013] Optionally, the obtaining a deviation vector according to all the deviation values includes:

[0014] If the length of the vector composed of all the deviation values does not meet the input length requirement of the clustering network, zeros are padded at the end of the vector until the length of the vector meets the input length requirement, obtaining the deviation vector.

[0015] Optionally, in the deviation vector, the values of 0 among all the deviation values are located after the non-zero values.

[0016] Optionally, after obtaining multiple skeletal key points, the method further includes:

[0017] Obtaining a first coordinate measured by a somatosensory answer device at the current moment, where the somatosensory answer device is arranged on a body part of the user and is used to measure the coordinate of the part;

[0018] Obtaining a second coordinate of a target skeletal key point corresponding to the part measured by the somatosensory answer device among the multiple skeletal key points at the current moment, where the origin of the first coordinate is the same as that of the second coordinate;

[0019] Correcting the second coordinate according to the first coordinate to obtain a corrected coordinate;

[0020] Replacing the second coordinate of the target skeletal key point among the multiple skeletal key points with the corrected coordinate.

[0021] Optionally, the somatosensory answer device is a wearable device arranged on the user's hand or a device worn on the user's finger.

[0022] Optionally, the correcting the second coordinate according to the first coordinate to obtain a corrected coordinate includes:

[0023] Determining the two-dimensional coordinate corresponding to the first coordinate;

[0024] Calculating a first proportional relationship between the distance of movement of the part and the image pixels according to the two-dimensional coordinate and the second coordinate at the current moment, and the two-dimensional coordinate and the second coordinate at the previous moment, where the previous moment is the moment when multiple skeletal key points were obtained by performing pose recognition through a pose recognition model last time;

[0025] Calculating a target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image according to the first proportional relationship;

[0026] Correcting the second coordinate according to the target coordinate to obtain the corrected coordinate.

[0027] Optionally, the first proportional relationship is the ratio of the first Euclidean distance to the second Euclidean distance, where the first Euclidean distance is the Euclidean distance between the second coordinates at the current moment and the second coordinates at the previous moment, and the second Euclidean distance is the Euclidean distance between the two-dimensional coordinates at the current moment and the two-dimensional coordinates at the previous moment.

[0028] Optionally, calculating the target coordinates corresponding to the two-dimensional coordinates at the current moment on the target image according to the first proportional relationship includes:

[0029] Calculating the average value of the first proportional relationship according to the first proportional relationship corresponding to each moment within a preset time range at the current moment and before, to obtain a second proportional relationship; where each moment is a moment when pose recognition is performed once through a pose recognition model to obtain multiple skeletal key points.

[0030] Calculating the target coordinates corresponding to the two-dimensional coordinates at the current moment on the target image according to the second proportional relationship.

[0031] Optionally, calculating the average value of the first proportional relationship according to the first proportional relationship corresponding to each moment within a preset time range at the current moment and before, to obtain a second proportional relationship, includes:

[0032] Calculating the weighted average value of the first proportional relationship according to the first proportional relationship corresponding to each moment within a preset time range at the current moment and before, to obtain the second proportional relationship; where the first proportional relationship at each moment corresponds to a weighting coefficient, the weighting coefficient is greater than 0 and less than 1, and the weighting coefficient is negatively correlated with the ratio of the first proportional relationship at the current moment to the first proportional relationship at the previous moment.

[0033] Optionally, the weighting coefficient is determined by the following formula:

[0034]

[0035] where α t is the weighting coefficient corresponding to the moment, is the first proportional relationship at this corresponding moment, is the first proportional relationship at the previous moment of this corresponding moment, 0 < m < 1.

[0036] Optionally, performing pose recognition on the target image through a pose recognition model to obtain multiple skeletal key points includes:

[0037] Detecting the target image through a target detection model to determine whether the area containing the user belongs to the whole body or the hand when the user is included in the target image.

[0038] If the area containing the user belongs to the whole body, the whole body pose of the target image is recognized through a human body pose recognition model to obtain multiple skeletal key points of the whole body;

[0039] If the area containing the user belongs to the hand, the hand pose of the target image is recognized through a hand pose recognition model to obtain multiple skeletal key points of the hand.

[0040] Optionally, obtaining the target image collected by the camera includes:

[0041] Writing the video collected by the camera into a queue to form a video frame image queue;

[0042] Taking out the images in the video frame image queue in sequence and transmitting them to the display unit for display;

[0043] Determining the image at the top of the video frame image queue at the current moment as the target image.

[0044] Optionally, the method further includes:

[0045] Receiving the identification information sent by the content trigger device;

[0046] According to the identification information, unlocking the preset content corresponding to the identification information, where the preset content includes virtual characters, virtual clothes or plots.

[0047] Optionally, the unlocking the preset content corresponding to the identification information according to the identification information includes:

[0048] Obtaining the device ID corresponding to the content trigger device;

[0049] Sending a verification request to the server, where the device ID is carried in the verification request;

[0050] Receiving the verification result returned by the server, where the verification result is generated by the server according to the unlocking times corresponding to the device ID and a preset threshold. Among them, when the unlocking times are greater than the preset threshold, the verification result is not passed, otherwise the verification result is passed;

[0051] If the verification result is passed, unlocking the preset content corresponding to the identification information;

[0052] After the unlocking is completed, sending an unlocking notification to the server so that the server increments the unlocking times corresponding to the device ID by 1.

[0053] In a second aspect, an interactive answering device provided by an embodiment of the present disclosure includes:

[0054] An image acquisition module, configured to acquire a target image captured by a camera;

[0055] A pose recognition module, configured to perform pose recognition on the target image through a pose recognition model to obtain a plurality of skeletal key points;

[0056] An answer option determination module, configured to perform clustering analysis on the plurality of skeletal key points through a clustering network to determine the answer option represented by the pose of the user in the target image in the image.

[0057] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0058] In a fourth aspect, an embodiment of the present disclosure provides a learning machine, including:

[0059] A camera, configured to capture an image;

[0060] A memory, on which a computer program is stored;

[0061] A processor, configured to execute the computer program in the memory to implement the method described in the first aspect.

[0062] Optionally, the learning machine further includes a somatosensory answer device, wherein the somatosensory answer device includes a gyroscope and a communication module, the somatosensory answer device is configured to be disposed on a body part of a user, the gyroscope is configured to measure a first coordinate of the part, and the communication module is configured to send the first coordinate measured by the gyroscope to the processor.

[0063] Optionally, the somatosensory answer device is a wearable device disposed on the user's hand or a device worn on the user's finger.

[0064] Through the above technical solutions, when the user uses the learning machine to answer questions, the learning machine can automatically recognize the user's current pose and determine the answer option selected by the user according to the user's current pose. For example, when the user answers a question and extends one finger, it can be determined that the answer option selected by the user is option 1 or option A, thereby improving the intelligence of the learning machine in answering questions, increasing the interactivity between the user and the learning machine and the fun of answering questions. At the same time, it can also provide another optional answering method for the user when the user is not convenient to click the answer option on the learning machine screen, such as when the hands are stained.

[0065] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0066] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0067] Figure 1 Flowchart of the interactive answering method provided by an embodiment of the present disclosure;

[0068] Figure 2 It is a specific flowchart of step S130 in an embodiment of the present disclosure;

[0069] Figure 3 It is a flowchart for correcting the skeletal key points obtained in step S120 according to a somatosensory answering device provided by an embodiment of the present disclosure;

[0070] Figure 4 It is a specific flowchart of step S230 in an embodiment of the present disclosure;

[0071] Figure 5 It is another flowchart of the interactive answering method provided by an embodiment of the present disclosure;

[0072] Figure 6 It is a specific flowchart of step S320 in an embodiment of the present disclosure;

[0073] Figure 7 It is a block diagram of an interactive answering device provided by an embodiment of the present disclosure;

[0074] Figure 8 It is a block diagram of a learning machine provided by an embodiment of the present disclosure. Detailed Description of the Specific Embodiment

[0075] The following provides a detailed description of the specific embodiments of the present disclosure in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure.

[0076] An embodiment of the present disclosure provides an interactive answering method applied to a learning machine. Through this interactive answering method, when a user uses the learning machine to answer questions, the learning machine can automatically recognize the user's current posture and determine the answering option selected by the user according to the user's current posture. For example, when the user answers a question and extends one finger, it is determined that the answering option selected by the user is option 1 or option A; when the user extends two fingers, it is determined that the answering option selected by the user is option 2 or option B; or when the user bends the body into the shape of "C", it is determined that the answering option selected by the user is option C. Thereby, the intelligence level of the learning machine for answering questions is improved, the interactivity between the user and the learning machine and the interestingness of answering questions are increased. At the same time, when it is inconvenient for the user to click on the answering options on the learning machine screen, such as when the hands are stained, another optional answering method can be provided for the user.

[0077] Figure 1 is a flowchart of the interactive answering method provided by an embodiment of the present disclosure. Please refer to Figure 1 , and the method includes the following steps:

[0078] S110, obtain a target image collected by a camera.

[0079] A camera is provided on the learning machine, and the learning machine can collect an image in front of the camera through the camera. The camera can collect an image at regular intervals as the target image, or can collect a video in real time and transmit it to the processor of the learning machine, and the processor processes the video to obtain the target image.

[0080] In an alternative embodiment, the specific process of obtaining the target image in step S110 includes: obtaining the video collected by the camera, writing the video collected by the camera into a queue to form a video frame image queue, sequentially taking out the images in the video frame image queue and transmitting them to the display unit for display, and determining the image at the top of the video frame image queue at the current moment as the target image. In this way, when performing interactive answering, the learning machine can not only display the video collected by the camera on the screen in real time, but also use only the image at the top of the video frame image queue at the current moment as the target image for subsequent pose recognition and clustering analysis processing, rather than processing each frame of the video frame image queue, thereby effectively reducing the serious power consumption caused by image processing. In particular, the longer the subsequent processing process takes, the more frames of images will be separated between two adjacent target images determined from the video frame image queue. Since the pose recognition and clustering analysis processing of these multiple frames of images are omitted, the power consumption can be greatly saved.

[0081] S120, perform pose recognition on the target image through a pose recognition model to obtain multiple skeletal key points.

[0082] Among them, when the user performs interactive answering, the user can answer the question by posing the whole body or hands in the shape of a specific option, that is, the user may answer the question by using the whole body pose or by using gestures.

[0083] First, the target image is detected by a target detection model to determine whether the user is included in the target image. If the user is not included in the target image, the target image is discarded, and then the image at the top of the video frame image queue is selected as the new target image. If the user is included in the target image, the target detection model can also determine whether the area containing the user in the target image belongs to the whole body or the hand. If the area containing the user belongs to the whole body, the whole body pose of the target image is recognized by a human pose recognition model (such as the OpenPose network, etc.) to obtain multiple skeletal key points of the whole body. Among them, the output of the human pose recognition model includes the coordinates of each skeletal key point of the whole body. If the area containing the user belongs to the hand, the hand pose of the target image is recognized by a hand pose recognition model (such as the MediaPipe model, etc.) to obtain multiple skeletal key points of the hand. Among them, the output of the hand pose recognition model includes the coordinates of each skeletal key point of the hand.

[0084] In a specific embodiment, the YOLO network can be used to detect whether the area containing the user in the target image belongs to the whole body or the hand. If the area containing the user belongs to the whole body, the OpenPose network can be used to perform pose recognition on the target image to obtain the skeletal key points of multiple parts of the user's whole body in the target image. The multiple parts include, for example, but are not limited to, the head, left shoulder, left elbow, left hand, right shoulder, right elbow, right hand, left knee, right knee, etc. If the area containing the user belongs to the hand, the MediaPipe model can be used to obtain the multiple skeletal key points of the hand, including the coordinates of each finger joint of the hand.

[0085] S130. Cluster analysis is performed on the multiple skeletal key points through a clustering network to determine the answer option represented by the user's pose in the target image.

[0086] After performing pose recognition on the target image to obtain multiple skeletal key points, cluster analysis is performed on the multiple skeletal key points through a clustering network, so as to determine the answer option represented by the user's pose in the target image. Among them, the clustering network can adopt a DBSCAN network or other networks for clustering.

[0087] Optionally, Figure 2 A specific flowchart showing the determination of the answer option represented by the user's pose in the target image in step S130 is shown. Please refer to Figure 2 , step S130 includes:

[0088] S131. Select one key point from the multiple skeletal key points as the reference point.

[0089] S132. Calculate the deviation values between each skeletal key point and the reference point respectively, and obtain a deviation vector based on all the deviation values.

[0090] Among them, assume that n skeletal key points are obtained. In the set A with n skeletal key points, A = {a1, a2, …, a n}, randomly select a key point a i ∈ A as the reference point, and calculate the coordinate deviations of all key points in the set A from a i to obtain a vector P:

[0091] P = [a1 - a i , a2 - a i , …, a n - a i .

[0092] Among them, if the length of the vector composed of all the calculated deviation values does not meet the input length requirement of the clustering network, then fill zeros at the end of the vector until the length of the vector after filling zeros meets the input length requirement of the clustering network, so as to obtain a deviation vector.

[0093] Among them, in the deviation vector, place 0 at the end of the vector. The purpose of this operation is to uniformly place meaningless information at the end so that it does not affect subsequent clustering operations.

[0094] It should be noted that the purpose of calculating the deviation vector is to normalize each skeletal key point. In the target image, the position where a person appears is random, and the distribution of the coordinates of each skeletal key point is also random. For example, the coordinates of the head key point may be (100, 100), or may be (10, 60). If the coordinates (absolute positions) of each skeletal key point are directly input into the clustering network, effective results may not be obtained. Therefore, select a skeletal key point as the reference point, subtract all skeletal key points from the reference point to obtain a deviation vector. In the mapping process from skeletal key points to answer options, cluster according to the relative position relationship between each skeletal key point and the reference point in the deviation vector, rather than clustering based on the absolute position of each skeletal key point, so as to solve the influence on clustering operations caused by the random position of the user in the image.

[0095] In addition, considering that when performing pose recognition on a target image through a pose recognition network, it cannot be guaranteed that all key points can be detected. If fixed skeletal key points are used as reference points, once a key point is missing and this key point happens to be the reference point, it will be very difficult to make subsequent judgments. For example, if the key points of the hand are fixedly selected as the reference points, then when the user's hand is outside the target image or the hand is blocked, resulting in the failure to recognize the skeletal key points of the hand, the subsequent steps cannot be continued. To increase the robustness of this solution, in step S131, a key point is randomly selected from the multiple skeletal key points as the reference point.

[0096] S133. Input the deviation vector into the clustering network to obtain the cluster corresponding to the deviation vector.

[0097] S134. Determine the corresponding answer option according to the label of the cluster.

[0098] Among them, each cluster will be pre-given a corresponding answer option label, and each label represents option A, option B, option C, option D respectively, or option 1, option 2, option 3, option 4, etc. After inputting the deviation vector into the clustering network, the clustering network outputs the cluster corresponding to the deviation vector, and then looks up the answer option label corresponding to the cluster to obtain the answer option represented by the pose of the user in the target image.

[0099] Furthermore, in the foregoing steps, the coordinates of the skeletal key points are obtained through pure vision recognition means, and there is often a certain error between the coordinates obtained by this pure vision recognition means and the real coordinates, and its accuracy is difficult to grasp, especially for some parts of the user's body that are easily blocked, and the coordinates of the corresponding skeletal key points are more likely to have errors. For example, in full-body pose recognition, since the hand itself has a small target and is the most easily blocked part, the coordinates of the hand are more likely to have errors among the multiple skeletal key points obtained. Similarly, in gesture recognition, the fingers far from the camera are easily blocked by the fingers close to the camera, so the coordinates of the finger joints far from the camera are more likely to have errors among the multiple skeletal key points obtained.

[0100] To solve the problem that the coordinates of the skeletal key points obtained by pure vision recognition means have errors, the embodiments of the present disclosure provide a somatosensory answer device. The somatosensory answer device is arranged on a body part of the user and is used to measure the coordinates of this part. The coordinates of the error-prone positions or some important positions in the pose recognition process can be corrected through the somatosensory answer device. For example, the somatosensory answer device is a wearable device arranged on the user's hand or a device worn on the user's finger. In the full-body pose recognition scenario, the somatosensory answer device may be a smart watch worn on the user's hand. In the gesture recognition scenario, the somatosensory answer device may be a ring worn on the user's finger.

[0101] In an alternative embodiment, the somatosensory answer device includes a gyroscope and a communication module. The gyroscope is used to collect the spatial coordinate information of the wearing part (such as the hand or a certain finger) of the somatosensory answer device. The communication module is used to transmit the spatial coordinate information collected by the gyroscope to the processor of the learning machine. The processor uses the spatial coordinate information collected by the somatosensory answer device to correct the coordinates of the corresponding key points among the multiple skeletal key points obtained by the pose recognition model. Among them, the communication module can be a Bluetooth module, a WIFI module, etc. Optionally, the somatosensory answer device may further include buttons, a recording module, a vibration module, an LED module, etc. Therefore, the somatosensory answer device has functions such as coordinate collection, coordinate transmission, button pressing, recording, vibration, and light-emitting reminder.

[0102] Figure 3 A flowchart showing the correction of a specific key point among the multiple skeletal key points obtained in the above step S120 according to the spatial coordinates measured by the somatosensory answer device is shown. The specific key point refers to the target skeletal key point corresponding to the measured part of the somatosensory answer device. Assume that the somatosensory answer device is a smart watch set on the user's hand, then the target skeletal key point among the multiple skeletal key points is the hand key point. Assume that the somatosensory answer device is a ring set at the root of the index finger, then the target skeletal key point among the multiple skeletal key points is the key point of the index finger root joint.

[0103] Please refer to Figure 3 , after obtaining the multiple skeletal key points, the coordinates of the specific key point are corrected through the following steps:

[0104] S210, obtain the first coordinate measured by the somatosensory answer device at the current moment.

[0105] S220, obtain the second coordinate of the target skeletal key point corresponding to the measured part of the somatosensory answer device among the multiple skeletal key points at the current moment.

[0106] Among them, the somatosensory answer device is used to collect the spatial position coordinates of the measured part, obtain the first coordinate, and transmit the first coordinate to the learning machine through a communication module such as Bluetooth.

[0107] Among them, the coordinate origin of the coordinate system adopted by the somatosensory answer device coincides with the coordinate origin of the coordinate system adopted by the camera. Therefore, the origin of the first coordinate measured by the somatosensory answer device at the current moment is the same as the origin of the second coordinate of the target skeletal key point at the current moment.

[0108] S230, correct the second coordinate according to the first coordinate to obtain the corrected coordinate.

[0109] S240, replace the second coordinate of the target skeletal key point among the multiple skeletal key points with the corrected coordinate.

[0110] Among them, the second coordinate of the target bone key point among multiple bone key points is corrected by the first coordinate measured by the body-sensing answer device at the current moment to obtain a corrected coordinate, and then the second coordinate corresponding to the target bone key point is replaced with this corrected coordinate. After that, when performing clustering analysis on multiple bone key points through the clustering network, the coordinate of the target bone key point is the corrected coordinate after being corrected by the body-sensing answer device. Therefore, the above steps solve the problem that the clustering analysis result is inaccurate due to the coordinate error of the bone key points output by the pose recognition model, and improve the correct rate of the answer options obtained by the clustering analysis.

[0111] Optionally, Figure 4 Fig. shows a specific flowchart of correcting the second coordinate according to the first coordinate in step S230 to obtain a corrected coordinate. Please refer to Figure 4 , step S230 includes:

[0112] S231, determining the two-dimensional coordinate corresponding to the first coordinate.

[0113] Among them, the first coordinate measured by the body-sensing answer device is a spatial position coordinate. Assuming that at time t, the first coordinate is The second coordinate of the target bone key point among multiple bone key points is a two-dimensional coordinate on the image plane of the camera imaging, which is (x t , y t ). Converting the first coordinate measured by the body-sensing answer device into a two-dimensional coordinate, specifically, removing the Z-axis information in the first coordinate and retaining the XY-axis information, the obtained two-dimensional coordinate is

[0114] S232, calculating a first proportional relationship between the moving distance of the part measured by the body-sensing answer device and the image pixels according to the two-dimensional coordinate and the second coordinate at the current moment, and the two-dimensional coordinate and the second coordinate at the previous moment.

[0115] Among them, the previous moment is the moment when the pose recognition model was used for pose recognition last time to obtain multiple bone key points. According to the two-dimensional coordinate at time t and the second coordinate (x t , y t ) of the target bone key point, and the two-dimensional coordinate at time t-1 and the second coordinate (x t-1 , y t-1 ), calculating the proportional relationship between the moving distance of the part measured by the body-sensing answer device and the image pixels to obtain the first proportional relationship R 1 at time t.

[0116] Among them, the first proportional relationship is the ratio of the first Euclidean distance to the second Euclidean distance. The first Euclidean distance is the Euclidean distance between the second coordinate at the current moment and the second coordinate at the previous moment, and the second Euclidean distance is the Euclidean distance between the two-dimensional coordinate at the current moment and the two-dimensional coordinate at the previous moment. The specific formula can be expressed as:

[0117]

[0118] Among them, ||.||2 represents the two-norm, that is, the Euclidean distance. The numerator is the first Euclidean distance between the second coordinate (x t , y t ) at time t and the second coordinate (x t-1 , y t-1 ) at time t-1, and the denominator is the second Euclidean distance between the two-dimensional coordinate at time t and the two-dimensional coordinate at time t-1 .

[0119] S233. According to this first proportional relationship, calculate the target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image.

[0120] Among them, after calculating the first proportional relationship R 1 , according to the first proportional relationship R 1 , calculate the target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image.

[0121] Optionally, the two-dimensional coordinate at the current moment can be directly multiplied by the first proportional relationship R 1 to obtain the target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image.

[0122] Optionally, in order to avoid the influence of the pose recognition model on the proportional relationship, continuously calculate the first proportional relationship R 1 at each moment within a certain time range, and then calculate its average value to obtain the second proportional relationship R. Then multiply the two-dimensional coordinate at the current moment by the second proportional relationship R to obtain the target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image. Among them, each moment is the moment when a pose recognition is performed through the pose recognition model to obtain multiple skeletal key points. The average value may be the arithmetic mean or the weighted mean.

[0123] In an alternative embodiment, according to the first proportional relationship R 1 corresponding to each moment within the preset time range at the current moment and before, calculate the first proportional relationship R 1The weighted average value is obtained to get the second proportional relationship R. Among them, each moment corresponds to a weighting coefficient, the weighting coefficient is greater than 0 and less than 1, and the weighting coefficient is negatively correlated with the ratio of the first proportional relationship R at the current moment to the previous moment. 1 The ratio is negatively correlated.

[0124] Exemplarily, the specific formula for the second proportional relationship R at the current moment can be expressed as:

[0125]

[0126]

[0127] Among them, N represents a total of N moments within a preset time range before and including the current moment, and α t is the weighting coefficient corresponding to the t-th moment among these N moments, is the first proportional relationship corresponding to the t-th moment, is the first proportional relationship corresponding to the moment before the t-th moment, 0 < m < 1.

[0128] Among them, 0 < m < 1, so is a function that decreases with the ratio of the first proportional relationship at the current moment to the previous moment, that is, a decreasing function.

[0129] It can be understood that when the first proportional relationship R at a certain moment 1 has a large gap with the first proportional relationship R at its previous moment 1 , it means that there is a great risk of error in the first proportional relationship R at this moment. According to the formula of the second proportional relationship R, it can be seen that a weighting coefficient α 1 is set before the first proportional relationship R. 1 Once the first proportional relationship R at a certain moment t has a large gap with the first proportional relationship R at its previous moment 1 , then the weighting coefficient α 1 at this moment will be very small, thereby reducing the impact of the error of the first proportional relationship at individual moments on the second proportional relationship. t

[0130] After calculating the second proportional relationship R, multiply the two-dimensional coordinates at the current moment by the second proportional relationship R to obtain the target coordinates corresponding to the two-dimensional coordinates at the current moment on the target image.

[0131] S234. Correct the second coordinate according to the target coordinates to obtain the corrected coordinates.

[0132] Among them, after obtaining the target coordinates, calculate the average value of the target coordinates and the second coordinates as the corrected coordinates. The average value of the X value of the target coordinates and the X value of the second coordinates will be used as the X value of the corrected coordinates, and the average value of the Y value of the target coordinates and the Y value of the second coordinates will be used as the Y value of the corrected coordinates. Thus, this solution solves the problem of coordinate errors in the skeletal key points output by the pose recognition model.

[0133] Furthermore, the embodiment of the present disclosure is also provided with a content trigger device, and the specific manifestation form of the content trigger device can be a card or a badge, etc. The content trigger device includes a storage module for storing identification information and a communication module for sending the identification information to the learning machine. The communication module can be an NFC (Near Field Communication) module, a WIFI module, a Bluetooth module, etc.

[0134] Figure 5 Another flowchart of the interactive answering method provided by the embodiment of the present disclosure is shown. Please refer to Figure 5 and the method further includes:

[0135] S310, receiving the identification information sent by the content trigger device.

[0136] S320, according to the identification information, unlocking the preset content corresponding to the identification information.

[0137] Among them, each identification information corresponds to a preset content, and the preset content includes virtual characters, virtual clothes or plots. Exemplarily, the content trigger device is provided with an NFC module. When the NFC module of the content trigger device approaches the NFC module of the learning machine, both the NFC module of the content trigger device and the NFC module of the learning machine are activated. The NFC module of the content trigger device sends the stored identification information to the learning machine, and the NFC module in the learning machine receives the identification information sent by the content trigger device and transmits it to the processor, and the processor performs the next processing according to the identification information.

[0138] Among them, after receiving the identification information, the learning machine unlocks the preset content corresponding to the identification information according to the identification information. After unlocking, the corresponding virtual characters, virtual clothes or plots can be displayed on the screen of the learning machine.

[0139] For example, when the user learns about the chapter related to "Li Bai", the user can choose to purchase the corresponding content trigger device. The content trigger device stores the identification information for introducing the plot of Li Bai's life experience and the identification information of Li Bai's virtual character image. When the user holds the content trigger device close to the learning machine, the learning machine obtains the identification information from the content trigger device, thereby unlocking the corresponding plot and virtual character image. Specifically, the virtual character of Li Bai can be displayed on the screen of the learning machine, and the corresponding character description can be provided to introduce Li Bai's relevant life experience.

[0140] For another example, the content trigger device stores the identification information of a certain virtual character. When the user holds the content trigger device close to the learning machine, the learning machine obtains the identification information from the content trigger device, thereby unlocking the corresponding virtual character and displaying the virtual character on the screen of the learning machine. The virtual character can be bound to the courseware. When the user studies the content of the courseware, the learning machine can interact with the user through the virtual character. For example, the virtual character explains the knowledge points of the courseware to the user, reads the questions aloud for the user in voice, and the virtual character makes corresponding expressions and encourages the user to continue answering questions when the user answers correctly or incorrectly.

[0141] Furthermore, in order to prevent the same content trigger device from being illegally shared by multiple people, Figure 6 Fig. shows a specific flowchart for unlocking preset content in step S320. Please refer to Figure 6 , step S320 includes:

[0142] S321, the learning machine obtains the device ID corresponding to the content trigger device.

[0143] S322, the learning machine sends a verification request to the server, and the device ID is carried in the verification request.

[0144] S323, the server obtains the unlocking times corresponding to the device ID carried in the verification request according to the verification request, and verifies whether the unlocking times are greater than a preset threshold, and generates a corresponding verification result.

[0145] When the unlocking times corresponding to the device ID are greater than the preset threshold, the server returns a "not passed" verification result to the learning machine; otherwise, the server returns a "passed" verification result to the learning machine.

[0146] S324, the server returns the verification result to the learning machine.

[0147] S325, when the verification result is passed, the learning machine unlocks the preset content corresponding to the identification information sent by the content trigger device.

[0148] S326, after the unlocking is completed, the learning machine sends an unlocking notification to the server.

[0149] Among them, the unlocking notification carries the device ID of the content trigger device.

[0150] S327, the server increments the unlocking times corresponding to the device ID by 1.

[0151] Each content trigger device corresponds to a unique device ID. Through the above process, the server manages the unlocking times of each content trigger device according to the device ID. When the unlocking times of a certain content trigger device exceed the preset threshold, a non-pass verification result is returned to the learning machine. Then, the learning machine will not unlock the corresponding preset content, thus ensuring that each content trigger device cannot be shared infinitely by several illegal users.

[0152] The embodiment of the present disclosure also provides an interactive answering device. Please refer to Figure 7 , the interactive answering device 400 includes:

[0153] An image acquisition module 410, configured to acquire a target image collected by a camera;

[0154] A pose recognition module 420, configured to perform pose recognition on the target image through a pose recognition model to obtain a plurality of skeletal key points;

[0155] An answering option determination module 430, configured to perform clustering analysis on the plurality of skeletal key points through a clustering network to determine the answering option represented by the pose of the user in the target image in the image.

[0156] Optionally, the answering option determination module 430 includes:

[0157] A deviation vector calculation sub-module, configured to select a key point from the plurality of skeletal key points as a reference point, calculate the deviation value between each skeletal key point and the reference point respectively, and obtain a deviation vector according to all the deviation values;

[0158] An option determination sub-module, configured to input the deviation vector into the clustering network, obtain the cluster corresponding to the deviation vector, and determine the corresponding answering option according to the label of the cluster.

[0159] Optionally, the deviation vector calculation sub-module is configured to, when the length of the vector composed of all the deviation values does not reach the input length requirement of the clustering network, fill zeros at the end of the vector until the length of the vector reaches the input length requirement to obtain the deviation vector.

[0160] Optionally, in the deviation vector, the values of 0 among all the deviation values are located after the non-zero values.

[0161] Optionally, the interactive answering device 400 further includes:

[0162] A first coordinate acquisition module, configured to acquire a first coordinate measured by a somatosensory answering device at the current moment. The somatosensory answering device is disposed on a body part of the user and is used to measure the coordinate of the part;

[0163] A second coordinate acquisition module, configured to acquire a second coordinate of a target bone key point corresponding to the measured part of the somatosensory answer device among the multiple bone key points at the current moment, where the origin of the first coordinate is the same as that of the second coordinate;

[0164] A coordinate correction module, configured to correct the second coordinate according to the first coordinate to obtain a corrected coordinate, and replace the second coordinate of the target bone key point among the multiple bone key points with the corrected coordinate.

[0165] Optionally, the somatosensory answer device is a wearable device provided on the user's hand or a device worn on the user's finger.

[0166] Optionally, the coordinate correction module includes:

[0167] A two-dimensional conversion sub-module, configured to determine the two-dimensional coordinate corresponding to the first coordinate;

[0168] A proportional relationship calculation sub-module, configured to calculate a first proportional relationship between the moving distance of the part and the image pixels according to the two-dimensional coordinate and the second coordinate at the current moment, and the two-dimensional coordinate and the second coordinate at the previous moment, where the previous moment is the moment when the pose recognition model is used for pose recognition for the previous time to obtain multiple bone key points;

[0169] A coordinate conversion sub-module, configured to calculate a target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image according to the first proportional relationship;

[0170] A coordinate correction sub-module, configured to correct the second coordinate according to the target coordinate to obtain the corrected coordinate.

[0171] Optionally, the first proportional relationship is a ratio of a first Euclidean distance to a second Euclidean distance, where the first Euclidean distance is the Euclidean distance between the second coordinate at the current moment and the second coordinate at the previous moment, and the second Euclidean distance is the Euclidean distance between the two-dimensional coordinate at the current moment and the two-dimensional coordinate at the previous moment.

[0172] Optionally, the coordinate conversion sub-module is configured to:

[0173] Calculate an average value of the first proportional relationship according to the first proportional relationship corresponding to each moment within a preset time range at the current moment and before, to obtain a second proportional relationship; where each moment is the moment when the pose recognition model is used for pose recognition once to obtain multiple bone key points;

[0174] Calculate a target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image according to the second proportional relationship.

[0175] Optionally, the coordinate conversion sub-module is configured to:

[0176] According to each first proportional relationship corresponding to each moment within a preset time range at and before the current moment, calculate a weighted average value of the first proportional relationships to obtain the second proportional relationship; wherein, each first proportional relationship at each moment corresponds to a weighting coefficient, the weighting coefficient is greater than 0 and less than 1, and the weighting coefficient is negatively correlated with the ratio of the first proportional relationship between the current moment and the previous moment.

[0177] Optionally, the weighting coefficient is determined by the following formula:

[0178]

[0179] wherein, α t is the weighting coefficient corresponding to the moment, is the first proportional relationship at the corresponding moment, is the first proportional relationship at the previous moment of the corresponding moment, 0 < m < 1.

[0180] Optionally, the pose recognition module 420 is configured to:

[0181] Detect the target image through a target detection model to determine whether the area containing the user belongs to the whole body or the hand in the case that the target image contains the user;

[0182] If the area containing the user belongs to the whole body, perform whole-body pose recognition on the target image through a human body pose recognition model to obtain multiple skeletal key points of the whole body;

[0183] If the area containing the user belongs to the hand, perform hand pose recognition on the target image through a hand pose recognition model to obtain multiple skeletal key points of the hand.

[0184] Optionally, the image acquisition module 410 is configured to:

[0185] Write the video collected by the camera into a queue to form a video frame image queue;

[0186] Take out the images in the video frame image queue in sequence and transmit them to the display unit for display;

[0187] Determine the image at the top of the video frame image queue at the current moment as the target image.

[0188] Optionally, the interactive question-and-answer device 400 further includes:

[0189] An identification information receiving module, configured to receive the identification information sent by the content triggering device;

[0190] A preset content unlocking module, configured to unlock preset content corresponding to the identification information according to the identification information, where the preset content includes virtual characters, virtual clothing, or plots.

[0191] Optionally, the preset content unlocking module includes:

[0192] A request verification sub-module, configured to obtain the device ID corresponding to the content trigger device and send a verification request to the server, where the device ID is carried in the verification request;

[0193] A result acquisition sub-module, configured to receive the verification result returned by the server, where the verification result is generated by the server according to the unlocking times corresponding to the device ID and a preset threshold. When the unlocking times are greater than the preset threshold, the verification result fails; otherwise, the verification result passes;

[0194] A content unlocking sub-module, configured to unlock the preset content corresponding to the identification information when the verification result passes;

[0195] An unlocking notification sub-module, configured to send an unlocking notification to the server after unlocking is completed, so that the server increments the unlocking times corresponding to the device ID by 1.

[0196] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0197] Figure 8 It is a block diagram of a learning machine 500 shown according to an exemplary embodiment. As Figure 8 shown, the learning machine 500 may include: a processor 501, a memory 502, and a camera 503. The camera 503 is configured to collect an image in front and transmit the image to the processor 501; the memory 502 is configured to store a computer program; the processor 501 is configured to execute the computer program in the memory 502 according to the image transmitted by the camera 503. When the computer program is executed by the processor 501, the steps of the interactive answering method provided by the embodiments of the present disclosure can be implemented. Optionally, the learning machine 500 may further include one or more of a multimedia component 504, an input / output (I / O) interface 505, and a communication component 506.

[0198] Among them, the processor 501 is used to control the overall operation of the learning machine 500 to complete all or part of the steps in the above interactive answering method. The memory 502 is used to store various types of data to support the operation of the learning machine 500. These data may include, for example, instructions for any application or method operating on the learning machine 500, as well as application-related data, such as courseware data, preset content data (including data on virtual characters, virtual costumes, and plots), contact data, sent and received messages, pictures, audio, video, and so on. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 504 may include a screen and an audio component. Through this multimedia component, the learning machine 500 can play courseware on the screen, display various preset contents, display a video frame image queue, and so on. Among them, the screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 505 provides an interface between the processor 501 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 506 is used for wired or wireless communication between the learning machine 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, NFC, 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them is not limited here. Therefore, the corresponding communication component 506 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc. Through this communication component, the learning machine can obtain the first coordinate from the somatosensory answering device, obtain the identification information and the corresponding device ID stored in the content trigger device from the content trigger device, send a verification request and an unlocking notification to the server, and obtain a verification result from the server.

[0199] Optionally, the learning machine 500 further includes a somatosensory answer device. The somatosensory answer device includes a gyroscope and a communication module. The somatosensory answer device is configured to be disposed on a body part of a user. The gyroscope is configured to measure a first coordinate of the part, and the communication module is configured to send the first coordinate measured by the gyroscope to the processor 501 of the learning machine 500. The somatosensory answer device may be a wearable device disposed on the user's hand, so that the gyroscope is configured to measure the first coordinate of the user's hand, or the somatosensory answer device may be a device worn on the user's finger, so that the gyroscope is configured to measure the first coordinate of the finger joint where it is worn.

[0200] In an exemplary embodiment, the learning machine 500 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is configured to execute the above interactive answering method.

[0201] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions. When the program instructions are executed by a processor, the steps of the above interactive answering method are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 502 including program instructions, and the above program instructions may be executed by the processor 501 of the learning machine 500 to complete the above interactive answering method.

[0202] In another exemplary embodiment, there is also provided a computer program product. The computer program product includes a computer program executable by a programmable device. The computer program has a code portion for executing the above interactive answering method when executed by the programmable device.

[0203] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0204] In addition, it should be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.

[0205] Furthermore, any combinations can be made among the various different embodiments of the present disclosure, as long as they do not violate the idea of the present disclosure, and they should equally be regarded as the content disclosed by the present disclosure.

Claims

1. An interactive question answering method, characterized in that, Including: Obtain a target image collected by a camera; Perform pose recognition on the target image through a pose recognition model to obtain a plurality of skeletal key points; Obtain a first coordinate measured by a somatosensory answer device at the current moment, where the somatosensory answer device is arranged on a body part of a user and is used to measure the coordinate of the part; Obtain a second coordinate of a target skeletal key point corresponding to the part measured by the somatosensory answer device among the plurality of skeletal key points at the current moment, and the origin of the first coordinate is the same as that of the second coordinate; Correct the second coordinate according to the first coordinate to obtain a corrected coordinate; Replace the second coordinate of the target skeletal key point among the plurality of skeletal key points with the corrected coordinate; Perform clustering analysis on the plurality of skeletal key points through a clustering network to determine an answer option represented by the pose of the user in the target image in the image; The performing clustering analysis on the plurality of skeletal key points through a clustering network to determine an answer option represented by the pose of the user in the target image in the image includes: Select a key point from the plurality of skeletal key points as a reference point; Calculate the deviation value between each skeletal key point and the reference point respectively, and obtain a deviation vector according to all the deviation values; Input the deviation vector into the clustering network to obtain a cluster corresponding to the deviation vector; Determine the corresponding answer option according to the label of the cluster.

2. The method according to claim 1, characterized in that, The obtaining a deviation vector according to all the deviation values includes: If the length of the vector composed of all the deviation values does not reach the input length requirement of the clustering network, fill zeros at the end of the vector until the length of the vector reaches the input length requirement to obtain the deviation vector.

3. The method according to claim 2, wherein In the deviation vector, the values of 0 among all the deviation values are located after the non-zero values.

4. The method according to claim 1, characterized in that The somatosensory answer device is a wearable device arranged on the user's hand or a device worn on the user's finger.

5. The method according to claim 1, characterized in that The correcting the second coordinate according to the first coordinate to obtain a corrected coordinate includes: Determine the two-dimensional coordinate corresponding to the first coordinate; Calculate a first proportional relationship between the distance moved by the part and the image pixels according to the two-dimensional coordinate and the second coordinate at the current moment, and the two-dimensional coordinate and the second coordinate at the previous moment, where the previous moment is the moment when pose recognition is performed through the pose recognition model to obtain a plurality of skeletal key points last time; Calculate a target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image according to the first proportional relationship; Correct the second coordinate according to the target coordinate to obtain the corrected coordinate.

6. The method according to claim 5, wherein The first proportional relationship is the ratio of the first Euclidean distance to the second Euclidean distance, where the first Euclidean distance is the Euclidean distance between the second coordinate at the current moment and the second coordinate at the previous moment, and the second Euclidean distance is the Euclidean distance between the two-dimensional coordinate at the current moment and the two-dimensional coordinate at the previous moment.

7. The method according to claim 6, wherein The calculating a target coordinate corresponding to the two-dimensional coordinate at the current moment on the target image according to the first proportional relationship includes: Calculate the average value of the first proportional relationships corresponding to each moment within the preset time range at the current moment and before, to obtain the second proportional relationship; where each moment is a moment when multiple skeletal key points are obtained through one pose recognition by the pose recognition model. Calculate the target coordinates on the target image corresponding to the two-dimensional coordinates at the current moment according to the second proportional relationship.

8. The method according to claim 7, characterized in that, The calculating the average value of the first proportional relationships corresponding to each moment within the preset time range at the current moment and before, to obtain the second proportional relationship includes: Calculate the weighted average value of the first proportional relationships corresponding to each moment within the preset time range at the current moment and before, to obtain the second proportional relationship; where each first proportional relationship at each moment corresponds to a weighting coefficient, the weighting coefficient is greater than 0 and less than 1, and the weighting coefficient is negatively correlated with the ratio of the first proportional relationship at the current moment to that at the previous moment.

9. The method according to claim 8, wherein The weighting coefficient is determined by the following formula: ; wherein, is the weighting coefficient at the corresponding moment, is the first proportional relationship at the corresponding moment, is the first proportional relationship at the previous moment of the corresponding moment, 0 < m < 1.

10. The method according to claim 1, wherein The obtaining multiple skeletal key points by performing pose recognition on the target image through the pose recognition model includes: Detect the target image through the target detection model to determine whether the area containing the user belongs to the whole body or the hand when the user is included in the target image. If the area containing the user belongs to the whole body, perform whole-body pose recognition on the target image through the human pose recognition model to obtain multiple skeletal key points of the whole body. If the area containing the user belongs to the hand, perform hand pose recognition on the target image through the hand pose recognition model to obtain multiple skeletal key points of the hand.

11. The method according to claim 1, wherein The obtaining the target image collected by the camera includes: Write the video collected by the camera into a queue to form a video frame image queue. Take out the images in the video frame image queue in sequence and transmit them to the display unit for display. Determine the image at the top of the video frame image queue at the current moment as the target image.

12. The method according to any one of claims 1-11, characterized in that, The method further includes: Receive the identification information sent by the content trigger device. Unlock the preset content corresponding to the identification information according to the identification information, where the preset content includes virtual characters, virtual clothing or plots.

13. The method according to claim 12, wherein The unlocking the preset content corresponding to the identification information according to the identification information includes: Obtain the device ID corresponding to the content trigger device. Send a verification request to the server, where the device ID is carried in the verification request. Receive the verification result returned by the server, where the verification result is generated by the server according to the unlocking times corresponding to the device ID and the preset threshold. When the unlocking times are greater than the preset threshold, the verification result is not passed, otherwise the verification result is passed. If the verification result is passed, unlock the preset content corresponding to the identification information. After the unlocking is completed, send an unlocking notification to the server so that the server increments the unlocking times corresponding to the device ID by 1.

14. An interactive question-and-answer device, characterized in that, including: An image acquisition module, configured to acquire the target image collected by the camera. The pose recognition module is used to perform pose recognition on the target image through a pose recognition model to obtain multiple skeletal key points; The first coordinate acquisition module is used to acquire the first coordinate measured by the somatosensory answer device at the current moment. The somatosensory answer device is arranged on a body part of the user and is used to measure the coordinate of the part; The second coordinate acquisition module is used to acquire the second coordinate of the target skeletal key point corresponding to the part measured by the somatosensory answer device among the multiple skeletal key points at the current moment. The origin of the first coordinate is the same as that of the second coordinate; The coordinate correction module is used to correct the second coordinate according to the first coordinate to obtain a corrected coordinate, and replace the second coordinate of the target skeletal key point among the multiple skeletal key points with the corrected coordinate; The answer option determination module is used to perform clustering analysis on the multiple skeletal key points through a clustering network to determine the answer option represented by the pose of the user in the target image; The answer option determination module includes: The deviation vector calculation sub-module is used to select a key point from the multiple skeletal key points as a reference point, calculate the deviation value between each skeletal key point and the reference point respectively, and obtain a deviation vector according to all the deviation values; The option determination sub-module is used to input the deviation vector into the clustering network, obtain the cluster corresponding to the deviation vector, and determine the corresponding answer option according to the label of the cluster.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-13.

16. A learning machine, characterized in that, Including: A camera for collecting images; A memory on which a computer program is stored; A processor for executing the computer program in the memory to implement the method according to any one of claims 1-13; A somatosensory answer device, wherein the somatosensory answer device includes a gyroscope and a communication module. The somatosensory answer device is used to be arranged on a body part of the user. The gyroscope is used to measure the first coordinate of the part, and the communication module is used to send the first coordinate measured by the gyroscope to the processor.

17. The learning machine according to claim 16, characterized in that, The somatosensory answer device is a wearable device arranged on the user's hand or a device worn on the user's finger.

Citation Information

Patent Citations

  • Human body action recognition method and visual enhancement processing system

    CN111274854A

  • Human body posture recognition method and device based on skeleton key points, storage medium and terminal

    CN111680562A

  • Real-time monitoring method, device and equipment for student discipline violation based on visual perception

    CN113298005A