Motion capture-based all-in-one machine interaction method and its visual interaction all-in-one machine

By applying an all-in-one interaction method based on motion capture in teaching, and using a visual interactive all-in-one machine to display the dynamic image of users with teaching videos, the problem of lack of interaction and fun in the existing teaching methods is solved, and students' learning enthusiasm and teaching effectiveness are improved.

CN119200835BActive Publication Date: 2025-05-27SHENZHEN YUANJUN INTEGRATED TECH DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411242724.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-05-27
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

The existing teaching methods lack interactivity and fun, which leads to a decrease in students' enthusiasm for learning, thereby reducing teaching effectiveness.

Method used

The all-in-one interaction method based on motion capture is adopted, and the dynamic images of users are displayed simultaneously with the teaching video through the visual interactive all-in-one. The feature pyramid network and multi-scale hollow convolution mask are used to perform multi-scale feature fusion and sparse processing to achieve target recognition and rhythmic display of dynamic images.

Benefits of technology

It enhances the interaction between teaching users and students, increases the fun of teaching, improves students' learning enthusiasm, and thus improves teaching effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119200835B_ABST
    Figure CN119200835B_ABST
Patent Text Reader

Abstract

The present invention discloses an integrated machine interaction method based on motion capture and its visual interaction integrated machine, including: responding to an online teaching request of a teaching user, acquiring a teaching video of the teaching user during the teaching process based on a camera in the visual integrated machine, and teaching action images corresponding to each frame in the teaching video; using a feature pyramid network to perform multi-scale feature fusion on the target features of the teaching action images to obtain feature maps of multiple levels, the feature maps of multiple levels including a first feature map of a target level and second feature maps of other levels; using a multi-scale dilated convolution mask to sparsify the first feature map to obtain a sparse feature map; performing target recognition on the second feature maps and the sparse feature map to obtain a user dynamic image corresponding to the teaching action image; and displaying the user dynamic image and the teaching video on the visual interaction integrated machine, and the user dynamic image rhythmically moves following the playback of the teaching video. The present invention improves the teaching effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an all-in-one machine interaction method based on motion capture and its visual interactive all-in-one machine. Background Art

[0002] With the rapid development of the Internet of Things, teaching methods have also diversified, such as online teaching and video teaching through the Internet of Things. For video teaching, mainly the teaching videos recording complete lectures or course content are played and displayed through terminal devices. Therefore, students can learn through teaching videos, resulting in a lack of interaction between teachers and students. For online teaching, mainly teachers conduct one-way knowledge transfer to students through online teaching, resulting in a lack of face-to-face discussion and communication between teachers and students. Therefore, the existing teaching methods lack interest, which is very likely to lead to a decline in students' learning enthusiasm, thus reducing the teaching effect. Summary of the Invention

[0003] The present invention provides an all-in-one machine interaction method based on motion capture and its visual interactive all-in-one machine to solve the problem that the existing teaching methods lack interest and improve the teaching effect.

[0004] In a first aspect, the present invention provides an all-in-one machine interaction method based on motion capture, which is applied to a visual interactive all-in-one machine. A fixed camera is installed in the visual interactive all-in-one machine. The all-in-one machine interaction method based on motion capture includes:

[0005] Respond to the online teaching request of a teaching user, and obtain the teaching video of the teaching user during the teaching process and the teaching action images corresponding to each frame in the teaching video based on the camera in the visual all-in-one machine;

[0006] Use a feature pyramid network to perform multi-scale feature fusion on the target features of the teaching action images to obtain feature maps of multiple levels, where the feature maps of multiple levels include the first feature map of the target level and the second feature maps of other levels;

[0007] Use a multi-scale dilated convolution mask to sparsify the first feature map to obtain a sparse feature map;

[0008] Perform target recognition on the second feature map and the sparse feature map to obtain the user dynamic image corresponding to the teaching action image;

[0009] Display the user dynamic image and the teaching video on the visual interactive all-in-one machine, where the user dynamic image moves rhythmically following the playback of the teaching video.

[0010] In a second aspect, the present invention provides a visual interactive all-in-one machine, in which a fixedly arranged camera and a chassis are installed. The visual interactive all-in-one machine includes an acquisition unit, a feature fusion unit, an image processing unit, an object recognition unit, and a display and interaction unit;

[0011] The acquisition unit is configured to respond to an online teaching request of a teaching user, and acquire a teaching video of the teaching user during the teaching process and a teaching action image corresponding to each frame in the teaching video based on the camera in the visual all-in-one machine;

[0012] The feature fusion unit is configured to perform multi-scale feature fusion on the target features of the teaching action image by using a feature pyramid network to obtain feature maps of multiple levels, and the feature maps of multiple levels include a first feature map of a target level and second feature maps of other levels;

[0013] The image processing unit is configured to perform sparsification processing on the first feature map by using a multi-scale dilated convolutional mask to obtain a sparse feature map;

[0014] The object recognition unit is configured to perform object recognition on the second feature map and the sparse feature map to obtain a user dynamic image corresponding to the teaching action image;

[0015] The display and interaction unit is configured to display the user dynamic image and the teaching video on the visual interactive all-in-one machine, wherein the user dynamic image moves rhythmically following the playback of the teaching video.

[0016] In a third aspect, the present invention further provides a visual interactive all-in-one machine, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the all-in-one machine interaction method based on motion capture as described in the first aspect.

[0017] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the all-in-one machine interaction method based on motion capture as described in the first aspect.

[0018] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the all-in-one machine interaction method based on motion capture as described in the first aspect.

[0019] The integrated machine interaction method based on motion capture provided by the present invention can jointly display the user's dynamic image and the teaching video on the visual interaction integrated machine, and the user's dynamic image rhythmically moves along with the playing of the teaching video, enhancing the interaction between teaching users and students, increasing the teaching interest, improving the learning enthusiasm of students, and thus improving the teaching effect. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of the integrated machine interaction method based on motion capture provided by the present invention;

[0022] Figure 2 It is a schematic structural diagram of the target recognition model provided by the present invention;

[0023] Figure 3 It is a schematic diagram of the position calibration process provided by the present invention;

[0024] Figure 4 It is a schematic structural diagram of the visual interaction integrated machine provided by the present invention;

[0025] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0027] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0028] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but rather to be in line with the broadest scope consistent with the principles and features disclosed in the present invention.

[0029] Figure 1 FIG. 4 is a schematic flow chart of an integrated machine interaction method based on motion capture provided by the present invention. The integrated machine interaction method based on motion capture is applied to a visual interaction integrated machine. A fixed camera is installed in the visual interaction integrated machine. The integrated machine interaction method based on motion capture includes:

[0030] Step 10: Respond to an online teaching request of a teaching user, and obtain a teaching video of the teaching user during the teaching process and a teaching action image corresponding to each frame in the teaching video based on the camera in the visual integrated machine.

[0031] It should be noted that when the teaching user needs to conduct online teaching, a request needs to be triggered in the visual interaction integrated machine. Therefore, the visual interaction integrated machine responds to the online teaching request of the teaching user. At this time, the visual interaction integrated machine calls the camera in the visual integrated machine to record the teaching process of the teaching user, so as to obtain the teaching video of the teaching user during the teaching process. At the same time, the visual interaction integrated machine obtains the teaching action image corresponding to each frame in the teaching video.

[0032] Step 20: Use a Feature Pyramid Network to perform multi-scale feature fusion on the target features of the teaching action images to obtain feature maps of multiple levels. The feature maps of multiple levels include a first feature map of the target level and second feature maps of other levels.

[0033] Among them, the Feature Pyramid Network (FPN) is a deep neural network widely used in tasks such as object recognition and semantic segmentation. The Feature Pyramid Network constructs a feature pyramid to extract and fuse feature information of different scales. The feature pyramid contains multiple levels, and each level corresponds to a feature map of a different scale.

[0034] The levels of the feature pyramid generally include: the bottom-up feature extraction levels (such as C2, C3, C4, C5) and the levels generated after top-down feature fusion (such as P2, P3, P4, P5). These levels provide rich multi-scale feature information for the object recognition task.

[0035] Exemplarily, the target level can be the P3 level. The feature map of the P3 level usually has a moderate resolution, which helps to capture the features of the object. Since the object in the teaching action image has a small pixel area, sufficient resolution is required to retain its key information. And the P3 level happens to be at the balance point between high resolution and strong semantic information. Moreover, the P3 level fuses the strong semantic information from high levels (such as P4, P5) and the high-resolution information from low levels (such as C3) through lateral connections. This fusion makes the feature map of the P3 level contain both rich semantic information and sufficient image details, which helps to improve the object recognition performance.

[0036] After extracting the object features of the teaching action image, the visual interactive all-in-one machine can use the feature pyramid network to perform multi-scale feature fusion on the object features to obtain feature maps of multiple levels. It can be understood that each level corresponds to a feature map of a different scale. The feature map corresponding to the target level (such as the P3 level) in multiple levels is the first feature map, and the feature maps corresponding to other levels except the target level in multiple levels are the second feature maps. There are as many second feature maps as there are other levels.

[0037] In one embodiment, the all-in-one machine interaction method based on motion capture may further include: the visual interactive all-in-one machine uses a backbone network to extract features from the teaching action image to obtain object features; wherein, the backbone network may include: multiple second convolutional layers and Cross Stage Partial networks (CSP). The visual interactive all-in-one machine can use the backbone network (including multiple second convolutional layers and the CSP) to extract features from the teaching action image. Specifically, the visual interactive all-in-one machine can use a series of convolutional operations and residual connections in the backbone network to extract the key feature information in the teaching action image, that is, the object features.

[0038] Step 30, sparsify the first feature map using a multi-scale dilated convolution mask to obtain a sparse feature map.

[0039] Among them, the multi-scale dilated convolution mask (MDCM) is a module proposed by the present invention for sparsifying a feature map to obtain a sparse feature map.

[0040] After determining the feature maps of multiple levels, the visual interactive all-in-one machine can use a multi-scale dilated convolution mask to sparsify the first feature map of the target level among the feature maps of multiple levels, obtaining a sparse feature map. This entire process can not only retain the key feature information in the first feature map but also effectively filter out redundant information. The obtained sparse feature map is relatively accurate and has a low redundancy, making the subsequent target recognition process more accurate and efficient.

[0041] In one embodiment, the multi-scale dilated convolution mask may include: a first convolutional layer and a first output layer; the visual interactive all-in-one machine using the multi-scale dilated convolution mask to sparsify the first feature map to obtain a sparse feature map may include: the visual interactive all-in-one machine using the first convolutional layer to process the first feature map to obtain a sparse mask vector; the visual interactive all-in-one machine using the first output layer to sparsify the first feature map based on the sparse mask vector to obtain a sparse feature map.

[0042] It should be noted that the sparse mask vector can accurately locate the foreground region.

[0043] In the process of determining the sparse feature map, the visual interactive all-in-one machine can first use the first convolutional layer to process the first feature map to obtain a sparse mask vector. In this way, the visual interactive all-in-one machine can use the first output layer to sparsify the first feature map based on the sparse mask vector to obtain a sparse feature map. While retaining the key feature information in the first feature map, it effectively filters out the redundant information in the first feature map, reduces the unnecessary computational burden, and effectively improves the efficiency of subsequent target recognition.

[0044] The first convolutional layer may include: dilated convolutional kernels with multiple different dilation rates, at least one channel compression convolutional layer, and a second output layer; the visual interactive all-in-one machine using the first convolutional layer to process the first feature map to obtain a sparse mask vector may include: the visual interactive all-in-one machine using dilated convolutional kernels with multiple different dilation rates to extract initial information of different scales from the first feature map; the visual interactive all-in-one machine using at least one channel compression convolutional layer to perform channel compression processing on the initial information of different scales to obtain target information; the visual interactive all-in-one machine using the second output layer to determine the sparse mask vector based on the target information.

[0045] Among them, dilated convolution (Dilated Convolution, abbreviated as Dilated Conv), also known as atrous convolution, is a method to expand the receptive field of the convolutional kernel without increasing the computational cost and is increasingly widely used in target recognition. To enhance the recognition ability of the recognition head for different large targets, by setting different dilation rates, dilated convolutional kernels with multiple different dilation rates are used to extract features at different scales.

[0046] In the process of determining the sparse mask vector, the visual interactive all-in-one machine can use the atrous convolution kernels with multiple different dilation rates in the first convolutional layer to extract initial information of different scales from the first feature map. Then, the visual interactive all-in-one machine can use at least one channel compression convolutional layer in the first convolutional layer to perform channel compression processing on the initial information of different scales (the purpose is to effectively filter the redundant information in the initial information of different scales to reduce the amount of calculation), and obtain the target information. Then, the visual interactive all-in-one machine can use the second output layer to determine the sparse mask vector based on the target information. In the whole process, through the atrous convolution kernels with multiple different dilation rates, diverse receptive fields can be obtained to capture the initial information of different scales. Moreover, through at least one channel compression convolutional layer, the amount of calculation can be effectively reduced.

[0047] Step 40: Perform object recognition on the second feature map and the sparse feature map to obtain the user's dynamic image corresponding to the teaching action image.

[0048] Among them, the Head is mainly responsible for predicting the bounding box, confidence, and class probability.

[0049] After determining the sparse feature map, the visual interactive all-in-one machine can send the second feature maps of other levels and the sparse feature Figure 1 map into the Head for object recognition to obtain the user's dynamic image corresponding to the teaching action image output by the Head. This user's dynamic image is relatively accurate. In the whole process, while ensuring the recognition efficiency, the accuracy of object recognition is effectively improved.

[0050] Exemplarily, Figure 2 is a schematic structural diagram of the object recognition model provided by the present invention. As can be seen from Figure 2 it, the object recognition model can adopt the YOLOv5 framework. YOLOv5 can include: an Input layer, a Backbone network, a Neck network, and a Head. Among them, the Backbone network can include multiple second convolutional layers (Conv) and Cross Stage Partial networks (CSP). The Cross Stage Partial network includes multiple CSP modules. The Neck network includes the Feature Pyramid Network (FPN) involved above. In Figure 2In the multi-scale dilated convolutional mask (MDCM), Conv represents the channel compression convolutional layer; Dilated Conv represents the dilated convolutional kernel, and the dilation rates of the two Dilated Convs are different. In this way, the visual interactive all-in-one machine can adopt a feature pyramid network to perform multi-scale feature fusion on the target features of the teaching action image, obtaining feature maps of multiple levels. The feature maps of multiple levels include the first feature map of the target level (such as the P3 level) and the second feature maps of other levels. For the P3 level, the visual interactive all-in-one machine can use the multi-scale dilated convolutional mask to sparsify the first feature map of the P3 level, obtaining a sparse feature map, and then sending the sparse feature map and the second features of other levels Figure 1 to the recognition head for target recognition, obtaining the user's dynamic image corresponding to the teaching action image output by the recognition head.

[0051] Step 50: Display the user's dynamic image and the teaching video on the visual interactive all-in-one machine, where the user's dynamic image moves rhythmically following the playback of the teaching video.

[0052] Furthermore, the visual interactive all-in-one machine displays the user's dynamic image and the teaching video on the visual interactive all-in-one machine, where the user's dynamic image moves rhythmically following the playback of the teaching video.

[0053] The embodiment of the present invention can display the user's dynamic image and the teaching video together on the visual interactive all-in-one machine, and the user's dynamic image moves rhythmically following the playback of the teaching video, enhancing the interaction between the teaching user and the students, increasing the teaching interest, improving the learning enthusiasm of the students, and thus improving the teaching effect.

[0054] In one embodiment, in order to improve the accuracy of motion capture, before performing motion capture, it is also necessary to calibrate the position of the camera in the visual interactive all-in-one machine. Therefore, the all-in-one machine interaction method based on motion capture provided by the embodiment of the present invention further includes:

[0055] Controlling the visual interactive all-in-one machine to capture the calibration object through the camera, obtaining a first image set; the first image set includes calibration object images at different shooting distances and / or different shooting angles; calibrating the pitch angle of the camera based on the first image set to determine the pitch angle calibration result of the camera; controlling the visual interactive all-in-one machine to move to the front of the calibration object; controlling the chassis to rotate and capturing the calibration object through the camera to obtain a second image set; the second image set includes calibration object images captured at different rotation angles; calibrating the relative position between the camera and the chassis based on the second image set and the rotation angle information of the chassis to determine the position calibration result of the camera.

[0056] The first image set includes calibration object images at different shooting distances and / or different shooting angles.

[0057] Specifically, first fix the calibration object at a definite position (such as on the wall surface), making the height of the center of the calibration object basically the same as that of the camera fixedly set on the visual interaction all-in-one machine, and ensure that there are no sundries in the surrounding area of the calibration object to avoid interference with the movement of the visual interaction all-in-one machine.

[0058] Optionally, the calibration object is a calibration board.

[0059] Further, control the visual interaction all-in-one machine to take pictures of the calibration object through the camera to obtain the first image set. In one embodiment, control the visual interaction all-in-one machine to be about 2 meters away from the calibration board, start the pitch angle calibration program, and the visual interaction all-in-one machine will move autonomously according to the set pitch angle calibration program, take pictures of the calibration board through the camera, and collect a series of calibration object images with different shooting distances and / or different shooting angles to form the first image set.

[0060] After obtaining the first image set, the pitch angle calibration program can process and analyze the collected first image set. By using the PNP algorithm to solve the external parameter data of the camera, combined with the internal parameter data of the camera and the calibration object data (i.e., the known parameters of the calibration object), calibrate the pitch angle of the camera, calculate the attitude angle of the camera relative to the calibration object, and the attitude angle is denoted as (Roll, Pitch, Yaw), where Roll represents the roll angle, Yaw represents the yaw angle, and Pitch represents the pitch angle. The solved Pitch angle is the pitch angle of the 3D camera relative to the moving plane of the visual interaction all-in-one machine. By calculating multiple groups of Pitch angle data and performing weighted aggregation calculation, the final pitch angle calibration result of the camera can be determined.

[0061] Further, control the visual interaction all-in-one machine to move to a position about 0.5 meters in front of the calibration object, making the visual interaction all-in-one machine directly face the calibration object.

[0062] Further, control the chassis to rotate and take pictures of the calibration object through the camera to obtain the second image set, and the second image set includes calibration object images taken at different rotation angles.

[0063] When the visual interaction all-in-one machine is directly facing the calibration object, start the position calibration program. The visual interaction all-in-one machine will automatically control the chassis to rotate according to the set position calibration program and take pictures of the calibration object through the camera to obtain the second image set.

[0064] The visual interaction all-in-one machine can first collect a reference image I according to a set position calibration program, control the chassis to rotate a small angle θ, collect the current calibration object image, and record the current rotation angle θ; by rotating different angles, take pictures of the calibration object, collect the calibration object images taken at different rotation angles, form a second image set, and generate multiple groups of data. Each group of data includes a calibration object image and the rotation angle corresponding to the calibration object image (i.e., the rotation angle of the chassis).

[0065] Further, based on the second image set and the rotation angle information of the chassis, calibrate the relative position between the camera and the chassis, and calculate the horizontal position offset of the 3D camera relative to the center of the chassis of the visual interaction all-in-one machine, that is, the position calibration result of the camera.

[0066] In the embodiment of the present invention, there is no need for technicians to use additional acquisition tooling for data acquisition and camera calibration. Only the images collected by the camera set on the visual interaction all-in-one machine itself are needed to complete the camera calibration, which improves the accuracy of camera calibration and thus improves the accuracy of motion capture.

[0067] In an embodiment, the rotation angle information includes the rotation angle corresponding to each calibration object image in the second image set; based on the second image set and the rotation angle information of the chassis, calibrate the relative position between the camera and the chassis, and determine the position calibration result of the camera, including: obtaining the camera internal parameter data and calibration object data of the camera; based on the camera internal parameter data, calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set, calibrate the relative position between the camera and the chassis, and determine multiple position deviation values; based on the multiple position deviation values, perform weighted aggregation calculation to obtain the position calibration result of the camera.

[0068] Specifically, obtain the camera internal parameter data and calibration object data of the camera.

[0069] Among them, the calibration object data refers to the known parameters of the calibration object.

[0070] In computer vision, especially in the fields of camera calibration and stereo vision, camera internal parameter data (intrinsic parameters) and camera external parameter data (extrinsic parameters) are very important concepts, and both are related to the geometric properties and poses of the camera.

[0071] Among them, the internal camera parameters are the parameters that describe the internal properties of the camera, including focal length, principal point (optical center) coordinates, distortion coefficients, etc. The internal camera parameters are usually determined during camera calibration. The internal camera parameters of a specific camera model are usually fixed and do not change over time. Once the internal camera parameters are determined, they usually remain unchanged during the use of the camera.

[0072] The external camera parameters are the parameters that describe the position and orientation of the camera in the world coordinate system, usually including the rotation matrix and the translation vector. The external camera parameters may change at different camera positions or shooting times. For example, in stereo vision, assuming there are two cameras, the relative position and direction between the two cameras will change every time the camera is moved, resulting in changes in the external camera parameters.

[0073] Furthermore, based on the internal camera parameters, calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set, the relative position between the camera and the chassis is calibrated to determine a plurality of position deviation values.

[0074] Furthermore, weighted aggregation calculation is performed based on a plurality of position deviation values to obtain the position calibration result of the camera.

[0075] In one embodiment, based on the internal camera parameters, calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set, the relative position between the camera and the chassis is calibrated to determine a plurality of position deviation values, including: selecting any two calibration object images in the second image set as the first calibration object image and the second calibration object image respectively; determining a first position vector and a second position vector based on the internal camera parameters and the calibration object data; the first position vector is the position vector between the calibration object in the first calibration object image and the camera, and the second position vector is the position vector between the calibration object in the second calibration object image and the camera; determining the distance between the first position and the second position based on the first position vector and the second position vector; the first position is the camera position when the first calibration object image is captured, and the second position is the camera position when the second calibration object image is captured; calibrating the relative position between the camera and the chassis based on the planar geometric relationship between the first position, the second position and the center point of the chassis, the rotation angle corresponding to the first calibration object image, and the rotation angle corresponding to the second calibration object image, and calculating a position deviation value; returning to the step of selecting any two calibration object images in the second image set until a plurality of position deviation values are calculated.

[0076] After obtaining the internal camera parameters and calibration object data of the camera, the calibration of the camera position can be started.

[0077] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the position calibration process provided by an embodiment of the present invention. Since in this embodiment, the camera is fixedly installed on the visual interaction all-in-one machine, when the visual interaction all-in-one machine performs a self-rotation action (i.e., controls the chassis to rotate), the camera will perform a circular motion with the center point o of the chassis of the visual interaction all-in-one machine as the center and the position deviation length length as the radius.

[0078] Specifically, after the visual interaction all-in-one machine rotates by an angle each time, the positions of the camera before and after rotation and the center point o will form an isosceles triangle.

[0079] Specifically, after obtaining the internal camera parameters and calibration object data of the camera, the three-axis translation vector of the camera relative to the calibration board can be calculated through the hand-eye calibration function of OpenCV.

[0080] Further, in the second image set, any two calibration object images are selected and used as the first calibration object image i and the second calibration object image j respectively. Based on the internal camera parameters and calibration object data, the first position vector and the second position vector

[0081] The first position vector is the position vector between the calibration object in the first calibration object image and the camera, and the second position vector is the position vector between the calibration object in the second calibration object image and the camera.

[0082] Further, from the spatial vector relationship, it can be known that based on the first position vector and the second position vector the distance l between the first position ci and the second position cj can be determined. The calculation formula for the distance l is as follows:

[0083]

[0084] where the first position is the camera position when the camera captures the first calibration object image, and the second position is the camera position when the camera captures the second calibration object image.

[0085] Further, based on the planar geometric relationship between the first position \(c_i\), the second position \(c_j\) and the center point \(o\) of the chassis, the rotation angle corresponding to the first calibration object image, and the rotation angle corresponding to the second calibration object image, the relative position between the camera and the chassis is calibrated, and a position deviation value is calculated.

[0086] As Figure 3 shown, in the second image set, any two calibration object images are selected, and are respectively denoted as the first calibration object image 1 and the second calibration object image 2. Based on the camera internal parameter data and the calibration object data, the first position vector and the second position vector

[0087] It can be known from the spatial vector relationship that based on the first position vector and the second position vector the distance between the first position \(c_1\) and the second position \(c_2\) can be determined

[0088] Assume that when the visual interaction all-in-one machine performs a self-rotation action (i.e., controls the chassis to rotate), the camera makes a circular motion with the center point \(o\) of the chassis of the visual interaction all-in-one machine as the center and the position deviation length \(length\) as the radius. Then, the first position \(c_1\), the second position \(c_2\) and the center point \(o\) of the chassis can form an isosceles triangle with the apex angle \(\theta\). The apex angle \(\theta\) is the self-rotation angle of the chassis and can be determined based on the rotation angle corresponding to the first calibration object image and the rotation angle corresponding to the second calibration object image.

[0089] As Figure 3 shown, at this time, based on the planar geometric relationship between the first position \(c_i\), the second position \(c_j\) and the center point \(o\) of the chassis, the base angle \(\alpha\) of this isosceles triangle can be obtained. The calculation formula for the base angle \(\alpha\) is as follows:

[0090] \(\alpha = (\pi - \theta) / 2\).

[0091] According to the planar geometric relationship, the relative position between the camera and the chassis can be calibrated, and a position deviation value \(length\) is calculated. The calculation formula for the position deviation value \(length\) is as follows:

[0092]

[0093] Further, return to the step of selecting any two calibration object images in the second image set until multiple position deviation values are calculated.

[0094] Specifically, based on the multiple calibration object images in the collected second image set, multiple position deviation values are respectively calculated, and are respectively denoted as \(length1\), \(length2\), \(length3\), \(length4\), etc.

[0095] Further, based on multiple position deviation values, weighted aggregation calculation is performed to obtain the position calibration result of the camera

[0096] In the embodiment of the present invention, the vision interactive all-in-one machine rotates itself to collect calibration object image data from different perspectives, combines the angle data of the chassis rotation, and utilizes the spatial and planar geometric relationships. Through simple calculations, the calibration of the relative position between the camera and the chassis can be achieved, and the accurate calculation of the position deviation of the camera relative to the chassis of the vision interactive all-in-one machine can be realized.

[0097] Based on multiple position deviation values, weighted aggregation calculation is performed to obtain the position calibration result of the camera, including: selecting one position deviation value as the reference data, and taking the remaining position deviation values as the comparison data; respectively determining the distance differences between each comparison data and the reference data; comparing each distance difference with a first preset threshold to obtain multiple comparison results; based on each comparison result, determining the error count of the reference data; returning to the step of selecting one position deviation value as the reference data and taking the remaining position deviation values as the comparison data until the error counts corresponding to all position deviation values are obtained; based on the error counts corresponding to each position deviation value, abnormal data is removed to obtain multiple normal position deviation values; based on the multiple normal position deviation values, weighted aggregation calculation is performed to obtain the position calibration result of the camera.

[0098] It can be understood that when collecting the calibration object images, when the vision interactive all-in-one machine has slight movement or the shooting angle is not good, there may be blurred or unclear images in the collected calibration object images, and such image data will affect the calculation of the position deviation values.

[0099] Therefore, it is necessary to remove the abnormal position deviation values and then perform weighted aggregation calculation to effectively improve the accuracy of the position calibration result.

[0100] In this embodiment, a statistical voting method is used to identify and remove abnormal position deviation values, and the specific steps are as follows:

[0101] (1) Select one position deviation value as the reference data, and take the remaining position deviation values as the comparison data.

[0102] Specifically, among the multiple position deviation values, the first position deviation value is selected as the reference data, and the remaining position deviation values are taken as the comparison data.

[0103] (2) Respectively determine the distance differences between each comparison data and the reference data.

[0104] Specifically, each remaining position deviation value is compared and analyzed with the reference data (i.e., the first position deviation value) in turn: for each remaining position deviation value, the distance difference between this remaining position deviation value and the reference data is calculated.

[0105] For example, calculate the distance difference Δdis between the 2nd position deviation value length 2 and the reference data length 1 . Then the calculation formula is as follows:

[0106] Δdis = |length 2 - length 1 |.

[0107] (3) Compare each distance difference Δdis with the first preset threshold respectively to obtain multiple comparison results.

[0108] Optionally, the first preset threshold is 0.1 m (meter).

[0109] For each distance difference Δdis, if this distance difference Δdis is greater than the first preset threshold, it can be considered that there is a large error between this remaining position deviation value and the reference data; if this distance difference Δdis is less than or equal to the first preset threshold, it can be considered that the error between this remaining position deviation value and the reference data is small.

[0110] (4) Based on each comparison result, determine the error count of the reference data.

[0111] Specifically, if the distance difference Δdis is greater than the first preset threshold, the error count of this reference data is incremented by 1. Based on each comparison result, the error count of the reference data can be determined.

[0112] It can be understood that the error count of a reference data should start incrementing from 0.

[0113] (5) Return to the step of selecting one position deviation value as the reference data and using the remaining position deviation values as the comparison data until the error counts corresponding to all position deviation values are obtained.

[0114] Specifically, for each reference data, loop through all the remaining position deviation values. After one round of comparison, replace with a new reference data, and successively use length2, length3, length4, etc. as the reference data, repeating the above comparison process until the error counts corresponding to all position deviation values are obtained.

[0115] (6) Based on the error count corresponding to each position deviation value, perform abnormal data elimination to obtain multiple normal position deviation values; based on the multiple normal position deviation values, perform weighted aggregation calculation to obtain the position calibration result of the camera.

[0116] Specifically, based on each position deviation value, sort them in descending order according to the error count value to obtain a sorting result; take the top 20% of the position deviation values in the sorting result as abnormal data, and remove the top 20% of the position deviation values in the sorting result. The remaining position deviation values are normal position deviation values.

[0117] Furthermore, based on multiple normal position deviation values, perform weighted aggregation calculation to obtain the position calibration result of the camera.

[0118] Considering that the calculation of abnormal position deviation values is based on images, when abnormal position deviation values occur, it may be that there are problems with the images. Therefore, the images corresponding to the abnormal position deviation values should be removed, and then the position calibration result of the camera should be recalculated based on the remaining normal images.

[0119] Preferably, based on each position deviation value, sort them in descending order according to the error count value to obtain a sorting result. After that, take the calibration object images corresponding to the top 20% of the position deviation values in the sorting result as abnormal images, remove all abnormal images, and retain multiple normal images; based on the normal images, calibrate the relative position between the camera and the chassis to determine the position calibration result of the camera.

[0120] It should be noted that in the prior art, there is little research on the verification of the accuracy of the data set. Most research defaults that the collected data set is normal and can be directly used. Therefore, the collected image data is not verified and screened. The present invention performs secondary screening on the collected images, traverses the position deviation values, compares the differences pairwise, removes the image data based on the statistical voting results, and calculates the final position calibration result according to the remaining images, which can further improve the accuracy of the calibration result.

[0121] The embodiment of the present invention adopts a data removal strategy to remove abnormal data, identifies and removes abnormal data by means of statistical voting, ensures the reliability and accuracy of the final result, and enhances the robustness of the algorithm.

[0122] In an embodiment, based on the first image set, calibrate the pitch angle of the camera to determine the pitch angle calibration result of the camera, including: obtaining the external camera parameters of the camera, the internal camera parameters of the camera, and the calibration object data; based on the external camera parameters of the camera, the internal camera parameters of the camera, the calibration object data, and the first image set, calibrate the pitch angle of the camera to determine multiple pitch angles; based on multiple pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera.

[0123] Specifically, after obtaining the first image set, the pitch angle calibration program can process and analyze the collected first image set, solve the external camera parameters of the camera by using the PNP algorithm, and obtain the internal camera parameters of the camera and the calibration object data (i.e., the known parameters of the calibration object).

[0124] Furthermore, based on the external camera parameters, internal camera parameters, calibration object data, and the first image set, calibrate the pitch angle of the camera, calculate multiple attitude angles of the camera relative to the calibration object. The attitude angles are denoted as (Roll, Pitch, Yaw), where Roll represents the roll angle, Yaw represents the yaw angle, and Pitch represents the pitch angle.

[0125] Furthermore, based on multiple pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera.

[0126] In the embodiment of the present invention, by collecting multiple groups of calibration object images and combining the external parameter calibration algorithm and the weighted average method, the pitch angle calibration of the camera is realized, with simple operation and low computational complexity.

[0127] In one embodiment, based on multiple pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera, including: selecting one pitch angle as the reference data, and using the remaining pitch angles as the comparison data; respectively determining the angle differences between each comparison data and the reference data; comparing each angle difference with a second preset threshold to obtain multiple comparison results; based on each comparison result, determining the error count of the reference data; returning to the step of selecting one pitch angle as the reference data and using the remaining pitch angles as the comparison data until the error counts corresponding to all pitch angles are obtained; based on the error counts corresponding to each pitch angle, perform abnormal data rejection to obtain multiple normal pitch angles; based on the multiple normal pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera.

[0128] It can be understood that when collecting calibration object images, when there is slight movement of the visual interaction all-in-one machine or the shooting angle is not good, the collected calibration object images may contain blurred or unclear images, and such image data will affect the calculation of the pitch angle.

[0129] Therefore, it is necessary to reject abnormal pitch angles and then perform weighted aggregation calculation to effectively improve the accuracy of the pitch angle calibration result.

[0130] In this embodiment, the method of statistical voting is used to identify and reject abnormal pitch angles. The specific steps are as follows:

[0131] (1) Select one pitch angle as the reference data, and use the remaining pitch angles as the comparison data.

[0132] Specifically, among multiple pitch angles, select the first pitch angle as the reference data, and use the remaining pitch angles as comparison data.

[0133] (2) Determine the angular difference between each piece of comparison data and the reference data respectively.

[0134] Specifically, compare each of the remaining pitch angles with the reference data (i.e., the first pitch angle) in turn: for each of the remaining pitch angles, calculate the angular difference between this remaining pitch angle and the reference data.

[0135] (3) Compare each angular difference with a second preset threshold respectively to obtain multiple comparison results.

[0136] For each angular difference, if the angular difference is greater than the second preset threshold, it can be considered that there is a large error between this remaining pitch angle and the reference data; if the angular difference is less than or equal to the second preset threshold, it can be considered that the error between this remaining pitch angle and the reference data is small.

[0137] (4) Based on each comparison result, determine the error count of the reference data.

[0138] Specifically, if the angular difference is greater than the second preset threshold, the error count of this reference data is incremented by 1, and the error count of the reference data can be determined based on each comparison result.

[0139] It can be understood that the error count of a reference data should start increasing from 0.

[0140] (5) Return to the step of selecting one pitch angle as the reference data and using the remaining pitch angles as comparison data until the error counts corresponding to all pitch angles are obtained.

[0141] (6) Based on the error count corresponding to each pitch angle, perform abnormal data rejection to obtain multiple normal pitch angles; based on the multiple normal pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera.

[0142] Specifically, based on each pitch angle, sort them in descending order according to the error count value to obtain a sorting result; take the first 20% of the pitch angles in the sorting result as abnormal data, and reject the first 20% of the pitch angles in the sorting result, and the remaining pitch angles are normal pitch angles.

[0143] Furthermore, based on the multiple normal pitch angles, perform weighted aggregation calculation to obtain the pitch angle calibration result of the camera.

[0144] Considering that the calculation of the abnormal pitch angle is based on the image, when the pitch angle appears, there may be a problem with the image. Therefore, the image corresponding to the abnormal pitch angle should be excluded, and then the pitch angle calibration result of the camera should be recalculated based on the remaining normal images.

[0145] Preferably, based on each pitch angle, sort them in descending order according to the error count value. After obtaining the sorting result, take the calibration object images corresponding to the top 20% of the pitch angles in the sorting result as abnormal images, exclude all abnormal images, and retain multiple normal images; based on the normal images, calibrate the pitch angle of the camera to determine the pitch angle calibration result of the camera.

[0146] The vision interactive all-in-one machine provided by the present invention will be described below. The vision interactive all-in-one machine described below can be mutually referred to the motion capture-based all-in-one machine interaction method described above. Please refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of the vision interactive all-in-one machine provided by an embodiment of the present invention. As Figure 4 shown, a fixedly installed camera and a chassis are installed in the vision interactive all-in-one machine. The vision interactive all-in-one machine includes an acquisition unit, a feature fusion unit, an image processing unit, a target recognition unit, and a display interaction unit;

[0147] The acquisition unit 410 is configured to respond to an online teaching request of a teaching user, and acquire a teaching video of the teaching user during the teaching process and the teaching action image corresponding to each frame in the teaching video based on the camera in the vision all-in-one machine;

[0148] The feature fusion unit 420 is configured to perform multi-scale feature fusion on the target features of the teaching action image by using a feature pyramid network to obtain feature maps of multiple levels. The feature maps of multiple levels include a first feature map of the target level and second feature maps of other levels;

[0149] The image processing unit 430 is configured to perform sparsification processing on the first feature map by using a multi-scale dilated convolution mask to obtain a sparse feature map;

[0150] The target recognition unit 440 is configured to perform target recognition on the second feature map and the sparse feature map to obtain the user dynamic image corresponding to the teaching action image;

[0151] The display interaction unit 450 is configured to display the user dynamic image and the teaching video on the vision interactive all-in-one machine, wherein the user dynamic image moves rhythmically following the playback of the teaching video.

[0152] The visual interactive all-in-one machine provided by the present invention can display the dynamic image of the user and the teaching video on the visual interactive all-in-one machine at the same time, and the dynamic image of the user moves rhythmically following the playback of the teaching video, enhancing the interaction between the teaching user and the students, increasing the teaching interest, improving the learning enthusiasm of the students, and thus improving the teaching effect.

[0153] Refer to Figure 5 , Figure 5 which is a schematic structural diagram of the electronic device provided by the present invention. An embodiment of the present invention provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored on the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the following steps are implemented:

[0154] Respond to the online teaching request of the teaching user, and based on the camera in the visual all-in-one machine, obtain the teaching video of the teaching user during the teaching process and the teaching action image corresponding to each frame in the teaching video;

[0155] Use a feature pyramid network to perform multi-scale feature fusion on the target features of the teaching action image to obtain feature maps at multiple levels, and the feature maps at multiple levels include a first feature map at the target level and second feature maps at other levels;

[0156] Use a multi-scale dilated convolution mask to sparsify the first feature map to obtain a sparse feature map;

[0157] Perform target recognition on the second feature map and the sparse feature map to obtain the dynamic image of the user corresponding to the teaching action image;

[0158] Display the dynamic image of the user and the teaching video on the visual interactive all-in-one machine, where the dynamic image of the user moves rhythmically following the playback of the teaching video.

[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that makes a contribution to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An all-in-one machine interaction method based on motion capture, characterized in that: Applied to a visual interactive all-in-one machine, the visual interactive all-in-one machine is equipped with a fixed camera, and the all-in-one machine interaction method based on motion capture includes: In response to an online teaching request from a teaching user, obtaining a teaching video of the teaching user during the teaching process and a teaching action image corresponding to each frame in the teaching video based on a camera in the visual interactive all-in-one machine; Using a feature pyramid network to perform multi-scale feature fusion on the target features of the teaching action image to obtain feature maps at multiple levels, wherein the feature maps at multiple levels include a first feature map at the target level and second feature maps at other levels; Using a multi-scale atrous convolution mask to perform sparse processing on the first feature map to obtain a sparse feature map; Performing target recognition on the second feature map and the sparse feature map to obtain a user dynamic image corresponding to the teaching action image; Displaying the user dynamic image and the teaching video on the visual interactive all-in-one machine, wherein the user dynamic image moves rhythmically following the playing of the teaching video; The visual interactive all-in-one machine further comprises a chassis, and the all-in-one machine interaction method based on motion capture further comprises: Controlling the visual interactive all-in-one machine to shoot the calibration object through the camera to obtain a first image set; the first image set includes images of the calibration object at different shooting distances and / or different shooting angles; Calibrate the pitch angle of the camera based on the first image set, and determine a calibration result of the pitch angle of the camera; Controlling the visual interactive integrated machine to move to the front of the calibration object; Controlling the chassis to rotate, and photographing the calibration object through the camera to obtain a second image set; the second image set includes images of the calibration object photographed at different rotation angles; Based on the second image set and the rotation angle information of the chassis, calibrate the relative position of the camera and the chassis to determine the position calibration result of the camera; The rotation angle information includes the rotation angle corresponding to each of the calibration object images in the second image set; calibrating the relative position of the camera and the chassis based on the second image set and the rotation angle information of the chassis, and determining the position calibration result of the camera, including: Acquire camera internal reference data and calibration object data of the camera; Based on the camera internal parameter data, the calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set, calibrate the relative position of the camera and the chassis to determine a plurality of position deviation values; Based on the multiple position deviation values, a weighted aggregation calculation is performed to obtain a position calibration result of the camera; The method of calibrating the relative position of the camera and the chassis based on the camera internal parameter data, the calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set to determine a plurality of position deviation values ​​includes: Select any two calibration object images from the second image set as the first calibration object image and the second calibration object image respectively; Based on the camera intrinsic parameter data and the calibration object data, respectively determine a first position vector and a second position vector; the first position vector is a position vector between the calibration object and the camera in the first calibration object image, and the second position vector is a position vector between the calibration object and the camera in the second calibration object image; Determine a distance between a first position and a second position based on the first position vector and the second position vector; the first position is a position of the camera when the camera takes the first calibration object image, and the second position is a position of the camera when the camera takes the second calibration object image; Based on the planar geometric relationship between the first position, the second position and the center point of the chassis, the rotation angle corresponding to the first calibration object image and the rotation angle corresponding to the second calibration object image, the relative position of the camera and the chassis is calibrated to calculate the position deviation value; Returning to the step of selecting any two calibration object images from the second image set until a plurality of position deviation values ​​are calculated; The step of performing weighted aggregation calculation based on the plurality of position deviation values ​​to obtain the position calibration result of the camera includes: Selecting one of the position deviation values ​​as reference data, and using the remaining position deviation values ​​as comparison data; respectively determining the distance difference between each of the comparison data and the reference data; Comparing each of the distance differences with a first preset threshold value to obtain a plurality of comparison results; determining an error count of the reference data based on each of the comparison results; Returning to the step of selecting one of the position deviation values ​​as reference data and using the remaining position deviation values ​​as comparison data, until error counts corresponding to all the position deviation values ​​are obtained; Based on the error count corresponding to each of the position deviation values, abnormal data is eliminated to obtain a plurality of normal position deviation values; Based on the multiple normal position deviation values, a weighted aggregation calculation is performed to obtain a position calibration result of the camera; The step of calibrating the pitch angle of the camera based on the first image set and determining the pitch angle calibration result of the camera includes: Acquire camera external parameter data of the camera, camera internal parameter data of the camera, and calibration object data; Calibrate the pitch angle of the camera based on the camera external parameter data, the camera internal parameter data, the calibration object data and the first image set to determine a plurality of pitch angles; Performing weighted aggregation calculation based on the plurality of pitch angles to obtain a pitch angle calibration result of the camera; The step of performing weighted aggregation calculation based on the plurality of pitch angles to obtain a pitch angle calibration result of the camera includes: Selecting one of the pitch angles as reference data, and using the remaining pitch angles as comparison data; respectively determining the angle difference between each of the comparison data and the reference data; Comparing each of the angle differences with a second preset threshold value to obtain a plurality of comparison results; determining an error count of the reference data based on each of the comparison results; Returning to the step of selecting one of the pitch angles as reference data and using the remaining pitch angles as comparison data, until error counts corresponding to all the pitch angles are obtained; Eliminate abnormal data based on the error counts corresponding to the pitch angles to obtain a plurality of normal pitch angles; Based on the multiple normal pitch angles, a weighted aggregation calculation is performed to obtain a pitch angle calibration result of the camera.

2. The all-in-one machine interaction method based on motion capture according to claim 1, characterized in that: The multi-scale hole convolution mask includes: a first convolution layer and a first output layer; the multi-scale hole convolution mask is used to perform sparse processing on the first feature map to obtain a sparse feature map, including: Processing the first feature map using the first convolutional layer to obtain a sparse mask vector; The first output layer is used to perform sparse processing on the first feature map based on the sparse mask vector to obtain the sparse feature map.

3. The all-in-one machine interaction method based on motion capture according to claim 2, characterized in that: The first convolution layer includes: a plurality of dilated convolution kernels with different dilation rates, at least one channel compression convolution layer, and a second output layer; the first feature map is processed by the first convolution layer to obtain a sparse mask vector, including: Extracting initial information of different scales from the first feature map using the plurality of dilated convolution kernels with different dilation rates; Using the at least one channel compression convolution layer to perform channel compression processing on the initial information of different scales to obtain target information; The second output layer is used to determine the sparse mask vector based on the target information.

4. A visual interactive all-in-one machine, characterized in that: The visual interactive all-in-one machine is equipped with a fixed camera and a chassis, and the visual interactive all-in-one machine includes an acquisition unit, a feature fusion unit, an image processing unit, a target recognition unit and a display interaction unit; The acquisition unit is used to respond to the online teaching request of the teaching user, and acquire the teaching video of the teaching user in the teaching process and the teaching action image corresponding to each frame in the teaching video based on the camera in the visual interactive all-in-one machine; The feature fusion unit is used to perform multi-scale feature fusion on the target features of the teaching action image using a feature pyramid network to obtain feature maps of multiple levels, wherein the feature maps of the multiple levels include a first feature map of the target level and second feature maps of other levels; The image processing unit is used to perform sparse processing on the first feature map by using a multi-scale hole convolution mask to obtain a sparse feature map; The target recognition unit is used to perform target recognition on the second feature map and the sparse feature map to obtain a user dynamic image corresponding to the teaching action image; The display interaction unit is used to display the user dynamic image and the teaching video on the visual interaction all-in-one machine, wherein the user dynamic image moves rhythmically following the playing of the teaching video; The visual interactive all-in-one machine further comprises a chassis, and the all-in-one machine interaction method based on motion capture further comprises: Controlling the visual interactive all-in-one machine to shoot the calibration object through the camera to obtain a first image set; the first image set includes images of the calibration object at different shooting distances and / or different shooting angles; Calibrate the pitch angle of the camera based on the first image set, and determine a calibration result of the pitch angle of the camera; Controlling the visual interactive integrated machine to move to the front of the calibration object; Controlling the chassis to rotate, and photographing the calibration object through the camera to obtain a second image set; the second image set includes images of the calibration object photographed at different rotation angles; Based on the second image set and the rotation angle information of the chassis, calibrate the relative position of the camera and the chassis to determine the position calibration result of the camera; The rotation angle information includes the rotation angle corresponding to each of the calibration object images in the second image set; calibrating the relative position of the camera and the chassis based on the second image set and the rotation angle information of the chassis, and determining the position calibration result of the camera, including: Acquire camera internal reference data and calibration object data of the camera; Based on the camera internal parameter data, the calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set, calibrate the relative position of the camera and the chassis to determine a plurality of position deviation values; Based on the multiple position deviation values, a weighted aggregation calculation is performed to obtain a position calibration result of the camera; The method of calibrating the relative position of the camera and the chassis based on the camera internal parameter data, the calibration object data, each calibration object image in the second image set, and the rotation angle corresponding to each calibration object image in the second image set to determine a plurality of position deviation values ​​includes: Select any two calibration object images from the second image set as the first calibration object image and the second calibration object image respectively; Based on the camera intrinsic parameter data and the calibration object data, respectively determine a first position vector and a second position vector; the first position vector is a position vector between the calibration object and the camera in the first calibration object image, and the second position vector is a position vector between the calibration object and the camera in the second calibration object image; Determine a distance between a first position and a second position based on the first position vector and the second position vector; the first position is a position of the camera when the camera takes the first calibration object image, and the second position is a position of the camera when the camera takes the second calibration object image; Based on the planar geometric relationship between the first position, the second position and the center point of the chassis, the rotation angle corresponding to the first calibration object image and the rotation angle corresponding to the second calibration object image, the relative position of the camera and the chassis is calibrated to calculate the position deviation value; Returning to the step of selecting any two calibration object images from the second image set until a plurality of position deviation values ​​are calculated; The step of performing weighted aggregation calculation based on the plurality of position deviation values ​​to obtain the position calibration result of the camera includes: Selecting one of the position deviation values ​​as reference data, and using the remaining position deviation values ​​as comparison data; respectively determining the distance difference between each of the comparison data and the reference data; Comparing each of the distance differences with a first preset threshold value to obtain a plurality of comparison results; determining an error count of the reference data based on each of the comparison results; Returning to the step of selecting one of the position deviation values ​​as reference data and using the remaining position deviation values ​​as comparison data, until error counts corresponding to all the position deviation values ​​are obtained; Based on the error count corresponding to each of the position deviation values, abnormal data is eliminated to obtain a plurality of normal position deviation values; Based on the multiple normal position deviation values, a weighted aggregation calculation is performed to obtain a position calibration result of the camera; The step of calibrating the pitch angle of the camera based on the first image set and determining the pitch angle calibration result of the camera includes: Acquire camera external parameter data of the camera, camera internal parameter data of the camera, and calibration object data; Calibrate the pitch angle of the camera based on the camera external parameter data, the camera internal parameter data, the calibration object data and the first image set to determine a plurality of pitch angles; Performing weighted aggregation calculation based on the plurality of pitch angles to obtain a pitch angle calibration result of the camera; The step of performing weighted aggregation calculation based on the plurality of pitch angles to obtain a pitch angle calibration result of the camera includes: Selecting one of the pitch angles as reference data, and using the remaining pitch angles as comparison data; respectively determining the angle difference between each of the comparison data and the reference data; Comparing each of the angle differences with a second preset threshold value to obtain a plurality of comparison results; determining an error count of the reference data based on each of the comparison results; Returning to the step of selecting one of the pitch angles as reference data and using the remaining pitch angles as comparison data, until error counts corresponding to all the pitch angles are obtained; Eliminate abnormal data based on the error counts corresponding to the pitch angles to obtain a plurality of normal pitch angles; Based on the multiple normal pitch angles, a weighted aggregation calculation is performed to obtain a pitch angle calibration result of the camera.

Citation Information

Patent Citations

  • Embedded breast ultrasonic image recognition method

    CN114842238A

  • Method for constructing interpretable deep network based on dynamic reasoning decision and information bottleneck

    CN115204358A

  • Auxiliary teaching method, device and equipment and storage medium

    CN116805458A

  • System and method for producing 3D animation based on motioncapture

    KR101757765B1