Motion interaction method, device and storage medium based on dual-end video stream
By using the picture capture and repair technology of dual-end video streams on mobile terminals, the problem of device and scene restrictions is solved, cross-platform motion interaction and analysis are realized, and the user's motion experience is improved.
Patent Information
- Application Number
- CN202310164909.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-02-16
AI Technical Summary
The existing sports and fitness equipment and platforms have equipment restrictions, scenario restrictions and platform restrictions, resulting in poor user exercise experience and the inability to flexibly use and take into account multi-platform functions.
By using a dual-ended video stream on the mobile terminal, the user's actions and reference videos are captured using the first and second camera devices, screen screening, repairing and motion trajectory comparison are performed, and interactive results are provided to assist the user's movement training.
It improves the flexibility and convenience of sports interaction. Users can use it in multiple scenarios, taking into account the functions and courses of different platforms, providing quantitative sports analysis results, and improving the sports experience effect.
Smart Images

Figure CN116311507B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video detection technology, and in particular to a method, device, electronic device and computer-readable storage medium for motion interaction based on dual-end video streams. Background Art
[0002] At present, the popularity of online sports is increasing year by year. In addition to traditional pre-made video courses, various manufacturers have also launched online live courses and corresponding AIOT equipment, such as fitness mirrors, and holographic fitness equipment based on artificial intelligence. Based on these devices, on the basis of recorded courses and live courses, basic judgments and comments on user movements can be added through relevant equipment, thereby helping users complete sports training, learning and review.
[0003] However, existing sports products in the industry still have obvious shortcomings and deficiencies. For example, there are equipment restrictions. Users must use equipment from fixed manufacturers to learn related sports courses, such as fitness mirrors, and the scope of application of the equipment is limited. In addition, there are scene restrictions. Users must rely on the fixed scene under the installed equipment to perform subsequent operations, resulting in poor flexibility and high dependence on equipment, which is not suitable for home scenes. There are also platform restrictions. Users can only use the platform under the current manufacturer and cannot take into account the functions and functions and courses of other platforms. The above defects will make the user's sports experience poor. Summary of the Invention
[0004] The present invention provides a motion interaction method, device, electronic device and computer-readable storage medium based on dual-end video streams, the main purpose of which is to improve the flexibility and convenience of motion interaction based on dual-end video streams.
[0005] To achieve the above objectives, the present invention provides a motion interaction method based on dual-end video streams, comprising:
[0006] A first camera device based on a mobile terminal captures a user action picture, and a second camera device based on the mobile terminal captures an action reference video; the first camera device and the second camera device respectively capture pictures on the front and back symmetrical sides of the mobile terminal;
[0007] Acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency;
[0008] Based on a preset rule, the first frame action picture set and the second frame action picture set are respectively screened to determine a corresponding target distorted picture set;
[0009] Performing image restoration processing on the target distorted image set to obtain corresponding restored images, and updating the first frame action image set and the second frame action image set based on the restored images;
[0010] The updated first action frame set and the second action frame set are compared in terms of their motion trajectories, and the motion trajectory comparison result is fed back to the user as an interaction result.
[0011] In addition, an optional technical solution is that the first frame action picture set and the second frame action picture set are screened based on a preset rule to determine the corresponding target distorted picture set, including:
[0012] Determining whether the size, resolution, pixel, hue, and color palette of the first frame of action picture set meet preset requirements; and
[0013] Determining whether the size, resolution, pixel, hue, and color palette of the second frame of action picture set meet the preset requirements;
[0014] When at least one of the size, resolution, pixel, hue, and color palette of a picture does not meet the preset requirement, the picture is determined to be a distorted picture, and the target distorted picture set is determined based on all distorted pictures.
[0015] In addition, an optional technical solution is that performing image restoration processing on the target distorted image set to obtain corresponding restored images includes:
[0016] The image in the target distorted image is repaired based on a preset image repair model; the preset process of the image repair model includes:
[0017] Obtaining sample data, and obtaining corresponding fuzzy data based on the sample data;
[0018] forming training data based on the fuzzy data, the clear data corresponding to the fuzzy data, and the sample data;
[0019] The generative adversarial network model constructed based on the training data is trained until the generative adversarial network model converges within a preset range to form the image restoration model.
[0020] In addition, an optional technical solution is to obtain a distortion category of each picture in the target distorted picture; the distortion category includes at least one of picture size, resolution, pixel, hue, and color palette not meeting the preset requirements;
[0021] The images are repaired accordingly based on the distortion category; the repairing includes lossless amplification, stretching, clarity enhancement, color enhancement, and contrast enhancement of the images.
[0022] In addition, an optional technical solution is that the comparing the motion trajectories of the updated first action frame set and the second action frame set includes:
[0023] Acquire a first action sequence of the updated first action frame set at a preset frequency, and a second action sequence of the updated second action frame set at the preset frequency;
[0024] Acquire the coordinates of the first human body posture key points of the first action sequence within a preset time period, and the coordinates of the second human body posture key points of the second action sequence within the preset time period;
[0025] Obtaining first angle information of the coordinates of the first human posture key point between preset key points, and second angle information of the coordinates of the second human posture key point between preset key points;
[0026] The motion trajectory comparison result is determined based on the similarity between the first angle information and the second angle information.
[0027] In addition, an optional technical solution is that the coordinates of the first human body posture key point are the average value of the coordinate values of the human body posture key points at the same position in the first action sequence within a preset time period; the coordinates of the second human body posture key point are the average value of the coordinate values of the human body posture key points at the same position in the second action sequence within the preset time period.
[0028] In addition, an optional technical solution is that the formula for obtaining the similarity is:
[0029]
[0030] Where S represents the similarity, W represents the weight threshold corresponding to each angle, W=[w1,w2,…,w n ], M represents the second angle information, M=[m1,m2,…,m n ], T represents the first angle information, T=[t1,t2,…,t n ], n represents the number of angles between key points of human posture.
[0031] In order to solve the above problems, the present invention further provides a motion interaction device based on dual-end video streams, the device comprising:
[0032] a picture capturing unit, configured to capture a user action picture using a first camera device of the mobile terminal, and to capture a reference action video using a second camera device of the mobile terminal; the first camera device and the second camera device respectively capture pictures of the front and back symmetrical sides of the mobile terminal;
[0033] A frame action acquisition unit, configured to acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency;
[0034] a distortion determination unit, configured to screen the first action frame set and the second action frame set based on a preset rule to determine a corresponding target distorted picture set;
[0035] a distortion restoration unit, configured to perform image restoration processing on the target distorted image set, obtain corresponding restoration images, and update the first frame action image set and the second frame action image set based on the restoration images;
[0036] The result feedback unit is used to compare the motion trajectories of the updated first frame action picture set and the second frame action picture set, and feed back the motion trajectory comparison result to the user as an interaction result.
[0037] In order to solve the above problem, the present invention further provides an electronic device, comprising:
[0038] a memory storing at least one instruction; and
[0039] The processor executes the instructions stored in the memory to implement the above-mentioned motion interaction method based on dual-end video streams.
[0040] In order to solve the above problems, the present invention also provides a computer-readable storage medium, which stores at least one instruction. The at least one instruction is executed by a processor in an electronic device to implement the above-mentioned motion interaction method based on dual-end video streams.
[0041] The embodiment of the present invention uses the first camera device and the second camera device of the mobile terminal to capture user action picture information and an action reference video respectively, and then screens the first frame action picture set of the user and the second frame action picture set of the action reference video based on preset rules to perform picture restoration processing on the target distorted picture, and then compares the action trajectories of the restored first frame action picture set and the second frame action picture set, and feeds back the interaction results between the user and the action reference video, which can assist the user in sports training and guidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1A schematic diagram of a flow chart of a motion interaction method based on dual-end video streams provided by one embodiment of the present invention;
[0043] Figure 2 A schematic diagram of the distribution of key points of human posture provided by one embodiment of the present invention;
[0044] Figure 3 A schematic diagram of angles of a frame image provided by an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a module of a motion interaction device based on dual-end video streams provided by an embodiment of the present invention;
[0046] Figure 5 A schematic diagram of the internal structure of an electronic device for implementing a motion interaction method based on dual-end video streams provided by an embodiment of the present invention;
[0047] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0048] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0049] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0050] In order to solve the problems existing in the prior art that users cannot timely understand or objectively know the standard degree of their own exercise movements in sports and fitness scenarios, and need to rely on professional equipment such as fitness mirrors and holographic fitness equipment based on artificial intelligence, which leads to limited applicable scenarios and poor user exercise experience, the present application provides an artificial intelligence-based dual-end video streaming sports interaction method, which can realize intelligent analysis of users' sports videos based on electronic devices such as mobile phones and tablets, and can provide quantitative analysis results to improve users' sports experience.
[0051] The present invention provides a motion interaction method based on dual-end video stream. Figure 1 FIG2 is a flow chart of a motion interaction method based on dual-end video streams provided by an embodiment of the present invention. The method can be executed by a device, which can be implemented by software and / or hardware.
[0052] In this embodiment, the motion interaction method based on dual-end video streams includes:
[0053] S100: A first camera device based on a mobile terminal captures a user action picture (or a user action video), and a second camera device based on the mobile terminal captures an action reference video; the first camera device and the second camera device respectively capture pictures on the front and back symmetrical sides of the mobile terminal.
[0054] Specifically, the above-mentioned mobile terminal can be a user's mobile phone, tablet or other electronic device with front and rear camera functions, and the first camera device and the second camera device can respectively capture the images on the front and back symmetrical sides of the mobile terminal. After the user independently selects the action reference video content, the second camera device of the mobile terminal can be used to shoot the teacher's action posture in the action reference video. At the same time, the first camera device of the mobile terminal is used to obtain the user's action picture information showing the user following the teacher's work, that is, the user's action posture. Then, based on the information shot by the second camera device, the user's exercise standard level can be judged so as to make optimization suggestions, assist the user to understand his or her own exercise situation in a timely manner, complete exercise interaction, and thus carry out standard and effective exercise training.
[0055] S200: Acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency.
[0056] Specifically, according to a preset sampling frequency, a group of frame actions corresponding to the user action pictures can be obtained to form a first frame action picture set. Similarly, a second real action picture set corresponding to the action reference video can be obtained. Among them, when the human body or user in the teaching video temporarily does not appear in the corresponding video, or the action is still, such as going to drink water or answering a phone call, that is, when no human body is detected in the video, the corresponding video segment can be deleted, and only the content with the human body movement trajectory is retained to reduce the processing amount of video data.
[0057] S300: Screening the first action frame set and the second action frame set respectively based on a preset rule to determine a corresponding target distorted frame set.
[0058] The first frame action picture set and the second frame action picture set are respectively screened based on a preset rule to determine a corresponding target distorted picture set, including:
[0059] Determine whether the size, resolution, pixel, hue, and color palette of the first frame of the action picture set meet preset requirements; and
[0060] Determining whether the size, resolution, pixel, hue, and color palette of the second frame of the action picture set meet the preset requirements;
[0061] When at least one of the size, resolution, pixel, hue, and color palette of a picture does not meet the preset requirement, the picture is determined to be a distorted picture, and the target distorted picture set is determined based on all distorted pictures.
[0062] Specifically, the above preset rules are mainly used to judge the above parameters of the picture. When any one of the parameters does not meet the corresponding requirements, it means that the current picture is a distorted picture, and then it is included in the target distorted picture set, so that each distorted picture can be repaired later to improve the accuracy of motion interaction.
[0063] As a specific example, if the size of the currently judged picture is smaller than the preset size, the repair of the current picture can be to losslessly enlarge the picture to the preset size, or stretch and restore it, etc. If the resolution of the currently judged picture is lower than the preset threshold, the current picture can be enhanced in clarity so that the resolution of the repaired picture meets the corresponding preset requirements.
[0064] It can be seen that the above preset rules are not limited to size, resolution, pixels, hue, color palette, etc., and can also be flexibly set according to specific usage scenarios or needs.
[0065] S400: Performing image restoration processing on the target distorted image set to obtain corresponding restored images, and updating the first frame action image set and the second frame action image set based on the restored images.
[0066] The target distorted picture set is subjected to picture restoration processing to obtain corresponding restored pictures, and pictures in the target distorted picture can be restored based on a preset image restoration model.
[0067] Specifically, the preset process of the above image restoration model includes:
[0068] S410: Acquire sample data, and acquire corresponding fuzzy data based on the sample data;
[0069] S420: forming training data based on the fuzzy data, the clear data corresponding to the fuzzy data, and the sample data;
[0070] S430: Training the constructed generative adversarial network model based on the training data until the generative adversarial network model converges within a preset range to form an image restoration model.
[0071] During the training process of the image restoration model, sample data can be used to annotate parameter features such as image size and clarity to further improve the model's learning and restoration capabilities. The image restoration model can restore the target distorted image by performing lossless magnification, stretching, clarity enhancement, color enhancement, and contrast enhancement.
[0072] Specifically, when the target distorted picture set is used, its resolution, pixels, hue, and color palette are judged to see whether they meet the preset requirements. This is also to ensure that the key point information of the human body posture in each picture can be accurately determined in the future, facilitating the comparison of the two sets of action sequences.
[0073] In addition, it should be noted that when repairing the images in the target distorted image set, corresponding image processing software, such as an image processor, can also be used to complete it, but the requirements for mobile terminals are high, and the operations during the processing will be more cumbersome. However, the use of the above-mentioned graphic repair model can simplify the processing process, and multiple parameters of the image can be repaired simultaneously at one time, thereby improving the portability of mobile terminals.
[0074] S500: performing motion track comparison on the updated first action frame set and the second action frame set, and feeding back the motion track comparison result to the user as an interaction result.
[0075] The step of comparing the updated first action frame set and the updated second action frame set with respect to the action trajectory includes:
[0076] S510: Acquire a first action sequence of an updated first action frame set at a preset frequency, and a second action sequence of an updated second action frame set at a preset frequency;
[0077] S520: Acquire the coordinates of the first human body posture key point of the first action sequence within a preset time period, and the coordinates of the second human body posture key point of the second action sequence within a preset time period;
[0078] S530: Acquire first angle information of the coordinates of the first human posture key point between preset key points, and second angle information of the coordinates of the second human posture key point between preset key points;
[0079] S540: Determine the motion trajectory comparison result based on the similarity between the first angle information and the second angle information, and feed it back to the user.
[0080] Among them, the coordinates of the first human body posture key point can be set to the average value of the coordinate values of the human body posture key points at the same position in the first action sequence within a preset time period; the coordinates of the second human body posture key point can be set to the average value of the coordinate values of the human body posture key points at the same position in the second action sequence within a preset time period, the first key point information includes the coordinate angle information (first angle) of the first human body posture key point between the preset positions, and the second key point information includes the coordinate angle information (second angle) of the second human body posture key point between the preset positions.
[0081] As a specific example, the preset time period can be set to 1s. Assuming that the user action picture and the action reference video have 24 frames of frame action pictures within 1s, and the movement time is 1h, the first frame action picture set includes all frame action pictures of the user within 1h, and then the target distorted picture set of the first frame action picture set is determined, and the target distorted picture set is repaired by the image repair model, and the first frame action picture set is updated by the repaired picture, and then the first human body posture key point group of all the first pictures of the updated first frame action picture set within 1s is obtained, and these human body posture key point groups are averaged to obtain the first key point information; similarly, the second key point information is obtained, and then the first key point information and the second key point information within multiple consecutive 1s are compared to obtain the degree of overlap between the user's motion trajectory and the motion teaching trajectory.
[0082] It should be noted that in the above process of determining the key point information, the coordinates of all corresponding positions in all the human body posture key point groups within 1s can be averaged to determine a complete set of human body posture key point information, or a repaired action picture within 1s can be determined first, and the human body posture key point information can be determined based on the action picture.
[0083] To overcome the time difference between the two videos, the preset time may be set to be slightly larger, for example, 2s or 3s.
[0084] In other words, the first key point information within multiple consecutive 1s forms the key point sequence of the first action sequence, and the second key point information forms the key point sequence of the second action sequence. During the comparison process, the comparison work can be completed through the angle information between the key points of the corresponding sequences. If the difference is small, it indicates that the overlap between the two is high. The difference can also be scored by a specific numerical value, so that users can understand their own exercise status more intuitively.
[0085] As a specific example, performing similarity calculation on the first angle information and the second angle information to obtain the degree of overlap between the user's motion trajectory and the motion teaching trajectory may further include:
[0086] Setting corresponding preset areas on the first screen and the second screen;
[0087] Obtaining similarity between first angle information and second angle information at the same position in the first action sequence and the second action sequence within the preset area;
[0088] The degree of coincidence is determined based on the similarity.
[0089] It can be seen that this application does not impose specific restrictions on the method of obtaining the motion sequence. It can ultimately obtain this information so that the user's motion trajectory can be explained and compared with the motion teaching trajectory. As an example, the preset time range can be set to 1s-3s, etc., and then if the entire exercise is 40 minutes, the motion sequence is a group of motion sequences obtained every 1s-3s.
[0090] In addition, the key points of human posture can include: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, right ankle and other key points, which can be flexibly set according to the amplitude of the movement or the exercise part.
[0091] Furthermore, the above angles may include the angle of the line between the left wrist, left elbow and left shoulder, the angle of the line between the left hip joint, left knee and left ankle, the angle of the line between the left ear, left eye and nose, the angle of the line between the right wrist, right elbow and right shoulder, the angle of the line between the right hip joint, right knee and right ankle, the angle of the line between the right ear, right eye and nose, etc.
[0092] Specifically, Figure 2 shows a distribution diagram of key points of human posture according to an embodiment of the present invention, Figure 3 Two angle examples are shown.
[0093] like Figure 2 and Figure 3 As shown together, in this embodiment, 0-nose, 1-left eye, 2-right eye, 3-left ear, 4-right ear, 5-left shoulder, 6-right shoulder, 7-left elbow, 8-right elbow, 9-left wrist, 10-right wrist, 11-left hip joint, 12-right hip joint, 13-left knee, 14-right knee, 15-left ankle, 16-right ankle.
[0094] Taking the angle of the line connecting the left wrist, left elbow and left shoulder (9-7-5) as an example, given the pixel positions of 9, 7, and 5 in the corresponding frame image, first calculate the unit vector from 7 to 9 Then calculate the unit vector of 7 value 5 c and s represent vector parameters, and then calculate the angle θ between vector u and vector v counterclockwise. The specific calculation method is as follows:
[0095] Assuming that the angle between vector u and the positive direction of the x-axis is α, and the angle between vector v and the positive direction of the x-axis is β, then c1 = cosα, s1 = sinα, c2 = cosβ, s2 = sinβ, θ = β-α±(2kπ); further deduction yields:
[0096] cosθ=cos(β-α)
[0097] =cosβcosα+sinβsinα
[0098] =c1c2+s1s2
[0099] as well as,
[0100] sinθ=sin(β-α)
[0101] =sinβcosα-cosβsinα
[0102] =c1s2-c2s1
[0103] Among them, the coordinate axis takes 7 as the origin, the y-axis is parallel to the direction of the frame image height, assuming that the vector u is parallel to the y-axis, then for Figure 3 In Example 1, suppose cosα=0, sinα=1, then α=90°, Then β=315°, so θ=β-α=225°. For example 2, α=90°, Then β=45°, then θ=β-α=-45°<0°, so θ=360°-45°=315°.
[0104] It can be seen that if sinθ<0, then If sinθ≥0, then The present invention uses the counterclockwise angle between two vectors to better distinguish directionality compared to the angle between two vectors.
[0105] Then, the formula for obtaining similarity is:
[0106]
[0107] Where S represents the similarity, W represents the weight threshold corresponding to each angle, W=[w1,w2,…,w n ], M represents the second angle information, M=[m1,m2,…,m n ], T represents the first angle information, T=[t1,t2,…,t n ], n represents the number of angles between key points of human posture.
[0108] The value range of the similarity is [-1, 1]. Determining the overlap based on the similarity includes: when the similarity is less than 0, the overlap is 0; when the similarity is greater than 0, the overlap is 100*S, where S is the overall similarity.
[0109] In another specific embodiment of the present invention, in addition to considering the overall similarity, it is also possible to specifically consider only the local similarity of the frame action picture. By setting the interval, when the threshold is not within the set interval, the corresponding point can be determined to output a reminder. At this time, the score can be not output, and only the similarity of the first angle information and the second angle information within the set interval is calculated and output.
[0110] It should be noted that in the motion interaction method based on dual-end video streaming of the present invention, assuming that the first action sequence of the nth second is currently selected, the standard action frame image of the same action sequence in the action reference video is searched in order from the sequence. If the first frame action picture (1 / 24) is blurred or has a large gap with the standard action frame image of the teaching video, or the key point detection points of the human body posture cannot be fully detected, the second frame action picture (2 / 24) is automatically compared in sequence, and the cycle is continued until the frame action picture closest to the standard action frame image is finally determined. If it is still not recognized in the end, the local similarity method can be enabled to overcome the problem of image distortion.
[0111] After the results of the exercise interaction are determined, the product functions can display evaluation scores, screenshots of incorrect movements, playback videos and other functions to help users understand the details of the exercise process, thereby enhancing the user's memory and proficiency of the movements.
[0112] According to the motion interaction method based on dual-end video streaming of the above-mentioned present invention, it is possible to perform complete or partial frame image comparison between the user motion video and the action reference video, and then obtain the overlap between the user motion trajectory and the motion teaching trajectory, and finally determine the action standard degree of the user motion trajectory based on the overlap. The threshold of the standard degree can also be set. When the standard degree does not meet the corresponding threshold, the user can be given a color or sound reminder, which can effectively take into account the functions of each platform and the course. Through the method of dual video acquisition, it is possible to ensure real-time synchronization of standard actions and user actions, and conduct real-time digital evaluation to judge, score, evaluate user actions and make optimization suggestions.
[0113] like Figure 4 , which is a schematic diagram of the functional modules of the motion interaction device based on dual-end video streams of the present invention.
[0114] The dual-end video stream-based motion interaction device 100 of the present invention can be installed in an electronic device. Depending on the functionality implemented, the dual-end video stream-based motion interaction device can include a picture capture unit 101, a frame motion acquisition unit 102, a distortion determination unit 103, a distortion restoration unit 104, and a result feedback unit 105. The units described herein, also known as modules, refer to a series of computer program segments that can be executed by an electronic device processor and perform fixed functions, and are stored in the memory of the electronic device.
[0115] In this embodiment, the functions of each module / unit are as follows:
[0116] The image capture unit 101 is used to capture user action image information based on the first camera device of the mobile terminal, and to capture action reference video based on the second camera device of the mobile terminal; the first camera device and the second camera device respectively capture images on the front and back symmetrical sides of the mobile terminal.
[0117] The frame action acquisition unit 102 is configured to acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency;
[0118] The distortion determination unit 103 is configured to screen the first action frame set and the second action frame set based on a preset rule to determine a corresponding target distorted picture set;
[0119] The distortion restoration unit 104 is configured to perform image restoration processing on the target distorted image set, obtain corresponding restoration images, and update the first frame action image set and the second frame action image set based on the restoration images;
[0120] The result feedback unit 105 is configured to compare the motion trajectories of the updated first action frame set and the second action frame set, and to feed back the interaction result between the user and the action reference video.
[0121] Specifically, the above-mentioned mobile terminal can be a user's mobile phone, tablet or other electronic device with front and rear camera functions, and the first camera device and the second camera device can respectively capture the images on the front and back symmetrical sides of the mobile terminal. After the user independently selects the action reference video content, the second camera device of the mobile terminal can be used to shoot the teacher's action posture in the action reference video. At the same time, the first camera device of the mobile terminal is used to obtain the user's action picture information showing the user following the teacher's work, that is, the user's action posture. Then, based on the information shot by the second camera device, the user's exercise standard level can be judged so as to make optimization suggestions, assist the user to understand his or her own exercise situation in a timely manner, complete exercise interaction, and thus carry out standard and effective exercise training.
[0122] Among them, in the frame action acquisition unit 102, a group of frame actions corresponding to the user action picture can be obtained according to the preset sampling frequency to form a first frame action picture set. Similarly, a second real action picture set corresponding to the action reference video is obtained; among them, when the human body or user in the teaching video temporarily does not appear in the corresponding video, or the action is still, such as going to drink water or answering the phone, that is, when the presence of a human body is not detected in the video, the corresponding video segment can be deleted, and only the content with the human body movement trajectory is retained to reduce the processing amount of video data.
[0123] In the distortion determination unit 103, based on a preset rule, the first frame action picture set and the second frame action picture set are respectively screened to determine a corresponding target distorted picture set, including:
[0124] Determine whether the size, resolution, pixel, hue, and color palette of the first frame of the action picture set meet preset requirements; and
[0125] Determining whether the size, resolution, pixel, hue, and color palette of the second frame of the action picture set meet the preset requirements;
[0126] When at least one of the size, resolution, pixel, hue, and color palette of a picture does not meet the preset requirement, the picture is determined to be a distorted picture, and the target distorted picture set is determined based on all distorted pictures.
[0127] Specifically, the above preset rules are mainly used to judge the above parameters of the picture. When any one of the parameters does not meet the corresponding requirements, it means that the current picture is a distorted picture, and then it is included in the target distorted picture set, so that each distorted picture can be repaired later to improve the accuracy of motion interaction.
[0128] As a specific example, if the size of the currently judged picture is smaller than the preset size, the repair of the current picture can be to losslessly enlarge the picture to the preset size, or stretch and restore it, etc. If the resolution of the currently judged picture is lower than the preset threshold, the current picture can be enhanced in clarity so that the resolution of the repaired picture meets the corresponding preset requirements.
[0129] It can be seen that the above preset rules are not limited to size, resolution, pixels, hue, color palette, etc., and can also be flexibly set according to specific usage scenarios or needs.
[0130] In the distortion restoration unit 104, the target distorted picture set is subjected to image restoration processing to obtain the corresponding restoration picture. The pictures in the target distorted picture can be restored based on a preset image restoration model. Specifically, the preset process of the above-mentioned image restoration model includes:
[0131] 1. Obtain sample data and obtain corresponding fuzzy data based on the sample data;
[0132] 2. forming training data based on the fuzzy data, the clear data corresponding to the fuzzy data, and the sample data;
[0133] 3. Train the constructed generative adversarial network model based on the training data until the generative adversarial network model converges within a preset range to form an image restoration model.
[0134] During the training process of the image restoration model, sample data can be used to annotate parameter features such as image size and clarity to further improve the model's learning and restoration capabilities. The image restoration model can restore the target distorted image by performing lossless magnification, stretching, clarity enhancement, color enhancement, and contrast enhancement.
[0135] In addition, it should be noted that when repairing the images in the target distorted image set, corresponding image processing software, such as an image processor, can also be used to complete it. However, the use of the above-mentioned graphic repair model can simplify the processing process and can simultaneously repair multiple parameters of the image at one time, thereby improving the portability of mobile terminals.
[0136] The result feedback unit 105 is configured to compare the motion trajectories of the updated first action frame set and the second action frame set, and feed back the motion trajectory comparison result to the user as an interaction result.
[0137] The step of comparing the updated first action frame set and the updated second action frame set with respect to the action trajectory includes:
[0138] 1. Obtaining a first action sequence of an updated first action frame set at a preset frequency, and a second action sequence of an updated second action frame set at a preset frequency;
[0139] 2. Obtaining the coordinates of the first human posture key point of the first action sequence within a preset time period, and the coordinates of the second human posture key point of the second action sequence within a preset time period;
[0140] 3. Obtaining first angle information of the coordinates of the first human posture key point between the preset key points, and second angle information of the coordinates of the second human posture key point between the preset key points;
[0141] 4. Based on the similarity between the first angle information and the second angle information, determine the motion trajectory comparison result and feed it back to the user.
[0142] Specifically, the process of obtaining the first action sequence may include:
[0143] Acquire a first human body posture key point group of each first frame in the updated first frame action frame set within a preset time period;
[0144] Performing averaging processing on the first human posture key point group to obtain first key point information;
[0145] The process of obtaining the second action sequence may include:
[0146] Acquire a second human body posture key point group of each second frame in the updated second frame action picture set within the preset time period;
[0147] Performing averaging processing on the second human posture key point group to obtain second key point information;
[0148] Among them, the coordinates of the first human body posture key point can be set to the average value of the coordinate values of the human body posture key points at the same position in the first action sequence within a preset time period; the coordinates of the second human body posture key point can be set to the average value of the coordinate values of the human body posture key points at the same position in the second action sequence within a preset time period, the first key point information includes the coordinate angle information (first angle) of the first human body posture key point between the preset positions, and the second key point information includes the coordinate angle information (second angle) of the second human body posture key point between the preset positions.
[0149] As a specific example, the preset time period can be set to 1s. Assuming that the user action picture and the action reference video have 24 frames of frame action pictures within 1s, and the movement time is 1h, the first frame action picture set includes all frame action pictures of the user within 1h, and then the target distorted picture set of the first frame action picture set is determined, and the target distorted picture set is repaired by the image repair model, and the first frame action picture set is updated by the repaired picture, and then the first human body posture key point group of all the first pictures of the updated first frame action picture set within 1s is obtained, and these human body posture key point groups are averaged to obtain the first key point information; similarly, the second key point information is obtained, and then the first key point information and the second key point information within multiple consecutive 1s are compared to obtain the degree of overlap between the user's motion trajectory and the motion teaching trajectory.
[0150] It should be noted that in the above process of determining the key point information, the coordinates of all corresponding positions in all the human body posture key point groups within 1s can be averaged to determine a complete set of human body posture key point information, or a repaired action picture within 1s can be determined first, and the human body posture key point information can be determined based on the action picture.
[0151] To overcome the time difference between the two videos, the preset time may be set to be slightly larger, for example, 2s or 3s.
[0152] In other words, the first key point information within multiple consecutive 1s forms the key point sequence of the first action sequence, and the second key point information forms the key point sequence of the second action sequence. During the comparison process, the comparison work can be completed through the angle information between the key points of the corresponding sequences. If the difference is small, it indicates that the overlap between the two is high. The difference can also be scored by a specific numerical value, so that users can understand their own exercise status more intuitively.
[0153] As a specific example, performing similarity calculation on the first angle information and the second angle information to obtain the degree of overlap between the user's motion trajectory and the motion teaching trajectory may further include:
[0154] Setting corresponding preset areas on the first screen and the second screen;
[0155] Obtaining similarity between first angle information and second angle information at the same position in the first action sequence and the second action sequence within the preset area;
[0156] The degree of coincidence is determined based on the similarity.
[0157] It can be seen that this application does not impose specific restrictions on the method of obtaining the motion sequence. It can ultimately obtain this information so that the user's motion trajectory can be explained and compared with the motion teaching trajectory. As an example, the preset time range can be set to 1s-3s, etc., and then if the entire exercise is 40 minutes, the motion sequence is a group of motion sequences obtained every 1s-3s.
[0158] In addition, the key points of human posture can include: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, right ankle and other key points, which can be flexibly set according to the amplitude of the movement or the exercise part.
[0159] Furthermore, the above angles may include the angle of the line between the left wrist, left elbow and left shoulder, the angle of the line between the left hip joint, left knee and left ankle, the angle of the line between the left ear, left eye and nose, the angle of the line between the right wrist, right elbow and right shoulder, the angle of the line between the right hip joint, right knee and right ankle, the angle of the line between the right ear, right eye and nose, etc.
[0160] Specifically, Figure 2 shows a distribution diagram of key points of human posture according to an embodiment of the present invention, Figure 3 Two angle examples are shown.
[0161] like Figure 2 and Figure 3 As shown together, in this embodiment, 0-nose, 1-left eye, 2-right eye, 3-left ear, 4-right ear, 5-left shoulder, 6-right shoulder, 7-left elbow, 8-right elbow, 9-left wrist, 10-right wrist, 11-left hip joint, 12-right hip joint, 13-left knee, 14-right knee, 15-left ankle, 16-right ankle.
[0162] Taking the angle of the line connecting the left wrist, left elbow and left shoulder (9-7-5) as an example, given the pixel positions of 9, 7, and 5 in the corresponding frame image, first calculate the unit vector from 7 to 9 Then calculate the unit vector of 7 value 5 c and s represent vector parameters, and then calculate the angle θ between vector u and vector v counterclockwise. The specific calculation method is as follows:
[0163] Assuming that the angle between vector u and the positive direction of the x-axis is α, and the angle between vector v and the positive direction of the x-axis is β, then c1 = cosα, s1 = sinα, c2 = cosβ, s2 = sinβ, θ = β-α±(2kπ); further deduction yields:
[0164] cosθ=cos(β-α)
[0165] =cosβcosα+sinβsinα
[0166] =c1c2+s1s2
[0167] as well as,
[0168] sinθ=sin(β-α)
[0169] =sinβcosα-cosβsinα
[0170] =c1s2-c2s1
[0171] Among them, the coordinate axis takes 7 as the origin, the y-axis is parallel to the direction of the frame image height, assuming that the vector u is parallel to the y-axis, then for Figure 3 In Example 1, suppose cosα=0, sinα=1, then α=90°, Then β=315°, so θ=β-α=225°. For example 2, α=90°, Then β=45°, then θ=β-α=-45°<0°, so θ=360°-45°=315°
[0172] It can be seen that if sinθ<0, then If sinθ≥0, then The present invention uses the counterclockwise angle between two vectors to better distinguish directionality compared to the angle between two vectors.
[0173] Then, the formula for obtaining similarity is:
[0174]
[0175] Where S represents the similarity, W represents the weight threshold corresponding to each angle, W=[w1,w2,…,w n ], M represents the second angle information, M=[m1,m2,…,m n ], T represents the first angle information, T=[t1,t2,…,t n ], n represents the number of angles between key points of human posture.
[0176] The value range of the similarity is [-1, 1]. Determining the overlap based on the similarity includes: when the similarity is less than 0, the overlap is 0; when the similarity is greater than 0, the overlap is 100*S, where S is the overall similarity.
[0177] In another specific embodiment of the present invention, in addition to considering the overall similarity, it is also possible to specifically consider only the local similarity of the frame action picture. By setting the interval, when the threshold is not within the set interval, the corresponding point can be determined to output a reminder. At this time, the score can be not output, and only the similarity of the first angle information and the second angle information within the set interval is calculated and output.
[0178] According to the motion interaction method based on dual-end video streaming of the above-mentioned present invention, it is possible to perform complete or partial frame comparison of the user motion video and the action reference video, and repair the distorted picture, thereby improving the accuracy of the motion interaction between the user motion trajectory and the motion teaching trajectory, and finally determine the action standard degree of the user motion trajectory according to the degree of overlap, and also set a threshold for the standard degree. When the standard degree does not meet the corresponding threshold, the user can be given a color or sound reminder, which can effectively take into account the individual platform functions and courses, and ensure real-time synchronization of standard actions and user actions through dual video acquisition, and conduct real-time digital evaluation to judge, score, evaluate user actions and make optimization suggestions.
[0179] like Figure 5 , which is a structural diagram of an electronic device for implementing a motion interaction method based on dual-end video streams according to the present invention.
[0180] The electronic device 1 may include a processor 10, a memory 11 and a bus, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a motion interaction program 12 based on dual-end video streams.
[0181] The memory 11 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 1. Furthermore, the memory 11 may include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a motion interaction program based on a dual-end video stream, but also to temporarily store data that has been output or is to be output.
[0182] In some embodiments, the processor 10 may be composed of an integrated circuit, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines. It executes or runs programs or modules stored in the memory 11 (such as a motion interaction program based on dual-end video streams) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.
[0183] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11 and at least one processor 10, etc.
[0184] Figure 5 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 5 The structure shown does not constitute a limitation on the electronic device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0185] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering the various components. Preferably, the power source may be logically connected to the at least one processor 10 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may further include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0186] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0187] Optionally, the electronic device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0188] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0189] The motion interaction program 12 based on the dual-end video stream stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve the following:
[0190] A first camera device based on a mobile terminal captures a user action picture, and a second camera device based on the mobile terminal captures an action reference video; the first camera device and the second camera device respectively capture pictures on the front and back symmetrical sides of the mobile terminal;
[0191] Acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency;
[0192] Based on a preset rule, the first frame action picture set and the second frame action picture set are respectively screened to determine a corresponding target distorted picture set;
[0193] Performing image restoration processing on the target distorted image set to obtain corresponding restored images, and updating the first frame action image set and the second frame action image set based on the restored images;
[0194] The updated first action frame set and the second action frame set are compared in terms of their motion trajectories, and the motion trajectory comparison result is fed back to the user as an interaction result.
[0195] In addition, an optional technical solution is that the first frame action picture set and the second frame action picture set are screened based on a preset rule to determine the corresponding target distorted picture set, including:
[0196] Determining whether the size, resolution, pixel, hue, and color palette of the first frame of action picture set meet preset requirements; and
[0197] Determining whether the size, resolution, pixel, hue, and color palette of the second frame of action picture set meet the preset requirements;
[0198] When at least one of the size, resolution, pixel, hue, and color palette of a picture does not meet the preset requirement, the picture is determined to be a distorted picture, and the target distorted picture set is determined based on all distorted pictures.
[0199] In addition, an optional technical solution is that performing image restoration processing on the target distorted image set to obtain corresponding restored images includes:
[0200] The image in the target distorted image is repaired based on a preset image repair model; the preset process of the image repair model includes:
[0201] Obtaining sample data, and obtaining corresponding fuzzy data based on the sample data;
[0202] forming training data based on the fuzzy data, the clear data corresponding to the fuzzy data, and the sample data;
[0203] The generative adversarial network model constructed based on the training data is trained until the generative adversarial network model converges within a preset range to form the image restoration model.
[0204] In addition, the optional technical solution is,
[0205] Obtaining a distortion category of each picture in the target distorted picture; the distortion category includes at least one of picture size, resolution, pixel, hue, and color palette not meeting the preset requirements;
[0206] The images are repaired accordingly based on the distortion category; the repairing includes lossless amplification, stretching, clarity enhancement, color enhancement, and contrast enhancement of the images.
[0207] In addition, an optional technical solution is that the preset time range is 1s-3s.
[0208] In addition, an optional technical solution is to obtain a first action sequence of the updated first action frame set at a preset frequency, and a second action sequence of the updated second action frame set at the preset frequency;
[0209] Acquire the coordinates of the first human body posture key points of the first action sequence within a preset time period, and the coordinates of the second human body posture key points of the second action sequence within the preset time period;
[0210] Obtaining first angle information of the coordinates of the first human posture key point between preset key points, and second angle information of the coordinates of the second human posture key point between preset key points;
[0211] The motion trajectory comparison result is determined based on the similarity between the first angle information and the second angle information.
[0212] In addition, an optional technical solution is that the coordinates of the first human body posture key point are the average value of the coordinate values of the human body posture key points at the same position in the first action sequence within the preset time period; the coordinates of the second human body posture key point are the average value of the coordinate values of the human body posture key points at the same position in the second action sequence within the preset time period.
[0213] The formula for obtaining the similarity is:
[0214]
[0215] Where S represents the similarity, W represents the weight threshold corresponding to each angle, W=[w1,w2,…,w n ], M represents the second angle information, M=[m1,m2,…,m n ], T represents the first angle information, T=[t1,t2,…,t n ], n represents the number of angles between key points of human posture.
[0216] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0217] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0218] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and other division methods may be used in actual implementation.
[0219] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.
[0220] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0221] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0222] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0223] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. Second-order terms are used to indicate names and do not imply any particular order.
[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A motion interaction method based on dual-end video streams, characterized in that: include: A first camera device based on a mobile terminal captures a user action picture, and a second camera device based on the mobile terminal captures an action reference video; The first camera device and the second camera device respectively capture images of the front and back symmetrical sides of the mobile terminal; Acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency; Based on a preset rule, the first frame action picture set and the second frame action picture set are respectively screened to determine a corresponding target distorted picture set; Performing image restoration processing on the target distorted image set to obtain corresponding restored images, and updating the first frame action image set and the second frame action image set based on the restored images; Comparing the motion trajectories of the updated first action frame set and the second action frame set, and feeding back the motion trajectory comparison result to the user as an interaction result; The comparing the motion trajectories of the updated first action frame set and the second action frame set includes: Acquire a first action sequence of the updated first action frame set at a preset frequency, and a second action sequence of the updated second action frame set at the preset frequency; Acquire the coordinates of the first human body posture key points of the first action sequence within a preset time period, and the coordinates of the second human body posture key points of the second action sequence within the preset time period; Obtaining first angle information of the coordinates of the first human posture key point between preset key points, and second angle information of the coordinates of the second human posture key point between preset key points; The motion trajectory comparison result is determined based on the similarity between the first angle information and the second angle information.
2. The motion interaction method based on dual-end video stream according to claim 1, characterized in that: The step of screening the first action frame set and the second action frame set based on a preset rule to determine a corresponding target distorted picture set includes: Determining whether the size, resolution, pixel, hue, and color palette of the first frame of action picture set meet preset requirements; and Determining whether the size, resolution, pixel, hue, and color palette of the second frame of action picture set meet the preset requirements; When at least one of the size, resolution, pixel, hue, and color palette of a picture does not meet the preset requirement, the picture is determined to be a distorted picture, and the target distorted picture set is determined based on all distorted pictures.
3. The motion interaction method based on dual-end video stream according to claim 2, characterized in that: The performing image restoration processing on the target distorted image set to obtain corresponding restored images includes: The image in the target distorted image is repaired based on a preset image repair model; wherein the preset process of the image repair model includes: Obtaining sample data, and obtaining corresponding fuzzy data based on the sample data; forming training data based on the fuzzy data, the clear data corresponding to the fuzzy data, and the sample data; The generative adversarial network model constructed based on the training data is trained until the generative adversarial network model converges within a preset range to form the image restoration model.
4. The motion interaction method based on dual-end video stream according to claim 3, characterized in that: The repairing of the target distorted image based on a preset image repair model includes: Obtaining a distortion category of each picture in the target distorted picture; the distortion category includes at least one of picture size, resolution, pixel, hue, and color palette not meeting the preset requirements; The images are repaired accordingly based on the distortion category; the repairing includes lossless amplification, stretching, clarity enhancement, color enhancement, and contrast enhancement of the images.
5. The motion interaction method based on dual-end video stream according to claim 1, characterized in that: The first human posture key point coordinates are the average value of the human posture key point coordinates at all the same positions in the first action sequence within the preset time period; The second human posture key point coordinates are the average value of the human posture key point coordinate values at all the same positions in the second action sequence within the preset time period.
6. The motion interaction method based on dual-end video stream according to claim 1, characterized in that: The formula for obtaining the similarity is: Where S represents the similarity, W represents the weight threshold corresponding to each angle, W=[w1,w2,…,w n ], M represents the second angle information, M=[m1,m2,…,m n ], T represents the first angle information, T=[t1,t2,…,t n ], n represents the number of angles between key points of human posture.
7. A motion interaction device based on dual-end video stream, characterized in that: include: a picture capturing unit, configured to capture a user action picture based on a first camera device of the mobile terminal, and to capture an action reference video based on a second camera device of the mobile terminal; The first camera device and the second camera device respectively capture images of the front and back symmetrical sides of the mobile terminal; A frame action acquisition unit, configured to acquire a first frame action picture set of the user action picture and a second frame action picture set of the action reference video according to a preset sampling frequency; a distortion determination unit, configured to screen the first action frame set and the second action frame set based on a preset rule to determine a corresponding target distorted picture set; a distortion restoration unit, configured to perform image restoration processing on the target distorted image set, obtain corresponding restoration images, and update the first frame action image set and the second frame action image set based on the restoration images; A result feedback unit is used to compare the motion trajectories of the updated first action frame set and the second action frame set, and feed back the motion trajectory comparison result to the user as an interaction result; The comparing the motion trajectories of the updated first action frame set and the second action frame set includes: Acquire a first action sequence of the updated first action frame set at a preset frequency, and a second action sequence of the updated second action frame set at the preset frequency; Acquire the coordinates of the first human body posture key points of the first action sequence within a preset time period, and the coordinates of the second human body posture key points of the second action sequence within the preset time period; Obtaining first angle information of the coordinates of the first human posture key point between preset key points, and second angle information of the coordinates of the second human posture key point between preset key points; The motion trajectory comparison result is determined based on the similarity between the first angle information and the second angle information.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps in the motion interaction method based on dual-end video streams as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the motion interaction method based on dual-end video streams as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Interaction gesture motion trail partition method based on multiple rules
CN103413137A
Motion video detection method and device, equipment and storage medium
CN115578786A