Video processing method and related electronic equipment
By adjusting the position of AR elements in AR videos to adapt to the user's body proportions, the problem of AR elements not matching with users is solved, and learning effect and security are improved.
Patent Information
- Application Number
- CN202310125211.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-02-02
AI Technical Summary
The AR elements in AR videos do not match the user's body shape, resulting in poor user learning and movement effects.
The user's image is collected through the camera, the bone node coordinates and motion space radius are obtained, the target coordinates of the AR element are calculated, to adapt to the user's body proportions, and the position of the AR element in the video.
It improves the effect of users learning actions based on AR elements, so that the proportion of AR elements to the user's body in the video is more consistent, and improves learning experience and security.
Smart Images

Figure CN118433471B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the AR field, and in particular to a video processing method and related electronic equipment. Background Art
[0002] With the continuous development of AR technology, users can learn gymnastics, aerobics, etc. through AR videos to achieve fitness effects. Specifically, users can play action videos processed by AR technology on electronic devices (for example, large-screen TVs). The electronic device will display AR elements to guide the user to perform corresponding actions. The action is completed only when and only when the user's specific joints "touch" the specific AR elements in the AR video. In this way, the user's interest in the learning process can be enhanced.
[0003] However, since the body proportions of the moving objects in the video may not match those of the user, and the position of the AR elements in the video is based on the body proportions of the moving objects in the video, this may cause the position of the AR elements in the video to not match the body proportions of the user. Summary of the Invention
[0004] An embodiment of the present application provides a video processing method that solves the problem that AR elements in AR videos do not match the body proportions of users, which in turn leads to poor learning results of users when learning actions in AR videos based on AR elements.
[0005] In a first aspect, an embodiment of the present application provides a video processing method, which is applied to an electronic device including a display screen and a camera, the method including: capturing a first image of a user through a camera; obtaining a first skeletal node coordinate and a first motion space radius of the user in the first image; calculating the target coordinates of an AR element in the i-th frame image based on the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius, and the coordinates of the AR element in the i-th frame image of the first video; the i-th frame image is the next frame image to be played of the first video; and the i-th frame image is displayed on the display screen; wherein the coordinates of the AR element in the displayed i-th frame image are the target coordinates.
[0006] In the above embodiment, before playing each frame of the first video, the electronic device will capture the image of the user and obtain the user's first skeletal node coordinates and the first motion space radius. The adjusted coordinates of the AR element, i.e., the target coordinates of the AR element, are calculated based on the position of the AR element in the frame image in the first video, the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, and the second motion space radius. In this way, the position of the adjusted AR element in the first video can be made more consistent with the user's body proportions, thereby improving the user's learning effect of learning the actions in the first video based on the AR element.
[0007] In combination with the first aspect, in one possible implementation, before capturing the user's first image through the camera, the method further includes: processing the second video to obtain all N frames of video images in the second video, where the second video is a video that does not include AR elements; if there is a P frame video image with a moving object in the N frame video image, processing the P frame video image to obtain a skeleton node image of the P frame video image, wherein the skeleton node image displays the skeleton node of the moving object, and the skeleton node image includes the position information of the skeleton node of the moving object; in response to the first operation, setting AR elements for the M frames of video images in the P frame video image based on the position information of the skeleton node of the moving object, recording the coordinates of the AR elements, and obtaining the first video; and saving the first video. In this way, users can add AR elements to videos without AR elements, thereby realizing personalized customization of AR action videos, learning the actions in the video based on the AR elements, increasing learning interest, and improving learning effects.
[0008] In combination with the first aspect, in one possible implementation method, the first bone node coordinates are the left sacrum coordinates or the right sacrum coordinates of the user in the first image; the first motion space radius is the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's left finger joints, or the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's right finger joints.
[0009] In conjunction with the first aspect, in one possible implementation, the target coordinates of the AR element in the i-th frame image are calculated based on the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius, and the coordinates of the AR element in the i-th frame image of the first video. Specifically, the first ratio value is calculated; and the target coordinates of the AR element in the i-th frame image are calculated based on the first ratio value, the first skeletal node coordinates, and the second skeletal node coordinates. In this way, the ratio value of the body shape of the moving object in the first video to the body shape of the user can be calculated, and then, based on the ratio value, the coordinates of the AR element in the first video are mapped to the target coordinates, so that the position where the AR element appears in the first video is more adapted to the body shape of the user.
[0010] In combination with the first aspect, in a possible implementation, calculating the first scale value specifically includes: calculating the first scale value Scale according to the formula Scale=L / l, where l is the first motion space radius and L is the second motion space radius.
[0011] In combination with the first aspect, in one possible implementation, calculating the first scale value specifically includes: calculating the left arm length l1, the right arm length l2, the left lower limb length l3, and the right lower limb length l4 of the user in the first image; calculating the left arm length L1, the right arm length L2, the left lower limb length L3, and the right lower limb length L4 of the moving object in the i-th frame image; and then calculating the first scale value according to the formula Scale i =l i / L i , calculate Scale1, Scale2, Scale3 and Scale4 respectively; among Scale1, Scale2, Scale3 and Scale4, select the Scale with the smallest value i As the first scale value. In this way, since the scale with the smallest value i Indicates the joint with the smallest difference between the moving object and the user in the first video. i As the first ratio value, the target coordinates of the AR element calculated by the first ratio value are more adapted to the user's body shape. In combination with the first aspect, in a possible implementation, the target coordinates of the AR element in the i-th frame image are calculated according to the first bone node coordinates, the second bone node coordinates, the first motion space radius and the second motion space radius, specifically including: according to the formula x′ j =x j *Scale+(X_center-x_center*Scale) calculates the target horizontal coordinate of the jth AR element in the i-th frame image; according to the formula y′ j =y j *Scale+(Y_center-y_center*Scale) calculates the target vertical coordinate of the j-th AR element in the i-th frame image; where x′ j is the target horizontal coordinate of the jth AR element, y′ j is the target vertical coordinate of the jth AR element, Scale is the first scale value, x j is the horizontal coordinate of the jth AR element, y j is the ordinate of the jth AR element, x_center is the abscissa of the second skeletal node, y_center is the ordinate of the second skeletal node, x_center is the abscissa of the first skeletal node, and y_center is the ordinate of the first skeletal node. In this way, the coordinates of the AR element in the first video are mapped to the target coordinates.
[0012] In combination with the first aspect, in one possible implementation, the target coordinates of the AR element in the i-th frame image are calculated according to the first skeleton node coordinates, the second skeleton node coordinates, the first motion space radius and the second motion space radius, specifically including: according to the formula x′ j =x j *Scale+(X_center-x_center) calculates the target horizontal coordinate of the jth AR element in the i-th frame image; according to the formula y′ j =y j *Scale+(Y_center-y_center) calculates the target vertical coordinate of the jth AR element in the i-th frame image; where x′ j is the target horizontal coordinate of the jth AR element, y′ j is the target vertical coordinate of the jth AR element, Scale is the first scale value, x j is the horizontal coordinate of the jth AR element, y j is the ordinate of the jth AR element, x_center is the abscissa of the second skeletal node, y_center is the ordinate of the second skeletal node, x_center is the abscissa of the first skeletal node, and y_center is the ordinate of the first skeletal node. In this way, the coordinates of the AR element in the first video are mapped to the target coordinates.
[0013] In combination with the first aspect, in one possible implementation method, the second skeletal node coordinates are the left sacrum coordinates or the right sacrum coordinates of the moving object in the i-th frame image; the second motion space radius is the Euclidean distance between the neck bone center node coordinates of the moving object in the i-th frame image and the user's right finger joint coordinates, or the Euclidean distance between the neck bone center node coordinates of the moving object in the i-th frame image and the user's left finger joint coordinates, or a fixed distance value stored in the electronic device.
[0014] In a second aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors, a display screen, a camera, and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, the one or more processors calling the computer instructions to enable the electronic device to execute: capturing a first image of the user through the camera; obtaining the first skeletal node coordinates and the first motion space radius of the user in the first image; calculating the target coordinates of the AR element in the i-th frame image based on the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius, and the coordinates of the AR element in the i-th frame image of the first video; the i-th frame image is the next frame image to be played of the first video; the i-th frame image is displayed on the display screen; wherein the coordinates of the AR element in the displayed i-th frame image are the target coordinates.
[0015] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to enable the electronic device to execute: before capturing the first image of the user through the camera, it also includes: processing the second video to obtain all N frames of video images in the second video, and the second video is a video that does not include AR elements; if there is a P frame video image with a moving object in the N frame video image, the P frame video image is processed to obtain a skeleton node image of the P frame video image, the skeleton node image displays the skeleton node of the moving object, and the skeleton node image includes the position information of the skeleton node of the moving object; in response to the first operation, AR elements are set for the M frame video images in the P frame video image based on the position information of the skeleton node of the moving object, the coordinates of the AR elements are recorded, and the first video is saved.
[0016] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to enable the electronic device to execute: calculating the target coordinates of the AR element in the i-th frame image based on the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius and the coordinates of the AR element in the i-th frame image of the first video, specifically: calculating the first proportional value; calculating the target coordinates of the AR element in the i-th frame image based on the first proportional value, the first skeletal node coordinates, and the second skeletal node coordinates.
[0017] In combination with the second aspect, in one possible implementation method, the one or more processors call the computer instructions to enable the electronic device to execute: calculate the first proportional value, specifically including: calculate the first proportional value Scale according to the formula Scale = L / l, l is the first motion space radius, and L is the second motion space radius.
[0018] In conjunction with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: calculating a first scale value, specifically including: calculating the left arm length l1, the right arm length l2, the left lower limb length l3, and the right lower limb length l4 of the user in the first image; calculating the left arm length L1, the right arm length L2, the left lower limb length L3, and the right lower limb length L4 of the moving object in the i-th frame image; and then calculating the first scale value according to the formula Scale i =l i / L i , calculate Scale1, Scale2, Scale3 and Scale4 respectively; among Scale1, Scale2, Scale3 and Scale4, select the Scale with the smallest value i as the first scale value.
[0019] In conjunction with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: calculating the target coordinates of the AR element in the i-th frame image according to the first skeletal node coordinates, the second skeletal node coordinates, the first motion space radius, and the second motion space radius, specifically including: according to the formula x′ j =x j *Scale+(X_center-x_center*Scale) calculates the target horizontal coordinate of the jth AR element in the i-th frame image; according to the formula y′ j =y j *Scale+(Y_center-y_center*Scale) calculates the target vertical coordinate of the j-th AR element in the i-th frame image; where x′ j is the target horizontal coordinate of the jth AR element, y′ j is the target vertical coordinate of the jth AR element, Scale is the first scale value, x j is the horizontal coordinate of the jth AR element, y j is the ordinate of the jth AR element, X_center is the abscissa of the second bone node, Y_center is the ordinate of the second bone node, x_center is the abscissa of the first bone node, and y_center is the ordinate of the first bone node.
[0020] In conjunction with the second aspect, in one possible implementation, the one or more processors call the computer instructions to cause the electronic device to execute: calculating the target coordinates of the AR element in the i-th frame image according to the first skeletal node coordinates, the second skeletal node coordinates, the first motion space radius, and the second motion space radius, specifically including: according to the formula x′ j =x j *Scale+(X_center-x_center) calculates the target horizontal coordinate of the jth AR element in the i-th frame image; according to the formula y′ j =y j *Scale+(Y_center-y_center) calculates the target vertical coordinate of the jth AR element in the i-th frame image; where x′ j is the target horizontal coordinate of the jth AR element, y′ j is the target vertical coordinate of the jth AR element, Scale is the first scale value, x j is the horizontal coordinate of the jth AR element, y j is the ordinate of the jth AR element, X_center is the abscissa of the second bone node, Y_center is the ordinate of the second bone node, x_center is the abscissa of the first bone node, and y_center is the ordinate of the first bone node.
[0021] In a third aspect, an embodiment of the present application provides an electronic device comprising: a touch screen, a camera, one or more processors and one or more memories; the one or more processors are coupled to the touch screen, the camera, and the one or more memories, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method described in the first aspect or any possible implementation method of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method described in the first aspect or any possible implementation method of the first aspect.
[0023] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute the method described in the first aspect or any possible implementation of the first aspect.
[0024] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1A-1B This is an example diagram of an AR scene provided in an embodiment of the present application;
[0026] Figure 2 This is an example diagram of AR element mapping provided in an embodiment of the present application;
[0027] Figure 3 This is a flowchart of a method for determining the position of an AR element provided in an embodiment of the present application;
[0028] Figure 4A This is an example diagram of a human skeleton node provided in an embodiment of the present application;
[0029] Figure 4B This is an example diagram of an interface of an electronic device 100 provided in an embodiment of the present application;
[0030] Figure 4C This is an example diagram of a motion space radius provided in an embodiment of the present application.
[0031] Figure 4DThis is an example diagram of a user image and a video frame image provided by an embodiment of the present application;
[0032] Figure 4E This is an example diagram of an interface of another electronic device 100 provided in an embodiment of the present application;
[0033] Figure 5 is a flowchart of another video processing method provided by an embodiment of the present application;
[0034] Figure 6A is a single frame image in the second video provided in an embodiment of the present application;
[0035] Figure 6B is an exemplary skeleton node image provided in an embodiment of the present application;
[0036] Figure 7 1 is a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application;
[0037] Figure 8 It is a software structure block diagram of the electronic device 100 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Mentioning "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present embodiment application. The appearance of this phrase in various positions in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It can be understood explicitly and implicitly by those skilled in the art that the embodiments described herein can be combined with other embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.
[0039] In the specification, claims, and accompanying drawings of this application, the terms "first," "second," "third," and the like are used to distinguish different objects and are not used to describe a particular order. Furthermore, the terms "including," "comprising," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a list of steps or elements may be included, or alternatively, steps or elements not listed may be included, or other steps or elements may be included that are inherent to the process, method, product, or apparatus.
[0040] Only part relevant to the present application is shown in the accompanying drawings, not all of it. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processing or methods depicted as flow charts. Although flow charts describe various operations (or steps) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. When its operation is completed, the processing can be terminated, but can also have additional steps not included in the accompanying drawings. The processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0041] As used in this specification, the terms "component," "module," "system," "unit," and the like are used to refer to computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media having various data structures stored thereon. Units can communicate, for example, through local and / or remote processes based on signals having one or more data packets (e.g., data from a second unit interacting with another unit in a local system, a distributed system, and / or a network. For example, the Internet interacts with other systems via signals).
[0042] With the continuous development of AR technology, the application scenarios of AR technology are becoming more and more extensive. Figure 1A The following is an exemplary application scenario of AR technology provided by an embodiment of the present application. Figure 1A The figure includes an electronic device 100, and a gymnastics video is played on the electronic device 100. The screen includes athletes performing gymnastics and AR elements 101 to 104. In addition to the electronic device 100, the figure also includes a user 106. The user 106 performs corresponding gymnastics movements based on the screen of the electronic device 100, thereby achieving the purpose of doing gymnastics. Among them, AR elements 101 to 104 are used to guide the user to complete the gymnastics movements in the electronic device 100. For example, AR element 101 is used to instruct the user's right arm to reach the position of AR element 101 in the screen, AR element 102 is used to instruct the user's left arm to reach the position of AR element 102 in the screen, AR element 103 is used to instruct the user's left toe to reach the position of AR element 103 in the screen, and AR element 104 is used to instruct the user's right toe to reach the position of AR element 104 in the screen. By marking AR elements in the video screen of the electronic device 100, the user can be guided to perform the gymnastics movements in the video screen more standardly, thereby achieving better fitness results.
[0043] In other application scenarios, the electronic device 100 can use the camera 105 to capture user images, analyze the user's movements based on the user images, and then score the user's movements based on the standardization of the user's movements. The scores are displayed on the interface of the electronic device 100, thereby prompting the user to learn the standardization of the movements. For example, Figure 1B In the figure, AR elements 101 to 104 all display scores. Among them, the score displayed by AR element 101 is 90, indicating that the movement of the user's right arm is relatively standard; the score displayed by AR element 102 is 60, indicating that the movement of the user's left arm is not very standard; the score displayed by AR element 103 is 40, indicating that the movement of the user's left toe is not up to standard and needs to be strengthened; the score displayed by AR element 104 is 100, indicating that the movement of the user's right toe is very standard. In addition, the electronic device 100 can also prompt the user whether the whole set of movements is standard based on the total score of the user's movements displayed on the interface. Taking the percentage system as an example, in Figure 2 In the video, the electronic device 100 displays the user's score of 78 points on the video interface, and displays a prompt message "The action is not standard, need to strengthen practice", which is used to remind the user that the action this time does not meet the standard.
[0044] However, for different AR sports videos, the athletes in the videos may be different, because the AR elements in the videos are generated based on the body proportions of the athletes in the videos. When the body proportions of the user and the athletes in the AR sports video are inconsistent, when the user learns the movements of the athletes in the video based on the AR elements in the video, the user's arms and other joints may not reach the position where the AR elements appear in the video, thereby reducing the effect of the user learning the movements in the AR sports video. For example, Figure 2As shown, when AR elements 101 and 102 in the video screen of electronic device 100 are mapped onto the same plane as the user, the corresponding locations are location 1 and location 2, respectively. Due to the user's body proportions, even if the user changes position, their left and right hands cannot reach the locations corresponding to location 1 and location 2. Therefore, when electronic device 100 uses camera 105 to capture the user's image, it analyzes the user's movements in the image and determines that the user's movements do not meet the requirements. If the user forcibly stretches their body to meet the required movements, it may cause injury. Therefore, to address the problem of AR elements and users not being compatible in AR motion videos, embodiments of the present application provide a method for determining the position of an AR element. The method includes: the electronic device obtains the user's skeletal node coordinates and the user's motion space radius. The electronic device obtains preset skeletal node coordinates and the user's motion space radius. Then, the electronic device calculates a ratio value based on the user's motion space radius and the preset user space radius. The electronic device calculates the coordinates of the AR element to be displayed on the electronic device interface based on the ratio value, the user's skeletal node coordinates, the preset skeletal node coordinates, and the coordinates of the AR element in the AR video. Finally, the electronic device displays the AR elements in the AR video at the corresponding calculated coordinate positions based on the calculated coordinates. In this way, the position of the AR elements in the AR video can be made to better match the user's body proportions, so that the user can better complete the actions in the video based on the AR elements in the video. This solves the problem that in the process of users learning actions in the video in the AR scene, the position of the AR elements in the video does not match the user due to the large difference between the user's body proportions and the body proportions of the task in the video, resulting in a decrease in the effectiveness of the user's learning of the video actions.
[0045] The following is a detailed description of the process of a method for determining the position of an AR element provided by an embodiment of the present application, with reference to the accompanying drawings. Figure 3 , Figure 3 This is a flow chart of a method for determining the position of an AR element provided in an embodiment of the present application. Figure 3 In the embodiment, only the electronic device determines the position of the AR element in a single frame image in the first video (AR motion video) is described. In fact, before the electronic device plays each frame image of the first video, it can execute Figure 3 The steps in the flowchart. The specific process is as follows:
[0046] S301: The electronic device captures a first image of a user through a camera, and determines first skeletal node coordinates and a first motion space radius of the user based on the first image.
[0047] For example, the electronic device may be the above Figure 1A The electronic device 100 in FIG.
[0048] Specifically, before playing each frame of video image, the electronic device can capture an image of the user through a camera, which is the first image. After acquiring the first image, the electronic device can identify the user's skeletal nodes, thereby obtaining the coordinate information of the skeletal nodes in the first image. Specifically, the skeletal node can be the user's left sacrum (for example, Figure 4A ), or the user's right sacrum (e.g., Figure 4A ), or the center node of the user's neck bone (e.g., Figure 4A The skeleton node 403 in the image may also be other skeleton nodes, which is not limited in the present embodiment. The present embodiment takes the first skeleton node coordinate (x_center, y_center) as the coordinate of the left sacrum of the user in the first image as an example for explanation.
[0049] In addition, after collecting the first image, the electronic device can also determine the motion space radius of the user in the first image through the bone nodes. The motion space radius of the user can be the motion space radius of the user's upper arm, or the motion space radius of the user's lower limbs, or the space radius of the user's height. Exemplarily, the electronic device can determine the distance between the central node of the user's neck bone in the first image and the fingertips of the user's left arm as the motion space radius of the user's left arm, the electronic device can determine the distance between the central node of the user's neck bone in the first image and the fingertips of the user's right arm as the motion space radius of the user's right arm, the electronic device can determine the distance between the left sacrum and the left toe in the first image as the motion space radius of the user's left lower limb, and the electronic device can determine the distance between the right sacrum and the right toe in the first image as the motion space radius of the user's right lower limb. In this embodiment of the present application, the distance between the central node of the user's neck bone in the first image and the fingertips of the user's left arm is determined as the first motion space radius l of the user (e.g. Figure 4C ) as an example for explanation.
[0050] Optionally, before capturing the first image, the electronic device may prompt the user to perform corresponding actions according to the actions displayed on the electronic device screen. Then, the electronic device captures the first image of the user performing the action through the camera, so that the electronic device can accurately calculate the coordinates of the user's skeletal nodes and the radius of the motion space in the first image. For example, Figure 4B As shown, the electronic device 100 displays "Please show the following gestures" on the screen. Then, the electronic device 100 displays the following gestures on the screen. Figure 4B The actions shown are for user convenience.
[0051] Optionally, before the electronic device captures the first image of the user via the camera, it may obtain the next frame of image to be displayed in the first video, i.e., the i-th frame of image. By identifying the i-th frame of image, it is determined whether an AR element exists in the i-th frame of image. If an AR element exists, S301 is executed. Otherwise, S301 is not executed.
[0052] S302: The electronic device obtains the coordinates of the second skeletal node and the radius of the second motion space.
[0053] Specifically, the second skeletal node coordinates and the second motion space radius can be the skeletal node coordinates and the motion space radius of the moving object in the i-th frame image of the first video. The first video is an AR motion video, which can be a fitness video, a limb rehabilitation training video, a gymnastics video, or other videos with AR elements. The embodiment of the present application does not limit this. In the AR video, AR elements are included. For example, the AR elements can be the above-mentioned Figure 1A AR element 101 in. The second skeletal node coordinates correspond to the first skeletal node coordinates. For example, when the first skeletal node coordinates are the coordinates of the user's left sacrum in the first image, the second skeletal node coordinates are the coordinates of the left sacrum of the moving object in the i-th frame image of the first video. When the first motion space radius is the distance between the center node of the user's neck bone and the fingertip of the user's left arm in the first image, the second motion space radius is the distance between the center node of the moving object's neck bone and the fingertip of the moving object's left arm in the i-th frame image of the first video. In an embodiment of the present application, the second skeletal node coordinates obtained by the electronic device are (X_center, Y_center).
[0054] In one possible implementation, the second skeletal node coordinates and the second motion space radius may be fixed skeletal node data coordinates and motion space radius pre-stored by the electronic device. The present embodiment of the application does not limit the manner in which the electronic device obtains the second skeletal node coordinates and the second motion space radius.
[0055] It should be understood that S301 can be executed before S302, or after S302, or simultaneously with S301, and the embodiments of the present application do not limit this.
[0056] S303: The electronic device calculates a first ratio value based on the first movement space radius and the second movement space radius.
[0057] Specifically, after obtaining the first movement space radius and the second movement space radius, the electronic device may calculate a ratio of the second movement space radius to the first movement space radius, where the ratio is the first ratio Scale, that is, Scale=L / l.
[0058] In one possible implementation, the electronic device may calculate ratios between the joints of the user in the first image and the joints of the moving object in the i-th frame of the first video. For example, the ratio of the length of the left limb of the moving object in the i-th frame to the length of the left limb of the user in the first image, the ratio of the length of the right limb of the moving object in the i-th frame to the length of the right limb of the user in the first image, and so on. The electronic device may then select the smallest of these calculated ratios as the first ratio.
[0059] S304: The electronic device obtains the coordinates of each AR element in the i-th frame image of the first video.
[0060] Specifically, the coordinates of the AR element can be the coordinates of the center point of the AR element in the i-th frame image, or the coordinates of any point in the AR element in the i-th frame image, and this embodiment of the application does not limit this. In this embodiment of the application, the coordinates of the j-th AR element in the i-th frame image of the first video are (x j ,y j ).
[0061] S305: The electronic device calculates the target coordinates of each AR element in the i-th frame image based on the first skeleton node coordinates, the first skeleton node coordinates, the coordinates of each AR element, the second skeleton node coordinates, the second skeleton node coordinates and the first ratio value.
[0062] Specifically, the electronic device can calculate the target horizontal coordinate x′ of the jth AR element in the i-th frame image of the first video by formula (1): j , the electronic device can calculate the target vertical coordinate of the j-th AR element in the i-th frame image of the first video by formula (2). Formula (1) and formula (2) are as follows:
[0063] x′ j =x j *Scale+(X_center-x_center*Scale)(1)
[0064] y′ j =y j *Scale+(Y_center-y_center*Scale)(2)
[0065] S306: The electronic device displays the i-th frame image of the first video, where the coordinates of each AR element in the i-th frame image are corresponding target coordinates.
[0066] Specifically, after calculating the target coordinates of each AR element in the i-th frame of the first video, the electronic device displays the i-th frame of the first video on the screen and displays each AR element at the corresponding position based on the target coordinates in the i-th frame. This allows the position of the AR element in the i-th frame to better match the user's body proportions, facilitating smoother and easier learning of the actions in the first video, thereby improving the user's learning effect.
[0067] Optionally, after S306, before displaying the i+1th frame image of the first video, the electronic device may capture an image of the user through a camera, and the image is the second image of the user. Then, the electronic device obtains the coordinates of the skeletal node corresponding to the AR element in the second image, and the Euclidean distance between the coordinates of the skeletal node and the coordinates of the AR element to which it corresponds, and scores the user's completion of the action corresponding to the AR element based on the Euclidean distance. The smaller the Euclidean distance, the higher the score, and the larger the Euclidean distance, the lower the score. For example, Figure 4D The first image of the user and the i-th frame image of the first video are included. In the first image, the coordinates of the user's left fingertip bone node are (x1, y1), the coordinates of the right fingertip bone node are (x2, y2), the coordinates of the left toe bone node are (x3, y3), and the coordinates of the right toe bone node are (x4, y4). In the i-th frame image, the coordinates of the AR element 511 are (x 11 ,y 11 ), the coordinates of the AR element 512 are (x 21 ,y 21 ), the coordinates of the AR element 513 are (x 31 ,y 31 ), the coordinates of the AR element 514 are (x 41 ,y 41). And AR element 511 corresponds to the left fingertip bone node, AR element 512 corresponds to the right fingertip bone node, AR element 513 corresponds to the left toe bone node, and AR element 514 corresponds to the right toe bone node. Then, the electronic device calculates the first Euclidean distance d1 according to the coordinates of the left fingertip bone node and the coordinates of the AR element 511, calculates the second Euclidean distance d2 according to the coordinates of the right fingertip bone node and the coordinates of the AR element 512, calculates the third Euclidean distance d3 according to the coordinates of the left toe bone node and the coordinates of the AR element 513, and calculates the fourth Euclidean distance d4 according to the coordinates of the right toe bone node and the coordinates of the AR element 514. Assume that d1~d4 are 12, 120, 47, and 99 respectively. If the Euclidean distance is in the range of (0, 10), the score range is (90, 100); when the Euclidean distance is in the range of (10, 30), the score range is (85, 90); when the Euclidean distance is in the range of (30, 60), the score range is (70, 85); when the Euclidean distance is in the range of (60, 100), the score range is (60, 70); when the Euclidean distance is greater than 100, the score is less than 60. Therefore, the electronic device can display the following in each AR element according to the Euclidean distance: Figure 4E The corresponding scores are shown. Figure 4E , the score displayed by AR element 511 is 88, the score displayed by AR element 512 is 48, the score displayed by AR element 513 is 80, and the score displayed by AR element 514 is 61.
[0068] In an embodiment of the present application, before displaying the i-th frame image in the first video, the electronic device will capture the user's image through a camera. Then, the user's first skeletal node coordinates and the first motion space radius will be calculated based on the captured user image. In addition, the coordinates of each AR element in the i-th frame image will also be obtained. Then, the electronic device calculates the target coordinates of each AR element based on the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius and the coordinates of the AR element. Finally, the electronic device displays the i-th frame image based on the target coordinates of each AR element in the i-th frame image. In this way, the position of the actual AR image in the i-th frame image is more consistent with the user's body proportions, which solves the problem that in the first video, when the difference in body proportions between the moving object and the user is too large, the position of the AR element appearing in the video does not match the user, resulting in poor learning effect when the user learns the actions in the first video according to the instructions of the AR element.
[0069] In the above Figure 3In the embodiment, a process is described in which the electronic device determines the target coordinates of the AR element in the i-th frame image that matches the user's body proportions before displaying the i-th frame image of the first video. The AR elements in the first video are pre-set by the video publisher. In some embodiments, the user can set one or more AR elements for any action of the moving object in the second video that does not contain AR elements. Before the electronic device displays each frame image of the second video with the AR element set, it calculates the target coordinates of each AR element in each frame image, and then displays each frame image of the second video according to the target coordinates of the AR element. Below, in combination with Figure 5 The above process is introduced. Figure 5 This section only describes how to set an AR element for a single frame image in the second video and calculate the target coordinates of the AR element. Figure 5 , Figure 5 This is a flow chart of another video processing method provided by an embodiment of the present application. The specific process is as follows:
[0070] S501: The electronic device processes the second video to obtain all N frames of video images in the second video.
[0071] Specifically, after receiving the second video, the electronic device may process the second video to obtain video images frame by frame, and the number of video images in the second video is N. The second video is a video that does not include AR elements. The second video is an AR motion video. The AR video may be a fitness video, a limb rehabilitation training video, a gymnastics video, or other videos with object motion, and this embodiment of the present application does not limit this.
[0072] S502: The electronic device identifies the moving object in each frame of video image, obtains the skeleton node image corresponding to each frame of video image based on the moving object in each frame of video image, displays the skeleton node of the moving object in the skeleton node image, and the skeleton node image includes the position information of the skeleton node.
[0073] For example, Figure 6A is a frame of video image in the second video, and the electronic device recognizes Figure 6A After finding the moving object in the video, the skeleton node image of the frame will be obtained. Figure 6B The image of the skeleton nodes of the moving object in the frame of video image may include the knee joint, finger joints and other skeletal joints of the moving object. The electronic device records the position information of each skeleton node in the skeleton node image (for example, the coordinates of the skeleton node in the skeleton node image).
[0074] S503: The first operation is detected. In response to the first operation, the electronic device sets AR elements for M frames of skeleton node images among N frames of skeleton node images, and stores the coordinates of each AR element in the corresponding skeleton node image to obtain a first video.
[0075] Specifically, the first operation can be an operation in which the user sets an AR element on a skeleton node image. M is less than or equal to N. After the electronic device sets the AR element for each frame of the M skeleton node images, it records the coordinates of each AR element in the corresponding skeleton node image. In this way, the AR element is added to the second video, and the second video is a video with a moving object and an AR element, that is, the second video is the first video.
[0076] S504: The electronic device saves the first video.
[0077] The above steps S501 to S504 describe the process of the electronic device adding AR elements to the second video. Next, in conjunction with steps S505 to S510, the process of the electronic device determining the position of the AR element in the second video while the user of the electronic device is learning the action in the second video with the AR element added is described.
[0078] S505: The electronic device captures a first image of the user through a camera, and determines first skeletal node coordinates and a first motion space radius of the user according to the first image.
[0079] S506: The electronic device obtains the second skeleton node coordinates and the second motion space radius.
[0080] S507: The electronic device calculates a first ratio value based on the first movement space radius and the second movement space radius.
[0081] S508: The electronic device obtains the coordinates of each AR element in the i-th frame image of the first video.
[0082] S509: The electronic device calculates the target coordinates of each AR element in the i-th frame image based on the first skeleton node coordinates, the first skeleton node coordinates, the coordinates of each AR element, the second skeleton node coordinates, the second skeleton node coordinates and the first ratio value.
[0083] S510: The electronic device displays an i-th frame image of the first video, where the coordinates of each AR element in the i-th frame image are target coordinates.
[0084] For the related descriptions of S505 to S510, please refer to the related descriptions of S301 to S306 above, which will not be repeated here.
[0085] The structure of the electronic device 100 is introduced below. Figure 7 , Figure 7 Schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application.
[0086] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0087] It is understood that the structure shown in the embodiment of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include Figure 7 The diagrams may show more or fewer components, combinations of certain components, separations of certain components, or different arrangements of components. Figure 7 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0088] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0089] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0090] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0091] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0092] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcast, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0093] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0094] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0095] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0096] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise and brightness. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0097] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0098] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0099] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0100] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0101] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0102] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0103] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0104] The pressure sensor 180A is used to sense the pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194 .
[0105] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.
[0106] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100. Figure 8 This is a block diagram of the software structure of the electronic device 100 according to an embodiment of the present application. A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0107] The application layer can include a series of application packages. Figure 8 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0108] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 3 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0109] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0110] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0111] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0112] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0113] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0114] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0115] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0116] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0117] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0118] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0119] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0120] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0121] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0122] A 2D graphics engine is a drawing engine for 2D drawings.
[0123] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0124] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described herein are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0125] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0126] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0127] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0128] In short, the above description is only an embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made based on the disclosure of the present invention should be included in the scope of protection of the present invention.
Claims
1. A video processing method, characterized in that: Applied to an electronic device including a display screen and a camera, the method includes: collecting a first image of the user through the camera; Obtaining first skeletal node coordinates and a first motion space radius of the user in the first image; Calculating target coordinates of the AR element in the i-th frame image according to the first skeletal node coordinates, the first motion space radius, the second skeletal node coordinates, the second motion space radius, and the coordinates of the AR element in the i-th frame image of the first video; the i-th frame image is the next frame image to be played of the first video; The i-th frame image is displayed on the display screen; wherein the coordinates of the AR element in the displayed i-th frame image are the target coordinates.
2. The method according to claim 1, wherein Before collecting the first image of the user through the camera, the method further includes: Processing the second video to obtain all N frames of video images in the second video, where the second video does not include an AR element; If there is a P-frame video image with a moving object among the N-frame video images, the P-frame video image is processed to obtain a skeleton node image of the P-frame video image, wherein the skeleton node image displays the skeleton nodes of the moving object, and the skeleton node image includes position information of the skeleton nodes of the moving object; In response to the first operation, an AR element is set on the M-frame video image in the P-frame video image based on the position information of the skeleton node of the moving object, and the coordinates of the AR element are recorded to obtain the first video; The first video is saved.
3. The method according to claim 1, wherein The first skeletal node coordinates are the left sacrum coordinates or the right sacrum coordinates of the user in the first image; The first motion space radius is the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's left finger joints, or the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's right finger joints.
4. The method according to claim 2, wherein The first skeletal node coordinates are the left sacrum coordinates or the right sacrum coordinates of the user in the first image; The first motion space radius is the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's left finger joints, or the Euclidean distance between the coordinates of the center node of the user's neck bone in the first image and the coordinates of the user's right finger joints.
5. The method according to any one of claims 1 to 4, characterized in that The calculating the target coordinates of the AR element in the i-th frame image according to the first skeleton node coordinates, the first motion space radius, the second skeleton node coordinates, the second motion space radius, and the coordinates of the AR element in the i-th frame image of the first video specifically includes: calculating a first ratio value; Calculate the target coordinates of the AR element in the i-th frame image according to the first ratio value, the first skeleton node coordinates, and the second skeleton node coordinates.
6. The method according to claim 5, wherein The calculating of the first ratio value specifically includes: The first proportional value Scale is calculated according to the formula Scale=L / l, where l is the radius of the first motion space, and L is the radius of the second motion space.
7. The method according to claim 5, wherein The calculating of the first ratio value specifically includes: Calculating the left arm length l1, the right arm length l2, the left lower limb length l3, and the right lower limb length l4 of the user in the first image; Calculate the left arm length L1, the right arm length L2, the left lower limb length L3, and the right lower limb length L4 of the moving object in the i-th frame image; Then according to the formula Scale i =l i / L i , calculate Scale1, Scale2, Scale3 and Scale4 respectively; Among Scale1, Scale2, Scale3 and Scale4, select the Scale with the smallest value. i as the first proportional value.
8. The method according to claim 5, wherein The calculating the target coordinates of the AR element in the i-th frame image according to the first skeleton node coordinates, the second skeleton node coordinates, the first motion space radius, and the second motion space radius specifically includes: According to the formula x j ′ =x j *Scale + (X_center - x_center * Scale) to calculate the target horizontal coordinate of the j-th AR element in the i-th frame image; According to the formula y j ′ =y j *Scale + (Y_center - y_center * Scale) calculates the target ordinate of the j-th AR element in the i-th frame image; Among them, the x j ′ is the target horizontal coordinate of the j-th AR element, the y j ′ is the target vertical coordinate of the j-th AR element, the Scale is the first scale value, and the x j is the horizontal coordinate of the j-th AR element, the y j is the vertical coordinate of the j-th AR element, X_center is the horizontal coordinate of the second bone node, Y_center is the vertical coordinate of the second bone node, the x_center is the horizontal coordinate of the first bone node, and the y_center is the vertical coordinate of the first bone node.
9. The method according to any one of claims 1 to 4, 6 to 8, characterized in that: The second skeletal node coordinates are the left sacrum coordinates or the right sacrum coordinates of the moving object in the i-th frame image; The radius of the second motion space is the Euclidean distance between the coordinates of the center node of the neck bone of the moving object in the i-th frame image and the coordinates of the user's right finger joints, or the Euclidean distance between the coordinates of the center node of the neck bone of the moving object in the i-th frame image and the coordinates of the user's left finger joints, or a fixed distance value stored in the electronic device.
10. An electronic device, characterized in that: include: Memory, processor, and touch screen; including: The touch screen is used to display content; the touch screen includes a display screen; The memory is used to store a computer program, wherein the computer program includes program instructions; The processor is configured to call the program instructions so that the electronic device executes the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Quick human movement identification method oriented to human-computer interaction
CN107908288A
Human motion evaluation method based on two-dimensional skeleton sequences
CN108597578A