An image generation method, apparatus, electronic device, and storage medium

By selecting the parent feature point in the video frame and calculating the offset value of the child feature point, the coordinates of the feature points of the virtual object are adjusted, thus solving the problem of low accuracy of 3D coordinates in the pose recognition algorithm and achieving higher coordinate calculation accuracy.

CN115457175BActive Publication Date: 2026-05-26BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING QIYI CENTURY SCI & TECH CO LTD
Filing Date
2022-09-23
Publication Date
2026-05-26

Smart Images

  • Figure CN115457175B_ABST
    Figure CN115457175B_ABST
Patent Text Reader

Abstract

This invention provides an image generation method, apparatus, electronic device, and storage medium, relating to the field of image processing technology. The method includes: selecting a parent feature point from the feature points of a target object; obtaining the target coordinates of the parent feature point corresponding to each video frame in the target scene; calculating the processing offset value of the target coordinates of each child feature point of the parent feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed; calculating the target coordinates of the child feature point corresponding to the video frame in the target scene based on the processing offset value and the target coordinates of the parent feature point corresponding to the video frame; and adjusting the coordinates of each feature point of a virtual object in each preset image according to the target coordinates of each feature point of the target object corresponding to each video frame in the target scene, thereby obtaining a target video in which the virtual object and the target object have the same actions, which can improve the accuracy of the calculated target coordinates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] When creating a video containing virtual objects (e.g., virtual characters, virtual animals, etc.), video footage of a character performing a specified action can be captured (this can be called the original video). For each frame of the original video, pose recognition is performed on the target object in that frame to obtain the 3D coordinates of each joint point of the corresponding character in the scene. Then, based on the 3D coordinates of each joint point of the character in each video frame, the 3D coordinates of each joint point of the virtual object in each preset image are adjusted to obtain the target video containing the virtual object performing the specified action.

[0003] Due to factors such as lighting and occlusion during the capture of the original video, when using a pose recognition algorithm to identify a video frame, it may be impossible to identify all the joints of the person corresponding to that video frame. This results in the loss of some joints, meaning the 3D coordinates of each joint cannot be determined. In related technologies, if the 3D coordinates of a particular joint (which can be called the target joint) of the person corresponding to that video frame cannot be determined, the average of the 3D coordinates of the target joint in the preceding and following video frames is calculated as the 3D coordinates of the target joint in the current video frame.

[0004] However, if the 3D coordinates of the target joint corresponding to this video frame differ significantly from those of the target joint corresponding to the previous video frame and the target joint corresponding to the next video frame, then calculating the average of the 3D coordinates of the target joint corresponding to the previous video frame and the target joint corresponding to the next video frame will yield a low accuracy in obtaining the 3D coordinates of the target joint corresponding to this video frame. Summary of the Invention

[0005] The purpose of this invention is to provide an image generation method, apparatus, electronic device, and storage medium to improve the accuracy of the determined target coordinates. The specific technical solution is as follows:

[0006] In a first aspect of the present invention, an image generation method is provided, the method comprising:

[0007] Select one feature point from the feature points of the target object in the original video as the parent feature point;

[0008] For each video frame in the original video, obtain the target coordinates of the parent feature point corresponding to that video frame in the target scene;

[0009] For each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed is calculated as the offset value to be processed; wherein, the child feature points of the parent feature point include: feature points connected to the parent feature point among the feature points of the target object; the video frame to be processed includes: the first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or, the second video frame whose timestamp is greater than the timestamp of the video frame;

[0010] Based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, calculate the target coordinates of the child feature point corresponding to the video frame in the target scene;

[0011] Based on the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, the coordinates of each feature point of the virtual object in each preset image are adjusted to obtain a target video in which the virtual object has the same action as the target object.

[0012] Optionally, after calculating the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, the method further includes:

[0013] The process of taking the sub-feature point of the target object as the parent feature point and returning to perform the steps of calculating the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point for each sub-feature point of the parent feature point, and using it as the offset value to be processed, continues until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are determined.

[0014] Optionally, the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and a second video frame whose timestamp is greater than the timestamp of the video frame;

[0015] For each child feature point of the parent feature point, the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed is calculated as the processing offset value, including:

[0016] For each child feature point of the parent feature point, calculate the offset value of the target coordinate of the child feature point corresponding to the first video frame relative to the target coordinate of the parent feature point corresponding to the first video frame, and use it as the offset value to be processed for the first video frame.

[0017] Calculate the offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the parent feature point corresponding to the second video frame, and use it as the offset value to be processed for the second video frame.

[0018] Optionally, calculating the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame includes:

[0019] Calculate the average value of the offset to be processed corresponding to the first video frame and the offset to be processed corresponding to the second video frame, and use it as the average offset value.

[0020] The sum of the average offset value and the target coordinates of the parent feature point corresponding to the video frame is calculated to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0021] Optionally, selecting a feature point from the feature points of the target object in the original video as the parent feature point includes:

[0022] From the feature points of the target object in the original video, determine the specified central feature point as the parent feature point;

[0023] The step of obtaining the target coordinates of the parent feature point in the target scene for each video frame in the original video includes:

[0024] Obtain the initial coordinates and corresponding first confidence level of the central feature point corresponding to the video frame; wherein, the first confidence level represents the probability that the central feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame;

[0025] If the first confidence level is less than a preset threshold, the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame is calculated to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, which is used as the target coordinates of the parent feature point corresponding to the video frame in the target scene; wherein, the third video frame includes: video frames in the original video with timestamps less than the timestamp of the video frame; the fourth video frame includes: video frames in the original video with timestamps greater than the timestamp of the video frame;

[0026] If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

[0027] Optionally, before calculating the offset value of the target coordinates of the sub-feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point for each sub-feature point of the parent feature point, and using this offset value as the processing offset value, the method further includes:

[0028] For each child feature point of the parent feature point, obtain the initial coordinates and corresponding second confidence level of the child feature point in the target scene corresponding to the video frame; wherein, the second confidence level represents the probability that the child feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the child feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the feature point in the video frame;

[0029] For each child feature point of the parent feature point, the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed is calculated as the processing offset value, including:

[0030] If the second confidence level is less than a preset threshold, calculate the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed.

[0031] The method further includes:

[0032] If the second confidence level is not less than the preset threshold, the initial coordinates of the sub-feature point corresponding to the video frame in the target scene are taken as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

[0033] Optionally, before obtaining the target coordinates of the parent feature point in the target scene for each video frame in the original video, the method further includes:

[0034] Obtain the original videos of the target scene from multiple perspectives, and extract multiple video frames with the same timestamp from each original video to obtain a video frame group;

[0035] For each video frame in the video frame group, pose recognition is performed on the video frame to obtain the two-dimensional coordinates of each feature point of the target object in the video frame.

[0036] For each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, the three-dimensional coordinates of the feature point corresponding to each video frame in the video frame group in the target scene are calculated, and used as the initial coordinates of the feature point corresponding to each video frame in the video frame group in the target scene.

[0037] Optionally, for each feature point of the target object, calculating the three-dimensional coordinates of the feature point in the target scene based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, and using these as the initial coordinates of the feature point in the target scene for each video frame of the video frame group, includes:

[0038] For each feature point of the target object, based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group, and the transformation relationships between the image coordinate system of the two video frames and the three-dimensional coordinate system of the target scene, the three-dimensional coordinates of the feature point in the target scene are calculated as the three-dimensional coordinates to be processed; the average value of each three-dimensional coordinate to be processed is calculated as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0039] Optionally, the step of performing pose recognition on each video frame in the video frame group to obtain the two-dimensional coordinates of each feature point of the target object in the video frame includes:

[0040] For each video frame in the video frame group, the video frame is input into a pre-trained pose recognition model to obtain the two-dimensional coordinates and corresponding confidence scores of each feature point of the target object in the video frame; wherein, the confidence score corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame;

[0041] The step of obtaining the second confidence score corresponding to the initial coordinates of the sub-feature point in the target scene of the video frame includes:

[0042] The mean of the confidence scores corresponding to the two-dimensional coordinates of the sub-feature point in each video frame of the video frame group to which the video frame belongs is calculated and used as the second confidence score.

[0043] The step of obtaining the first confidence score corresponding to the initial coordinates of the center feature point of the video frame includes:

[0044] The mean confidence score of the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs is calculated and used as the first confidence score.

[0045] Optionally, before calculating the three-dimensional coordinates of the feature point in the target scene corresponding to each video frame in the video frame group for each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, and using these as the initial coordinates of the feature point in the target scene corresponding to each video frame of the video frame group, the method further includes:

[0046] Based on the two-dimensional coordinates of feature points of multiple target objects in each video frame of the video frame group, the same target objects in each video frame of the video frame group are determined.

[0047] Optionally, determining the same target object in each video frame of the video frame group based on the two-dimensional coordinates of feature points of multiple target objects in each video frame of the video frame group includes:

[0048] For each target object, the epipolar distance from each feature point of the target object to the corresponding epipolar plane is calculated, and the mean value of each epipolar distance corresponding to each feature point of the target object is calculated to obtain the mean distance corresponding to the target object; wherein, the epipolar plane corresponding to a feature point represents the plane in which the feature point is located in the target scene;

[0049] For every two video frames in the video frame group, the similarity between the two target objects is calculated based on the average distance between them, resulting in a first similarity matrix. An element in the first similarity matrix represents the probability that the two target objects in the two video frames are the same.

[0050] The two video frames are input into a pre-trained object matching model to obtain the similarity between every two target objects in the two video frames, resulting in a second similarity matrix; where each element in the second similarity matrix represents the probability that the two corresponding target objects in the two video frames are the same.

[0051] The first similarity matrix and the second similarity matrix are fused to obtain the target similarity matrix;

[0052] Based on the target similarity matrix, the same target objects in each video frame of the video frame group are identified.

[0053] In a second aspect of the invention, an image generation apparatus is also provided, the apparatus comprising:

[0054] The selection module is used to select one feature point from the feature points of the target object in the original video as the parent feature point;

[0055] The first acquisition module is used to acquire the target coordinates of the parent feature point in the target scene for each video frame in the original video.

[0056] The first determining module is used to calculate, for each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, as the offset value to be processed; wherein, the child feature points of the parent feature point include: feature points connected to the parent feature point among the feature points of the target object; the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or a second video frame whose timestamp is greater than the timestamp of the video frame;

[0057] The second determining module is used to calculate the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame.

[0058] The generation module is used to adjust the coordinates of each feature point of the virtual object in each preset image according to the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, so as to obtain a target video in which the virtual object has the same action as the target object.

[0059] Optionally, the device further includes:

[0060] The processing module is configured to, after the second determining module performs the following steps: calculates the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame; then, takes the child feature point of the target object as the parent feature point and returns to perform the following steps: for each child feature point of the parent feature point, calculates the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed, and uses it as the offset value to be processed, until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are determined.

[0061] Optionally, the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and a second video frame whose timestamp is greater than the timestamp of the video frame;

[0062] The first determining module is specifically used to calculate, for each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the first video frame relative to the target coordinate of the parent feature point corresponding to the first video frame, and use it as the offset value to be processed for the first video frame.

[0063] Calculate the offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the parent feature point corresponding to the second video frame, and use it as the offset value to be processed for the second video frame.

[0064] Optionally, the second determining module is specifically used to calculate the average of the offset value to be processed corresponding to the first video frame and the offset value to be processed corresponding to the second video frame, as the average offset value.

[0065] The sum of the average offset value and the target coordinates of the parent feature point corresponding to the video frame is calculated to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0066] Optionally, the selection module is specifically used to determine a specified central feature point from the feature points of the target object in the original video, as the parent feature point;

[0067] The first acquisition module is specifically used to acquire the initial coordinates and the corresponding first confidence level of the central feature point corresponding to the video frame; wherein, the first confidence level represents the probability that the central feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame;

[0068] If the first confidence level is less than a preset threshold, the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame is calculated to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, which is used as the target coordinates of the parent feature point corresponding to the video frame in the target scene; wherein, the third video frame includes: video frames in the original video with timestamps less than the timestamp of the video frame; the fourth video frame includes: video frames in the original video with timestamps greater than the timestamp of the video frame;

[0069] If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

[0070] Optionally, the device further includes:

[0071] The second acquisition module is configured to, before the first determining module performs the following steps for each sub-feature point of the parent feature point: calculating the offset value of the target coordinates of the sub-feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point of the video frame to be processed, and using this offset value as the processing offset value, perform the following steps for each sub-feature point of the parent feature point: acquiring the initial coordinates and corresponding second confidence level of the sub-feature point corresponding to the video frame in the target scene; wherein, the second confidence level represents the probability that the sub-feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the sub-feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the feature point in the video frame;

[0072] The first determining module is specifically used to calculate, if the second confidence level is less than a preset threshold, the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed.

[0073] The device further includes:

[0074] The third determining module is used to, if the second confidence level is not less than the preset threshold, take the initial coordinates of the sub-feature point corresponding to the video frame in the target scene as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

[0075] Optionally, the device further includes:

[0076] The third acquisition module is used to acquire the original video of the target scene from multiple perspectives before the first acquisition module acquires the target coordinates of the parent feature point corresponding to each video frame in the original video. The module then acquires multiple video frames with the same timestamp from each original video to obtain a video frame group.

[0077] The recognition module is used to perform pose recognition on each video frame in the video frame group to obtain the two-dimensional coordinates of each feature point of the target object in the video frame.

[0078] The fourth determining module is used to calculate, for each feature point of the target object, the three-dimensional coordinates of the feature point in the target scene based on the two-dimensional coordinates of the feature point in each video frame of the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, and use these as the initial coordinates of the feature point in the target scene corresponding to each video frame of the video frame group.

[0079] Optionally, the fourth determining module is specifically used to, for each feature point of the target object, calculate the three-dimensional coordinates of the feature point in the target scene based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group, and the transformation relationships between the image coordinate system of the two video frames and the three-dimensional coordinate system of the target scene, as the three-dimensional coordinates to be processed; and calculate the average value of each three-dimensional coordinate to be processed as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0080] Optionally, the recognition module is specifically used to input each video frame in the video frame group into a pre-trained pose recognition model to obtain the two-dimensional coordinates and corresponding confidence scores of each feature point of the target object in the video frame; wherein, the confidence score corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame;

[0081] The second acquisition module is specifically used to calculate the average confidence level of the two-dimensional coordinates of the sub-feature point in each video frame of the video frame group to which the video frame belongs, and use it as the second confidence level.

[0082] The first acquisition module is specifically used to calculate the average confidence level of the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs, and use it as the first confidence level.

[0083] Optionally, the device further includes:

[0084] The matching module is used to perform the following steps in the fourth determining module: for each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, calculate the three-dimensional coordinates of the feature point in each video frame of the video frame group in the target scene, and use these coordinates as the initial coordinates of the feature point in each video frame of the video frame group in the target scene. Before this, the module performs the following steps: based on the two-dimensional coordinates of the feature points of multiple target objects in each video frame of the video frame group, determine the same target objects in each video frame of the video frame group.

[0085] Optionally, the matching module is specifically used to calculate the epipolar distance from each feature point of the target object to the corresponding epipolar plane for each target object, and to calculate the average value of each epipolar distance corresponding to each feature point of the target object, so as to obtain the average distance corresponding to the target object; wherein, the epipolar plane corresponding to a feature point represents the plane in which the feature point is located in the target scene;

[0086] For every two video frames in the video frame group, the similarity between the two target objects is calculated based on the average distance between them, resulting in a first similarity matrix. An element in the first similarity matrix represents the probability that the two target objects in the two video frames are the same.

[0087] The two video frames are input into a pre-trained object matching model to obtain the similarity between every two target objects in the two video frames, resulting in a second similarity matrix; where each element in the second similarity matrix represents the probability that the two corresponding target objects in the two video frames are the same.

[0088] The first similarity matrix and the second similarity matrix are fused to obtain the target similarity matrix;

[0089] Based on the target similarity matrix, the same target objects in each video frame of the video frame group are identified.

[0090] In another aspect of the present invention, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0091] Memory, used to store computer programs;

[0092] The processor, when executing a program stored in memory, implements any of the image generation method steps described above.

[0093] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of any of the above-described image generation methods.

[0094] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the image generation methods described above.

[0095] This invention provides an image generation method that, from the feature points of a target object in an original video, selects one feature point as a parent feature point; for each video frame in the original video, obtains the target coordinates of the parent feature point in the target scene corresponding to that video frame; for each child feature point of the parent feature point, calculates the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed, as the offset value to be processed; the child feature points of the parent feature point include: feature points of the target object connected to the parent feature point; the video frames to be processed include: a first video frame in the original video with a timestamp less than the timestamp of the video frame, and / or a second video frame with a timestamp greater than the timestamp of the video frame; based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, calculates the target coordinates of the child feature point corresponding to the video frame in the target scene; according to the target coordinates of the feature points of the target object corresponding to each video frame in the target scene, adjusts the coordinates of the feature points of virtual objects in each preset image to obtain a target video in which the virtual objects and the target objects have the same actions.

[0096] Based on the above processing, for each child feature point of the parent feature point of the target object, the offset value to be processed is the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the child feature point corresponding to the video frame to be processed. Since the parent feature point and the child feature point of the target object are connected, it means that the offset value of the child feature point relative to the parent feature point is fixed in each video frame. Accordingly, based on the offset value to be processed and the target coordinate of the parent feature point corresponding to the video frame, the target coordinate of the child feature point corresponding to the video frame in the target scene can be calculated, which can improve the accuracy of the calculated target coordinates. Attached Figure Description

[0097] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0098] Figure 1 This is a flowchart of an image generation method provided in an embodiment of the present invention;

[0099] Figure 2(a) is a schematic diagram of the joints of a target object provided in an embodiment of the present invention;

[0100] Figure 2(b) is a schematic diagram of the joints of another target object provided in an embodiment of the present invention;

[0101] Figure 3 This is a flowchart of another image generation method provided in an embodiment of the present invention;

[0102] Figure 4 This is a flowchart of a method for determining the initial coordinates of each feature point of a target object according to an embodiment of the present invention;

[0103] Figure 5 This is a schematic diagram illustrating the principle of camera calibration provided in an embodiment of the present invention;

[0104] Figure 6(a) is a schematic diagram of a video frame containing a target scene provided in an embodiment of the present invention;

[0105] Figure 6(b) is a schematic diagram of another video frame containing a target scene provided in an embodiment of the present invention;

[0106] Figure 7 This is a flowchart of another method for determining the initial coordinates of feature points of a target object provided in an embodiment of the present invention;

[0107] Figure 8 A flowchart illustrating another method for determining the initial coordinates of feature points of a target object provided in this embodiment of the invention.

[0108] Figure 9 This is a schematic diagram of a target similarity matrix provided in an embodiment of the present invention;

[0109] Figure 10(a) is a schematic diagram of the joints of a target object provided in an embodiment of the present invention;

[0110] Figure 10(b) is a schematic diagram of the joints of another target object provided in an embodiment of the present invention;

[0111] Figure 10(c) is a schematic diagram of the joints of another target object provided in an embodiment of the present invention;

[0112] Figure 10(d) is a schematic diagram of the joints of another target object provided in an embodiment of the present invention;

[0113] Figure 11 This is a flowchart of a method for determining the target coordinates of a target object provided in an embodiment of the present invention;

[0114] Figure 12 This is a flowchart of another method for determining the target coordinates of a target object provided in an embodiment of the present invention;

[0115] Figure 13 This is a structural diagram of an image generation device provided in an embodiment of the present invention;

[0116] Figure 14 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0117] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0118] In related technologies, if the 3D coordinates of the target joint point of the target object corresponding to a video frame differ significantly from the 3D coordinates of the target joint point corresponding to the previous video frame and the 3D coordinates of the target joint point corresponding to the next video frame, then the accuracy of the 3D coordinates of the target joint point corresponding to the video frame obtained by calculating the average of the 3D coordinates of the target joint point corresponding to the previous video frame and the 3D coordinates of the target joint point corresponding to the next video frame is low.

[0119] To address the aforementioned problems, this invention provides an image generation method applied to an electronic device, which can be a server. The electronic device can select one feature point from the feature points of a target object in the original video as a parent feature point. For each video frame in the original video, the target coordinates of the parent feature point in the target scene corresponding to that video frame are obtained. For each child feature point of the parent feature point, a processing offset value is calculated relative to the target coordinates of the parent feature point of the video frame to be processed. Based on the processing offset value and the target coordinates of the parent feature point of the video frame, the target coordinates of the child feature point of the video frame in the target scene are calculated. Furthermore, a target video can be generated based on the target coordinates of each child feature point of the target object in the target scene corresponding to the video frame.

[0120] Since the parent feature point and child feature point of the target object are connected, it means that the offset value of the child feature point relative to the parent feature point is fixed in each video frame. Accordingly, based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, the target coordinates of the child feature point in the target scene corresponding to the video frame can be calculated, which can improve the accuracy of the calculated target coordinates.

[0121] See Figure 1 , Figure 1 A flowchart of an image generation method provided in an embodiment of the present invention, the method may include the following steps:

[0122] S101: Select one feature point from the feature points of the target object in the original video as the parent feature point.

[0123] S102: For each video frame in the original video, obtain the target coordinates of the parent feature point corresponding to that video frame in the target scene.

[0124] S103: For each child feature point of the parent feature point, calculate the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed.

[0125] Among them, the child feature points of the parent feature point include: the feature points of the target object that are connected to the parent feature point; the video frames to be processed include: the first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or the second video frame whose timestamp is greater than the timestamp of the video frame.

[0126] S104: Based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, calculate the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0127] S105: Adjust the coordinates of each feature point of the virtual object in each preset image according to the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, so as to obtain a target video in which the virtual object and the target object have the same action.

[0128] Based on the image generation method provided in this embodiment of the invention, for each child feature point of the parent feature point of the target object, the offset value to be processed is the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the child feature point corresponding to the video frame to be processed. Since the parent feature point and the child feature point of the target object are connected, it means that the offset value of the child feature point relative to the parent feature point is fixed in each video frame. Accordingly, based on the offset value to be processed and the target coordinate of the parent feature point corresponding to the video frame, the target coordinate of the child feature point corresponding to the video frame in the target scene is calculated, which can improve the accuracy of the calculated target coordinates.

[0129] For step S101, the original video is a video of the target scene from any perspective, and the target object is a moving object in the target scene. For example, the target scene can be a street, and the target object can be vehicles, animals, and people in the street. Alternatively, the target scene can be a stage, and the target object can be people.

[0130] The feature points of a target object can be any points within the target object that can characterize it. For example, when the target object is a vehicle, the feature points can include points on the vehicle's outline. Or, when the target object is an animal, the feature points can include the joints of the animal's skeleton. Or, when the target object is a person, the feature points can include the joints of the human skeleton.

[0131] For example, referring to Figure 2(a), the target object is a human figure, and the target object's joints include 25 joints from joint 0 to joint 24. Furthermore, in addition to the joints of the human skeleton shown in Figure 2(a), the target object's joints also include joints of specific parts of the human body. For example, Figure 2(b) shows 21 joints of one hand of the human figure.

[0132] In one implementation, the electronic device can select any one of the feature points of the target object as the current parent feature point. For example, in the embodiment of Figure 2(a), when the target object is a person, the electronic device can select joint point 1 as the current parent feature point. Alternatively, the electronic device can select joint point 0 as the current parent feature point.

[0133] In another implementation, the electronic device can determine a specified central feature point from among the feature points of the target object, and use it as the current parent feature point. For example, in the embodiment of Figure 2(a), when the target object is a person, the electronic device can determine the hip joint (i.e., joint point 8) as the central feature point, that is, determine joint point 8 as the current parent feature point.

[0134] For step S102, for each video frame in the original video frame, the target coordinates of the current parent feature point corresponding to the video frame represent the position of the current parent feature point of the target object in the target scene at the time the video frame was acquired.

[0135] The original video contains multiple video frames. For each feature point of the target object, the target coordinates of that feature point in the target scene corresponding to a video frame are determined based on the two-dimensional coordinates of that feature point in that video frame.

[0136] The target coordinates of a feature point in the target scene can be either the two-dimensional coordinates of the feature point in the target scene or the three-dimensional coordinates of the feature point in the target scene.

[0137] In one implementation, when the target coordinates of a feature point in the target scene are two-dimensional coordinates of the feature point in the target scene, the electronic device can obtain the two-dimensional coordinates of the feature point in the video frame and use them as the target coordinates of the feature point in the target scene.

[0138] In another implementation, for each feature point of the target object, the electronic device acquires the three-dimensional coordinates (i.e., the initial coordinates in subsequent embodiments) of that feature point in the target scene corresponding to the video frame, and determines the target coordinates of that feature point in the target scene based on the acquired initial coordinates. The method by which the electronic device acquires the initial coordinates of each feature point of the target object in the target scene can be found in the relevant description in subsequent embodiments.

[0139] When the target coordinates of a feature point in the target scene are two-dimensional coordinates of the feature point in the target scene, for the current parent feature point of the target object, the electronic device can select any two coordinate values ​​from the three-dimensional coordinates of the current parent feature point in the target scene as the target coordinates of the current parent feature point in the target scene.

[0140] When the target coordinates of a feature point in the target scene are the three-dimensional coordinates of the feature point in the target scene, the electronic device can determine the target coordinates of the current parent feature point in the following way.

[0141] Method 1

[0142] Electronic devices can directly obtain the initial coordinates of the current parent feature point in the target scene corresponding to the video frame, and use them as the target coordinates of the current parent feature point in the target scene corresponding to the video frame group.

[0143] Method 2

[0144] When an electronic device determines that a specified center feature point is the current parent feature point, the electronic device can determine the target coordinates of the current parent feature point in the following manner, and correspondingly, in Figure 1 Based on this, see Figure 3 Step S101 may include the following steps:

[0145] S1011: Determine the specified central feature point from the feature points of the target object in the original video, and use it as the parent feature point.

[0146] Accordingly, step S102 may include the following steps:

[0147] S1021: Obtain the initial coordinates and the corresponding first confidence level of the center feature point corresponding to the video frame.

[0148] The first confidence level represents the probability that the central feature point is located at the position indicated by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame.

[0149] S1022: If the first confidence level is less than the preset threshold, calculate the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, and use it as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

[0150] The third video frame includes video frames in the original video whose timestamps are less than the timestamp of this video frame; the fourth video frame includes video frames in the original video whose timestamps are greater than the timestamp of this video frame.

[0151] S1023: If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

[0152] For cases where the target coordinates of a feature point in the target scene are its three-dimensional coordinates, the initial coordinates of the current parent feature point corresponding to this video frame are determined based on the two-dimensional coordinates of the current parent feature point in each video frame within the video frame group to which this video frame belongs. The video frame group to which this video frame belongs includes other video frames of the target scene with different perspectives from this video frame.

[0153] The first confidence level corresponding to the initial coordinates of the central feature point in the video frame represents the probability that the central feature point is located at the position indicated by the initial coordinates in the target scene.

[0154] In some embodiments, for each video frame, if the electronic device determines the initial coordinates of each feature point of the target object based on the two-dimensional coordinates of each feature point of the target object in each video frame of the video frame group to which the video frame belongs, the electronic device can input each video frame into a pre-trained pose recognition model to obtain the two-dimensional coordinates of each feature point of the target object in the video frame and the corresponding confidence score. Furthermore, the electronic device can calculate the first confidence score based on the following method.

[0155] Accordingly, step S1021 may include the following steps: calculating the mean of the confidence scores corresponding to the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs, and using it as the first confidence score.

[0156] The confidence level corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame.

[0157] If the confidence level of the two-dimensional coordinates of a feature point of a target object in the video frame is low, it means that the accuracy of the two-dimensional coordinates of the feature point of the target object in the video frame is low. Consequently, the accuracy of the initial coordinates of the feature point of the target object determined based on the two-dimensional coordinates of the feature point of the target object in the video frame is also low.

[0158] The electronic device can calculate the mean of the confidence scores corresponding to the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs, and use this as the first confidence score. Furthermore, the first confidence score can represent the probability that the central feature point is located at the position indicated by the initial coordinates in the target scene.

[0159] If the initial confidence level is less than a preset threshold, meaning the probability that the central feature point is located at the position indicated by the initial coordinates in the target scene is low, it indicates that the accuracy of the initial coordinates of the central feature point is low. The electronic device can obtain the target coordinates of the central feature point corresponding to the third video frame and the target coordinates of the central feature point corresponding to the fourth video frame, and calculate the average of the target coordinates of the central feature point corresponding to the third video frame and the target coordinates of the central feature point corresponding to the fourth video frame as the target coordinates of the central feature point in the target scene, which is also the target coordinates of the current parent feature point in the target scene.

[0160] A video frame's timestamp indicates its position within the original video. For example, if the original video has a frame rate of 25 FPS (Frames Per Second), then a video frame lasts for 40ms within the original video. Therefore, the timestamp of the first video frame in the original video is 40ms, the timestamp of the second video frame is 80ms, the timestamp of the third video frame is 120ms, and so on.

[0161] The third video frame includes any video frame in the original video whose timestamp is less than the timestamp of this video frame; that is, the third video frame is located before this video frame in the original video. The fourth video frame includes any video frame in the original video whose timestamp is less than the timestamp of this video frame; that is, the fourth video frame is located after this video frame in the original video.

[0162] For example, the video frame is the 5th video frame in the original video. When the first confidence level corresponding to the center feature point of the 5th video frame is less than a preset threshold, the electronic device determines the video frame in the original video whose target coordinates of the corresponding center feature point have been obtained from the video frames located before the 5th video frame and uses it as the third video frame.

[0163] The video frames preceding the 5th video frame include the 1st to 4th video frames in the original video. If the target coordinates of the center feature point corresponding to the 4th video frame have been determined, then the 4th video frame is determined as the 3rd video frame. If the target coordinates of the center feature point corresponding to the 4th video frame have not been determined, but the target coordinates of the center feature point corresponding to the 3rd video frame have been determined, then the 3rd video frame is determined as the 3rd video frame, and so on, until the 3rd video frame whose target coordinates of the corresponding center feature point have been determined is determined.

[0164] The electronic device determines the video frame from the original video frames after the 5th video frame where the target coordinates of the corresponding center feature point have been acquired, and uses it as the fourth video frame.

[0165] The video frames following the 5th video frame include the 6th to 10th video frames in the original video. If the target coordinates of the center feature point corresponding to the 6th video frame have been determined, then the 6th video frame is determined as the 4th video frame. If the target coordinates of the center feature point corresponding to the 6th video frame have not been determined, but the target coordinates of the center feature point corresponding to the 7th video frame have been determined, then the 7th video frame is determined as the 4th video frame, and so on, until the 4th video frame whose target coordinates of the corresponding center feature point have been determined is determined.

[0166] If the first confidence level is not less than the preset threshold, that is, the probability that the central feature point is located at the position represented by the initial coordinates in the target scene is high, indicating that the accuracy of the initial coordinates of the central feature point is high, then the electronic device can directly obtain the initial coordinates of the central feature point corresponding to the video frame as the target coordinates of the central feature point corresponding to the video frame in the target scene, that is, the target coordinates of the current parent feature point in the target scene.

[0167] In some embodiments, for each video frame in the original video, the video frame may contain multiple target objects, and the same target objects may be contained in different video frames. For each pair of target objects in two adjacent video frames, the electronic device can determine whether the two target objects are the same target object based on the initial coordinates of the feature points of the two target objects corresponding to the two adjacent video frames, so as to realize object tracking between different video frames, that is, to determine the same target objects contained in different video frames.

[0168] The electronic device can calculate the difference between the initial coordinates of each feature point of the two target objects corresponding to two adjacent video frames, and calculate the mean of each difference to obtain the feature point error between the two target objects. Accordingly, for each target object, the target object with the smallest feature point error between it and the target object is identified.

[0169] For example, the first video frame in the original video corresponds to object 1, object 2, and object 3, and the second video frame corresponds to object A, object B, and object C. The electronic device calculates the difference between the initial coordinates of the first feature point of object 1 corresponding to the first video frame and the initial coordinates of the first feature point of object A corresponding to the second video frame. Then it calculates the difference between the initial coordinates of the second feature point of object 1 corresponding to the first video frame and the initial coordinates of the second feature point of object A corresponding to the second video frame, and so on, to obtain multiple differences.

[0170] Then, the average of each difference is calculated to obtain the feature point error between object 1 and object A. Similarly, the feature point error between object 1 and object B, and the feature point error between object 1 and object C can be calculated. Furthermore, if the feature point error between object A and object 1 is the smallest among object A, object B, and object C, then object 1 and object A can be determined to be the same object.

[0171] Subsequently, for each target object, the electronic device can calculate the target coordinates of the target object in the target scene based on the initial coordinates of each feature point of the target object corresponding to different video frames.

[0172] For steps S103 and S104, the child feature points of the parent feature point include: each feature point of the target object that is connected to the parent feature point. For example, if the target object is a person, the feature points of the target object include each joint point of the person's skeleton. In the embodiment of Figure 2(a), if joint point 1 is the current parent feature point, then the child feature points of the current parent feature point include: joint point 2, joint point 0, joint point 5, and joint point 8. Alternatively, if joint point 0 is the current parent feature point, then the child feature points of the current parent feature point include: joint point 1, joint point 15, and joint point 16.

[0173] In one implementation, the video frame to be processed includes: the first video frame in the original video whose timestamp is less than the timestamp of the video frame; that is, the video frame to be processed includes: the first video frame in the original video that is located before the video frame.

[0174] For each child feature point of the current parent feature point, the electronic device calculates the difference between the target coordinates of the child feature point corresponding to the first video frame and the target coordinates of the current parent feature point corresponding to the first video frame, and obtains the offset value (i.e. the offset value to be processed) of the target coordinates of the child feature point corresponding to the first video frame relative to the target coordinates of the current parent feature point corresponding to the first video frame.

[0175] Then, the electronic device calculates the sum of the offset value to be processed corresponding to the first video frame and the target coordinates of the current parent feature point corresponding to the video frame, and obtains the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0176] In another implementation, the video frames to be processed include: a second video frame in the original video whose timestamp is greater than that of the video frame; that is, the video frames to be processed include: a second video frame in the original video that is located after the video frame.

[0177] For each child feature point of the current parent feature point, the electronic device calculates the difference between the target coordinates of the child feature point corresponding to the second video frame and the target coordinates of the current parent feature point corresponding to the second video frame, and obtains the offset value (i.e. the offset value to be processed) of the target coordinates of the child feature point corresponding to the second video frame relative to the target coordinates of the current parent feature point corresponding to the second video frame.

[0178] Then, the electronic device calculates the sum of the offset value to be processed corresponding to the second video frame and the target coordinates of the current parent feature point corresponding to the video frame, and obtains the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0179] In another implementation, the video frames to be processed include: a first video frame in the original video with a timestamp less than the timestamp of this video frame, and a second video frame with a timestamp greater than the timestamp of this video frame. That is, the video frames to be processed include: a first video frame in the original video preceding this video frame, and a second video frame in the original video following this video frame.

[0180] Accordingly, step S103 may include the following steps:

[0181] Step 1: For each child feature point of the parent feature point, calculate the offset value of the target coordinate of the child feature point corresponding to the first video frame relative to the target coordinate of the parent feature point corresponding to the first video frame, and use it as the offset value to be processed for the first video frame.

[0182] Step 2: Calculate the offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the parent feature point corresponding to the second video frame, and use it as the offset value to be processed for the second video frame.

[0183] For each child feature point of the current parent feature point, the electronic device can calculate the difference between the target coordinates of the child feature point corresponding to the first video frame and the target coordinates of the current parent feature point corresponding to the first video frame, and obtain the target offset value of the child feature point corresponding to the first video frame relative to the target coordinates of the current parent feature point corresponding to the first video frame.

[0184] The electronic device can also calculate the difference between the target coordinates of the sub-feature point corresponding to the second video frame and the target coordinates of the current parent feature point corresponding to the second video frame, and obtain the unprocessed offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the current parent feature point corresponding to the second video frame.

[0185] In some embodiments, step S104 may include the following steps: calculating the average of the offset value to be processed corresponding to the first video frame and the offset value to be processed corresponding to the second video frame, as the average offset value; calculating the sum of the average offset value and the target coordinates of the parent feature point corresponding to the video frame, to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0186] Alternatively, the electronic device can calculate the weighted sum of the offset value to be processed corresponding to the first video frame and the offset value to be processed corresponding to the second video frame, and calculate the sum of this sum and the target coordinates of the current parent feature point corresponding to the video frame to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0187] In some embodiments, prior to step S103, the method may further include the following steps: for each child feature point of the parent feature point, obtaining the initial coordinates and corresponding second confidence level of the child feature point in the target scene corresponding to the video frame. The second confidence level represents the probability that the child feature point is located at the position indicated by the initial coordinates in the target scene; the initial coordinates of the child feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the feature point in the video frame.

[0188] Accordingly, step S103 may include the following steps: for each child feature point of the parent feature point, if the second confidence level is less than a preset threshold, calculate the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed.

[0189] Accordingly, the method may further include the following steps: if the second confidence level is not less than a preset threshold, the initial coordinates of the sub-feature point corresponding to the video frame in the target scene are used as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

[0190] For each child feature point of the current parent feature point, the initial coordinates of that child feature point in this video frame are determined based on the two-dimensional coordinates of that child feature point in each video frame within the video frame group to which this video frame belongs. The video frame group to which this video frame belongs includes other video frames of the target scene with different perspectives from this video frame.

[0191] For each video frame, if the electronic device determines the initial coordinates of each feature point of the target object based on the two-dimensional coordinates of each feature point of the target object in each video frame within the video frame group to which the video frame belongs, the electronic device can input each video frame into a pre-trained pose recognition model to obtain the two-dimensional coordinates of each feature point of the target object in that video frame and the corresponding confidence score. Then, the electronic device can calculate the second confidence score based on the following method.

[0192] Accordingly, the steps of obtaining the initial coordinates and the corresponding second confidence level of the sub-feature point in the target scene corresponding to the video frame include: calculating the average confidence level of the two-dimensional coordinates of the sub-feature point in each video frame of the video frame group to which the video frame belongs, and using it as the second confidence level.

[0193] For each child feature point of the current parent feature point, the electronic device calculates the mean of the confidence scores corresponding to the two-dimensional coordinates of that child feature point in each video frame of the video frame group to which the current video frame belongs, as the second confidence score. The second confidence score can then represent the probability that the child feature point is located at the position indicated by the initial coordinates in the target scene.

[0194] If the second confidence level is not less than the preset threshold, that is, the probability that the sub-feature point is located at the position represented by the initial coordinates in the target scene is high, it indicates that the accuracy of the initial coordinates of the sub-feature point corresponding to the video frame is high. Then the electronic device can directly use the initial coordinates of the sub-feature point corresponding to the video frame in the target scene as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

[0195] If the second confidence level is less than the preset threshold, that is, the probability that the sub-feature point is located at the position represented by the initial coordinates in the target scene is low, it indicates that the accuracy of the initial coordinates of the sub-feature point corresponding to the video frame is low. The electronic device can obtain the target offset value of the sub-feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed.

[0196] Furthermore, the electronic device can calculate the target coordinates of the child feature point in the target scene based on the offset value to be processed and the target coordinates of the current parent feature point corresponding to the video frame.

[0197] In some embodiments, after step S104, the method may further include the following steps:

[0198] The process involves taking the child feature point of the target object as the parent feature point, and then returning to perform the following steps for each child feature point of the parent feature point: calculating the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point of the video frame to be processed, and using this offset value as the processing offset value, until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are determined.

[0199] After calculating the target coordinates of the child feature points of the current parent feature point in the target scene for the current video frame, for each child feature point of the current parent feature point, the electronic device can take the child feature point as the current parent feature point, and recalculate the offset value to be processed for each child feature point of the current parent feature point, and calculate the target coordinates of the child feature point in the target scene for the current video frame based on the offset value to be processed, and so on, until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are obtained.

[0200] In some embodiments, after calculating the target coordinates of each feature point of the target object corresponding to each video in the target scene, the electronic device can also use Savgol filtering to smooth the target coordinates of each feature point of the determined target object in order to improve the accuracy of the determined target coordinates.

[0201] Regarding step S105, the electronic device stores a preset image containing a virtual object that has not performed any action. After calculating the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, the electronic device can adjust the coordinates of each feature point of the virtual object in the preset image according to the target coordinates of each feature point of the target object corresponding to each video frame, thereby obtaining a target video in which the virtual object's actions are the same as those of the target object.

[0202] In some embodiments, the target object can be a person. After calculating the target coordinates of each feature point of the target object in the target scene corresponding to each video, the electronic device can also generate a 3D (three-dimensional) person mesh model in SMSLX (Skinned Multi-Person Linear Model) format based on the target coordinates of each feature point of the target object. Then, the electronic device can use the Blender tool to convert the 3D person mesh model into a BVH (a universal human feature animation file) format file and save the generated file.

[0203] In some embodiments, for each video frame in the original video, when the target coordinates of each sub-feature point of the target object corresponding to the video frame in the target scene are three-dimensional coordinates, the electronic device can calculate the initial coordinates of each feature point of the target object corresponding to the video frame based on the two-dimensional coordinates of each feature point of the target object in the video frame in the following manner.

[0204] Method 1,

[0205] Electronic devices can set up a camera in the target scene to capture images of the target scene. This camera can be an RGB-D (Red Green Blue-Deep) camera. By using this camera to capture video containing the target scene, the original video can be obtained.

[0206] For each video frame of the original video, the electronic device obtains the three-dimensional coordinates (i.e., initial coordinates) of each feature point of the target object in the target scene based on the two-dimensional coordinates of each feature point of the target object in the video frame, the depth information of each feature point of the target object, and the transformation relationship between the image coordinate system of the video frame and the three-dimensional coordinate system of the target scene.

[0207] Method 2,

[0208] In some embodiments, see Figure 4 , Figure 4 A flowchart of a method for determining the initial coordinates of feature points of a target object, provided in an embodiment of the present invention, is included. The method may include the following steps:

[0209] S401: Obtain the original video of the target scene from multiple perspectives, and extract multiple video frames with the same timestamp from each original video to obtain a video frame group.

[0210] S402: For each video frame in the video frame group, perform pose recognition on the video frame to obtain the two-dimensional coordinates of each feature point of the target object in the video frame.

[0211] S403: For each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, calculate the three-dimensional coordinates of the feature point in the target scene corresponding to each video frame in the video frame group, and use them as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0212] The original video in the foregoing embodiments can be any one of the original videos of the target scene from multiple perspectives.

[0213] Accordingly, multiple cameras can be set up at different locations in the target scene, and these multiple cameras can capture the target objects in the target scene. These multiple cameras can be RGB (Red, Green, Blue) cameras. For example, when the target scene is a stage, multiple cameras can be set up around the stage, ensuring that most cameras can capture the people on the stage.

[0214] For example, see Figure 5 The target scene is a stage. A checkerboard calibration board can be laid flat at the center of the stage. This board can be used to calibrate the intrinsic and extrinsic parameters of each camera around the stage. Consequently, each camera can clearly capture the people on the stage. For example, Figure 5 The images shown are of the target scene from different perspectives. Figure 5 Each of the images shown contains a clear and complete image of the checkerboard calibration board.

[0215] Furthermore, by simultaneously capturing images of the target scene from multiple perspectives using multiple cameras, multiple videos (i.e., original videos) are obtained. Since each camera has the same acquisition rate—meaning each camera captures one frame of the target scene at the same time—the number of video frames in each original video is consistent. For example, if each camera's acquisition rate is 30 FPS (30 frames per second), and the acquisition time for each camera is 2 seconds, multiple original videos containing 60 frames each can be obtained.

[0216] Correspondingly, the electronic device acquires the original videos of the target scene from multiple perspectives, and extracts multiple video frames with the same timestamp from each original video to obtain a video frame group.

[0217] For example, each original video includes: original videos of the target scene from three different perspectives, namely: original video 1, original video 2, and original video 3, each of which contains 50 video frames. The electronic device acquires the first frame from original video 1, the first frame from original video 2, and the first frame from original video 3 to obtain the first video frame group. Then, the electronic device acquires the second frame from original video 1, the second frame from original video 2, and the second frame from original video 3 to obtain the second video frame group, and so on, to obtain 50 video frame groups.

[0218] In some embodiments, step S402 may include the following steps: for each video frame in the video frame group, input the video frame into a pre-trained pose recognition model to obtain the two-dimensional coordinates and corresponding confidence scores of each feature point of the target object in the video frame. The confidence score corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame.

[0219] For each video frame in a video frame group, an electronic device can perform pose recognition on that video frame to obtain the two-dimensional coordinates of each feature point of the target object within that video frame. For example, the electronic device can input the video frame into a pre-trained pose recognition model to obtain the two-dimensional coordinates of each feature point of the target object within that video frame and the corresponding confidence level, as output by the pose recognition model. The pose recognition model can be a 2D pose recognition model provided by OpenPose.

[0220] The first confidence level and the second confidence level in the foregoing embodiments can be determined by the electronic device based on the two-dimensional coordinates of the corresponding feature point in each video frame of the video frame group to which the video frame belongs and the corresponding confidence level.

[0221] For example, referring to Figures 6(a) and 6(b), Figures 6(a) and 6(b) show four video frames in a video frame group, which are images of the target scene from different perspectives. Each video frame contains three target objects, and each rectangular region contains one target object. The four video frames are input into the pose recognition model to obtain the two-dimensional coordinates of the feature points of the multiple target objects in the four video frames.

[0222] For each video frame group, if each video frame in the video frame group contains a target object, the electronic device calculates the initial coordinates of each feature point of the target object in the target scene based on the two-dimensional coordinates of each feature point of the target object in each video frame in the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene.

[0223] If each video frame in the video frame group contains multiple target objects, the electronic device matches the multiple target objects in each video frame to obtain the same target objects in each video frame of the video frame group. Then, for each target object, the electronic device can determine the initial coordinates of each feature point of the target object in the target scene corresponding to the video frame group.

[0224] Accordingly, in some embodiments, in Figure 4 Based on this, see Figure 7 Before step S403, the method may further include the following steps:

[0225] S404: Based on the two-dimensional coordinates of feature points of multiple target objects in each video frame of the video frame group, determine the same target objects in each video frame of the video frame group.

[0226] For each group of video frames, the electronic device can determine the same target object in each video frame of that group in the following manner.

[0227] Method 1

[0228] For each target object, the epipolar distance from each feature point of the target object to the corresponding epipolar plane is calculated, and the mean of each epipolar distance corresponding to each feature point of the target object is calculated to obtain the mean distance for the target object. For every two video frames in the video frame group, based on the mean distances corresponding to every two target objects in the two video frames, the similarity between the two target objects is calculated to obtain the first similarity matrix.

[0229] Here, the epipolar plane corresponding to a feature point represents the plane in which that feature point lies within the target scene. An element in the first similarity matrix represents the probability that two corresponding target objects in the two video frames are identical.

[0230] For each video frame group, each video frame in the video frame group is an image of the target scene captured by multiple cameras from different perspectives. In other words, each video frame in the video frame group contains images of the target object from different perspectives.

[0231] According to the multi-view geometry principle in computer vision, the same feature point of the same target object lies within the same epipolar plane in the target scene from different viewpoints. Therefore, from different viewpoints, the same feature point of each target object corresponds to an epipolar plane, and the epipolar distance from the feature point to the epipolar plane indicates whether the feature point lies within the epipolar plane.

[0232] Accordingly, for each video frame in the video frame group, the electronic device selects a target object from that video frame. For each feature point of the target object, the electronic device calculates the straight line between the feature point and the optical center of the camera that acquired the video frame, based on the two-dimensional coordinates of the feature point in the video frame and the intrinsic and extrinsic parameters of the camera that acquired the video frame, thus obtaining the ray corresponding to the feature point.

[0233] Furthermore, the electronic device can calculate the intersection point of the rays corresponding to the feature point of each two target objects in the target scene, obtaining multiple intersection points. Then, based on the three-dimensional coordinates of each intersection point in the target scene, the plane equation of the epipolar plane corresponding to the feature point in the target scene is calculated.

[0234] For each video frame in the video frame group, the electronic device selects a target object from that video frame. For each feature point of the target object, the electronic device can calculate the epipolar distance from each feature point of the target object to the epipolar plane corresponding to that feature point, based on the two-dimensional coordinates of the feature point in the video frame, the transformation relationship between the image coordinate system of the video frame and the three-dimensional coordinate system of the target scene, and the plane equation of the epipolar plane corresponding to the feature point in the target scene.

[0235] Then, the electronic device calculates the mean distance of each epipolar line corresponding to each feature point of the target object, thus obtaining the mean distance of the target object. Furthermore, for every two video frames in the video frame group, the sum of the mean distances corresponding to each pair of target objects in those two video frames is calculated to obtain the similarity between the two target objects.

[0236] An element in the first similarity matrix represents the probability that two corresponding target objects in two video frames are the same. Accordingly, the electronic device determines the same target objects in each video frame within the video frame group based on the first similarity matrix.

[0237] For example, the electronic device can cluster multiple target objects in each video frame of the video frame group based on the first similarity matrix to obtain the same target objects in each video frame of the video frame group. Alternatively, the electronic device can calculate the same target objects in each video frame of the video frame group based on the Hungarian algorithm and the first similarity matrix.

[0238] Method 2

[0239] For each pair of video frames in this video frame group, the electronic device can input these two video frames into a pre-trained object matching model to obtain the similarity between each pair of target objects in the two video frames, thus obtaining a second similarity matrix. The object matching model can be a network model based on ReID (Person Re-Identification) technology.

[0240] An element in the second similarity matrix represents the probability that two corresponding target objects in the two video frames are the same. Accordingly, the electronic device uses the second similarity matrix to determine the same target objects in each video frame within the video frame group.

[0241] The way in which an electronic device determines the same target object in each video frame of a video frame group based on a second similarity matrix is ​​similar to the way in which an electronic device determines the same target object in each video frame of a video frame group based on a first similarity matrix, and can be referred to the relevant description in the foregoing embodiments.

[0242] Method 3

[0243] To improve the accuracy of identifying the same target object in each video frame within the identified video frame group, Figure 7 Based on this, see Figure 8 Step S404 may include the following steps:

[0244] S4041: For each target object, calculate the epipolar distance from each feature point of the target object to the corresponding epipolar plane, and calculate the mean value of each epipolar distance corresponding to each feature point of the target object to obtain the mean value of the distance corresponding to the target object.

[0245] In this context, the epipolar plane corresponding to a feature point represents the plane in which that feature point is located in the target scene.

[0246] S4042: For every two video frames in the video frame group, calculate the similarity between the two target objects based on the average distance between the two target objects in the two video frames, and obtain the first similarity matrix.

[0247] In the first similarity matrix, an element represents the probability that the two target objects in the two video frames are the same.

[0248] S4043: Input the two video frames into the pre-trained object matching model to obtain the similarity between each pair of target objects in the two video frames, and obtain the second similarity matrix.

[0249] In the second similarity matrix, an element represents the probability that the two target objects in the two video frames are the same.

[0250] S4044: Merge the first similarity matrix and the second similarity matrix to obtain the target similarity matrix.

[0251] S4045: Based on the target similarity matrix, identify the same target objects in each video frame of the video frame group.

[0252] The method by which the electronic device obtains the first similarity matrix and the second similarity matrix can be referred to the relevant description in the foregoing embodiments.

[0253] After obtaining the first similarity matrix and the second similarity matrix, the electronic device can fuse the first similarity matrix and the second similarity matrix to obtain the target similarity matrix.

[0254] In one implementation, the electronic device can calculate the weighted sum of each element in the first similarity matrix and the corresponding element in the second similarity matrix to obtain the target similarity matrix.

[0255] In another implementation, the electronic device can fuse each element in the first similarity matrix with the corresponding element in the second similarity matrix according to the following formula (1) to obtain the target similarity matrix.

[0256]

[0257] A i,j This represents the element in the i-th row and j-th column of the target similarity matrix; a i,j b represents the element in the i-th row and j-th column of the first similarity matrix; i,j w1 represents the element in the i-th row and j-th column of the second similarity matrix; w2 represents the weight of the element in the i-th row and j-th column of the first similarity matrix; w3 represents the weight of the element in the i-th row and j-th column of the second similarity matrix.

[0258] For example, see Figure 9 , Figure 9 This is a schematic diagram of a target similarity matrix provided in an embodiment of the present invention. Figure 9 The target similarity matrix shown corresponds to 11 target objects horizontally and 11 target objects vertically. The rectangular area corresponding to each pair of target objects represents the similarity between the two target objects. The darker the color of the rectangular area, the higher the similarity between the two target objects. Furthermore, based on the target similarity matrix, the electronic device identifies the same target objects in each video frame within the video frame group.

[0259] The way in which an electronic device determines the same target object in each video frame of a video frame group based on a target similarity matrix is ​​similar to the way in which an electronic device determines the same target object in each video frame of a video frame group based on a first similarity matrix. For details, please refer to the relevant descriptions in the foregoing embodiments.

[0260] In some embodiments, step S403 may include the following steps: for each feature point of the target object, based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group, and the transformation relationships between the image coordinate system of the two video frames and the three-dimensional coordinate system of the target scene, calculate the three-dimensional coordinates of the feature point in the target scene as the three-dimensional coordinates to be processed; calculate the average value of each three-dimensional coordinate to be processed as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0261] For each feature point of the target object, the electronic device calculates the three-dimensional coordinates of the feature point in the target scene based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group and the following formula (2).

[0262]

[0263] (um v m () represents the two-dimensional coordinates of the m-th feature point of the target object in a video frame, K represents the intrinsic parameters of the camera that acquired the video frame, and [R|t] represents the extrinsic parameters of the camera that acquired the video frame. m Y m Z m ) represents the three-dimensional coordinates of the m-th feature point of the target object in the target scene.

[0264] The intrinsic and extrinsic parameters of the camera that acquired the video frame represent the transformation relationship between the image coordinate system of the video frame and the three-dimensional coordinate system of the target scene.

[0265] The electronic device uses the two-dimensional coordinates of the feature point of the target object in every two video frames of the video frame group as (u) in the above formula (2). m v m ), and by taking the intrinsic and extrinsic parameters of the camera that captured the two video frames as K[R|t] in the above formula (2), we can obtain two parameters about (X). m Y m Z m By solving the system of equations, we can obtain the three-dimensional coordinates of the m-th feature point in the target scene.

[0266] Based on the two-dimensional coordinates of the feature point in every two video frames of the video frame group, multiple three-dimensional coordinates of the feature point can be obtained. Then, the electronic device calculates the average of these three-dimensional coordinates to obtain the initial coordinates of the feature point in the target scene corresponding to the video frame group. These initial coordinates of the feature point in the target scene are also the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0267] Based on the above processing, the target coordinates of each feature point of the target object in the target scene can be determined by using the two-dimensional coordinates of the target object in each video frame from multiple perspectives, and the transformation relationships between the image coordinate system of each video frame and the three-dimensional coordinate system of the target scene. The target coordinates of each feature point of a target object in the target scene can represent the three-dimensional pose of the target object, meaning that the three-dimensional pose of the target object can be determined without the target object wearing motion capture equipment. Furthermore, generating target videos based on the three-dimensional pose of the target object can reduce the time and labor costs of video generation and improve the efficiency of video generation.

[0268] For example, the target object is a person. See Figure 10(a), which is a schematic diagram of the joints of a target object provided by an embodiment of the present invention. From left to right, the images in Figure 10(a) are schematic diagrams of the joints of the target object corresponding to each video frame. The target coordinates of the missing joints of the target object in the second, third, and fourth images from left to right in Figure 10(a) can be determined according to the method provided by the embodiment of the present invention, resulting in the schematic diagram of the joints of the target object shown in Figure 10(d).

[0269] For example, for the image in the middle of Figure 10(a), a joint point can be selected from the joint points of the target object as the current parent feature point, and the target coordinates of the current parent feature point and the target coordinates of each child feature point of the current parent feature point can be determined according to the method provided in the embodiments of the present invention, thereby obtaining a schematic diagram of the joint points of the target object as shown in Figure 10(b).

[0270] Then, the child feature points of the current parent feature point are taken as the current parent feature point, and the target coordinates of the current parent feature point and the target coordinates of each child feature point of the current parent feature point are determined according to the method provided in the embodiment of the present invention. Thus, a schematic diagram of the joint points of the target object as shown in Figure 10(c) can be obtained. This process is repeated until the target coordinates of each joint point of the target object are obtained, thus a schematic diagram of the joint points of the target object as shown in Figure 10(d) can be obtained.

[0271] See Figure 11 , Figure 11 A flowchart illustrating a method for determining the target coordinates of a target object, as provided in an embodiment of the present invention.

[0272] S1101: Multi-camera calibration.

[0273] The electronic device calibrates multiple cameras at different locations in the target scene to determine the intrinsic and extrinsic parameters of the multiple cameras, and then uses these multiple cameras to acquire original videos of the target scene from different perspectives.

[0274] S1102: 2D (two-dimensional) pose recognition.

[0275] The electronic device acquires multiple video frames with the same timestamp from various original videos of the target scene from multiple perspectives, resulting in multiple video frame groups. The target object is a person. For each video frame group, the electronic device performs 2D pose recognition on each video frame in that group to obtain the two-dimensional coordinates of each key point of the target object within that video frame.

[0276] S1103: Multi-view character selection.

[0277] For each video frame group, if each video frame in the video frame group contains multiple target objects, the electronic device performs multi-view character selection, that is, the electronic device matches multiple target objects in each video frame to obtain the same target object in each video frame in the video frame group.

[0278] S1104: 3D pose generation.

[0279] For each target object, the electronic device performs 3D pose generation on the target object. That is, based on the two-dimensional coordinates of each joint of the target object in each video frame of the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, the initial coordinates of each joint of the target object in the target scene corresponding to the video frame group are calculated.

[0280] S1105: Frame interpolation smoothing.

[0281] After determining the initial coordinates of each joint of the target object in the target scene for each video frame group, the electronic device determines the target coordinates of each joint of the target object based on the initial coordinates of each joint. Then, the electronic device uses savgol filtering to smooth the determined target coordinates of each joint of the target object, obtaining the final target coordinates of each joint of the target object.

[0282] S1106: Save action.

[0283] After determining the target coordinates of each joint of the target object in the target scene for each video frame group, the electronic device can generate a 3D character mesh model in SMSLX format based on the target coordinates of each joint of the target object. Then, the electronic device can use the Blender tool to convert the 3D character mesh model into a BVH format file and save the generated file.

[0284] Based on the above processing, the target coordinates of each joint of the target object in the target scene can be determined by using the two-dimensional coordinates of the target object in each video frame from multiple perspectives, and the transformation relationships between the image coordinate system of each video frame and the three-dimensional coordinate system of the target scene. The target coordinates of each joint of a target object in the target scene can represent the three-dimensional pose of the target object, meaning that the three-dimensional pose of the target object can be determined without the target object wearing motion capture equipment. Furthermore, generating target videos based on the three-dimensional pose of the target object can reduce the time and labor costs of video generation and improve the efficiency of video generation.

[0285] See Figure 12 , Figure 12A flowchart of another method for determining the target coordinates of a target object provided in an embodiment of the present invention.

[0286] S1201: Average interpolation of the central node.

[0287] The target object is a person, and the feature points of the target object are also the joint points of the target object. The central node is the central feature point in this embodiment of the invention. The electronic device performs central node average interpolation, which means that the electronic device determines the specified central joint point from the joint points of the target object. For each video frame, the electronic device obtains the first confidence level of the central joint point corresponding to that video frame. If the first confidence level is less than a preset threshold, the average of the target coordinates of the central joint point corresponding to the third video frame and the target coordinates of the central joint point corresponding to the fourth video frame is calculated to obtain the target coordinates of the central node corresponding to that video frame in the target scene. If the first confidence level is not less than the preset threshold, the electronic device obtains the initial coordinates of the central joint point corresponding to that video frame to obtain the target coordinates of the central joint point corresponding to that video frame in the target scene.

[0288] S1202: Calculate the relative position offset between the unknown node and its parent node.

[0289] The electronic device identifies the central joint point from all joint points of the target object as the current parent joint point, which is also the current parent feature point. Unknown nodes are child joint points of the current parent joint point, which are also child feature points of the current parent feature point. For each child joint point of the current parent joint point, the electronic device calculates the second confidence score corresponding to the two-dimensional coordinates of that child joint point in each video frame within the video frame group to which the video frame belongs. If the second confidence score is less than a preset threshold, the device calculates the processing offset value of the target coordinates of the child joint point corresponding to the video frame to be processed relative to the target coordinates of the parent joint point corresponding to the video frame to be processed.

[0290] S1203: Calculate the current node position.

[0291] The current node is also the child node of the current parent node. Based on the offset value to be processed and the target coordinates of the parent node corresponding to the video frame, the electronic device calculates the target coordinates of the child node in the target scene.

[0292] Then, the electronic device takes the current node as the parent node and continues to calculate the relationship between its child nodes and the parent node, which is the current parent joint. In other words, the electronic device takes the child joint of the target object as the current parent joint and calculates the second confidence level corresponding to each child joint of the current parent joint again, and calculates the target coordinates of each child joint of the current parent joint until the target coordinates of each joint of the target object in the target scene corresponding to the video frame are obtained.

[0293] S1204: Savgol filter.

[0294] After calculating the target coordinates of each joint of the target object in the target scene corresponding to each video, the electronic device uses Savgol filtering to smooth the target coordinates of each joint of the determined target object in order to improve the accuracy of the determined target coordinates.

[0295] S1205: Calculation results.

[0296] The electronic device uses Savgol filtering to smooth the target coordinates of each joint of the identified target object, and obtains the final calculation result, which is the target coordinates of each joint of the target object corresponding to each video frame.

[0297] Based on the above processing, for each child feature point of the parent feature point of the target object, the offset value to be processed is the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the child feature point corresponding to the video frame to be processed. Since the parent feature point and the child feature point of the target object are connected, it means that the offset value of the child feature point relative to the parent feature point is fixed in each video frame. Accordingly, based on the offset value to be processed and the target coordinate of the parent feature point corresponding to the video frame, the target coordinate of the child feature point corresponding to the video frame in the target scene can be calculated, which can improve the accuracy of the calculated target coordinates.

[0298] and Figure 1 For the corresponding method implementation examples, see [link to relevant documentation]. Figure 13 , Figure 13 This is a structural diagram of an image generation apparatus provided in an embodiment of the present invention. The apparatus includes:

[0299] The selection module 1301 is used to select one feature point from the feature points of the target object in the original video as the parent feature point.

[0300] The first acquisition module 1302 is used to acquire the target coordinates of the parent feature point corresponding to each video frame in the original video in the target scene.

[0301] The first determining module 1303 is used to calculate, for each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, as the offset value to be processed; wherein, the child feature points of the parent feature point include: feature points connected to the parent feature point among the feature points of the target object; the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or a second video frame whose timestamp is greater than the timestamp of the video frame;

[0302] The second determining module 1304 is used to calculate the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame.

[0303] The generation module 1305 is used to adjust the coordinates of each feature point of the virtual object in each preset image according to the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, so as to obtain a target video in which the virtual object has the same action as the target object.

[0304] Optionally, the device further includes:

[0305] The processing module is configured to, after the second determining module 1304 performs the following steps: based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, calculate the target coordinates of the child feature point corresponding to the video frame in the target scene; take the child feature point of the target object as the parent feature point; and return to perform the following steps: for each child feature point of the parent feature point, calculate the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed, until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are determined.

[0306] Optionally, the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and a second video frame whose timestamp is greater than the timestamp of the video frame;

[0307] The first determining module 1303 is specifically used to calculate, for each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the first video frame relative to the target coordinate of the parent feature point corresponding to the first video frame, and use it as the offset value to be processed for the first video frame.

[0308] Calculate the offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the parent feature point corresponding to the second video frame, and use it as the offset value to be processed for the second video frame.

[0309] Optionally, the second determining module 1304 is specifically used to calculate the average of the offset value to be processed corresponding to the first video frame and the offset value to be processed corresponding to the second video frame, as the average offset value.

[0310] The sum of the average offset value and the target coordinates of the parent feature point corresponding to the video frame is calculated to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

[0311] Optionally, the selection module 1301 is specifically used to determine a specified central feature point from the feature points of the target object in the original video, as the parent feature point;

[0312] The first acquisition module 1302 is specifically used to acquire the initial coordinates and the corresponding first confidence level of the central feature point corresponding to the video frame; wherein, the first confidence level represents the probability that the central feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame;

[0313] If the first confidence level is less than a preset threshold, the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame is calculated to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, which is used as the target coordinates of the parent feature point corresponding to the video frame in the target scene; wherein, the third video frame includes: video frames in the original video with timestamps less than the timestamp of the video frame; the fourth video frame includes: video frames in the original video with timestamps greater than the timestamp of the video frame;

[0314] If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

[0315] Optionally, the device further includes:

[0316] The second acquisition module is configured to, before the first determination module 1303 performs the following steps for each sub-feature point of the parent feature point: calculating the offset value of the target coordinates of the sub-feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point of the video frame to be processed, and using this offset value as the processing offset value, perform the following steps for each sub-feature point of the parent feature point: acquiring the initial coordinates and corresponding second confidence level of the sub-feature point corresponding to the video frame in the target scene; wherein, the second confidence level represents the probability that the sub-feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the sub-feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the feature point in the video frame;

[0317] The first determining module 1303 is specifically used to calculate the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed if the second confidence level is less than a preset threshold, and use it as the offset value to be processed.

[0318] The device further includes:

[0319] The third determining module is used to, if the second confidence level is not less than the preset threshold, take the initial coordinates of the sub-feature point corresponding to the video frame in the target scene as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

[0320] Optionally, the device further includes:

[0321] The third acquisition module is used to acquire the original video of the target scene from multiple perspectives before the first acquisition module 1302 acquires the target coordinates of the parent feature point corresponding to each video frame in the original video. The module acquires multiple video frames with the same timestamp from each original video to obtain a video frame group.

[0322] The recognition module is used to perform pose recognition on each video frame in the video frame group to obtain the two-dimensional coordinates of each feature point of the target object in the video frame.

[0323] The fourth determining module is used to perform pose recognition on each video frame in the video frame group to obtain the two-dimensional coordinates of each feature point of the target object in the video frame.

[0324] Optionally, the fourth determining module is specifically used to, for each feature point of the target object, calculate the three-dimensional coordinates of the feature point in the target scene based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group, and the transformation relationships between the image coordinate system of the two video frames and the three-dimensional coordinate system of the target scene, as the three-dimensional coordinates to be processed; and calculate the average value of each three-dimensional coordinate to be processed as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

[0325] Optionally, the recognition module is specifically used to input each video frame in the video frame group into a pre-trained pose recognition model to obtain the two-dimensional coordinates and corresponding confidence scores of each feature point of the target object in the video frame; wherein, the confidence score corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame;

[0326] The second acquisition module is specifically used to calculate the average confidence level of the two-dimensional coordinates of the sub-feature point in each video frame of the video frame group to which the video frame belongs, and use it as the second confidence level.

[0327] The first acquisition module 1302 is specifically used to calculate the average confidence level of the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs, and use it as the first confidence level.

[0328] Optionally, the device further includes:

[0329] The matching module is used to perform the following steps in the fourth determining module: for each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, calculate the three-dimensional coordinates of the feature point in each video frame of the video frame group in the target scene, and use these coordinates as the initial coordinates of the feature point in each video frame of the video frame group in the target scene. Before this, the module performs the following steps: based on the two-dimensional coordinates of the feature points of multiple target objects in each video frame of the video frame group, determine the same target objects in each video frame of the video frame group.

[0330] Optionally, the matching module is specifically used to calculate the epipolar distance from each feature point of the target object to the corresponding epipolar plane for each target object, and to calculate the average value of each epipolar distance corresponding to each feature point of the target object, so as to obtain the average distance corresponding to the target object; wherein, the epipolar plane corresponding to a feature point represents the plane in which the feature point is located in the target scene;

[0331] For every two video frames in the video frame group, the similarity between the two target objects is calculated based on the average distance between them, resulting in a first similarity matrix. An element in the first similarity matrix represents the probability that the two target objects in the two video frames are the same.

[0332] The two video frames are input into a pre-trained object matching model to obtain the similarity between every two target objects in the two video frames, resulting in a second similarity matrix; where each element in the second similarity matrix represents the probability that the two corresponding target objects in the two video frames are the same.

[0333] The first similarity matrix and the second similarity matrix are fused to obtain the target similarity matrix;

[0334] Based on the target similarity matrix, the same target objects in each video frame of the video frame group are identified.

[0335] Based on the image generation apparatus provided in this embodiment of the invention, for each child feature point of the parent feature point of the target object, the offset value to be processed is the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the child feature point corresponding to the video frame to be processed. Since the parent feature point and the child feature point of the target object are connected, it means that the offset value of the child feature point relative to the parent feature point is fixed in each video frame. Accordingly, based on the offset value to be processed and the target coordinate of the parent feature point corresponding to the video frame, the target coordinate of the child feature point corresponding to the video frame in the target scene is calculated, which can improve the accuracy of the calculated target coordinates.

[0336] This invention also provides an electronic device, such as... Figure 14 As shown, it includes a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404, wherein the processor 1401, the communication interface 1402, and the memory 1403 communicate with each other through the communication bus 1404.

[0337] Memory 1403 is used to store computer programs;

[0338] When the processor 1401 executes the program stored in the memory 1403, it implements the image generation method steps described in any of the above embodiments.

[0339] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0340] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0341] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0342] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0343] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the image generation methods described in the above embodiments.

[0344] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the image generation methods described in the above embodiments.

[0345] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0346] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0347] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0348] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An image generation method, characterized in that, The method includes: Select one feature point from the feature points of the target object in the original video as the parent feature point; For each video frame in the original video, obtain the target coordinates of the parent feature point corresponding to that video frame in the target scene; For each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed is calculated as the offset value to be processed; wherein, the child feature points of the parent feature point include: feature points connected to the parent feature point among the feature points of the target object; the video frame to be processed includes: the first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or, the second video frame whose timestamp is greater than the timestamp of the video frame; Based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, calculate the target coordinates of the child feature point corresponding to the video frame in the target scene; According to the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, adjust the coordinates of each feature point of the virtual object in each preset image to obtain a target video in which the virtual object has the same action as the target object; The step of selecting a feature point from the feature points of the target object in the original video as the parent feature point includes: From the feature points of the target object in the original video, determine the specified central feature point as the parent feature point; The step of obtaining the target coordinates of the parent feature point in the target scene for each video frame in the original video includes: Obtain the initial coordinates and corresponding first confidence level of the central feature point corresponding to the video frame; wherein, the first confidence level represents the probability that the central feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame; If the first confidence level is less than a preset threshold, the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame is calculated to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, which is used as the target coordinates of the parent feature point corresponding to the video frame in the target scene; wherein, the third video frame includes: video frames in the original video with timestamps less than the timestamp of the video frame; the fourth video frame includes: video frames in the original video with timestamps greater than the timestamp of the video frame; If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

2. The method according to claim 1, characterized in that, After calculating the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame, the method further includes: The process of taking the sub-feature point of the target object as the parent feature point and returning to perform the steps of calculating the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point for each sub-feature point of the parent feature point, and using it as the offset value to be processed, continues until the target coordinates of each feature point of the target object corresponding to the video frame in the target scene are determined.

3. The method according to claim 1, characterized in that, The video frames to be processed include: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and a second video frame whose timestamp is greater than the timestamp of the video frame; For each child feature point of the parent feature point, the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed is calculated as the processing offset value, including: For each child feature point of the parent feature point, calculate the offset value of the target coordinate of the child feature point corresponding to the first video frame relative to the target coordinate of the parent feature point corresponding to the first video frame, and use it as the offset value to be processed for the first video frame. Calculate the offset value of the target coordinates of the sub-feature point corresponding to the second video frame relative to the target coordinates of the parent feature point corresponding to the second video frame, and use it as the offset value to be processed for the second video frame.

4. The method according to claim 3, characterized in that, The step of calculating the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame includes: Calculate the average value of the offset to be processed corresponding to the first video frame and the offset to be processed corresponding to the second video frame, and use it as the average offset value. The sum of the average offset value and the target coordinates of the parent feature point corresponding to the video frame is calculated to obtain the target coordinates of the child feature point corresponding to the video frame in the target scene.

5. The method according to claim 1, characterized in that, Before calculating the offset value of the target coordinates of the sub-feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point for each sub-feature point of the parent feature point, and using this offset value as the processing offset value, the method further includes: For each child feature point of the parent feature point, obtain the initial coordinates and corresponding second confidence level of the child feature point in the target scene corresponding to the video frame; wherein, the second confidence level represents the probability that the child feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the child feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the feature point in the video frame; For each child feature point of the parent feature point, the offset value of the target coordinates of the child feature point corresponding to the video frame to be processed relative to the target coordinates of the parent feature point corresponding to the video frame to be processed is calculated as the processing offset value, including: If the second confidence level is less than a preset threshold, calculate the offset value of the target coordinate of the sub-feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, and use it as the offset value to be processed. The method further includes: If the second confidence level is not less than the preset threshold, the initial coordinates of the sub-feature point corresponding to the video frame in the target scene are taken as the target coordinates of the sub-feature point corresponding to the video frame in the target scene.

6. The method according to claim 5, characterized in that, Before obtaining the target coordinates of the parent feature point in the target scene for each video frame in the original video, the method further includes: Obtain the original videos of the target scene from multiple perspectives, and extract multiple video frames with the same timestamp from each original video to obtain a video frame group; For each video frame in the video frame group, pose recognition is performed on the video frame to obtain the two-dimensional coordinates of each feature point of the target object in the video frame. For each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group, and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, the three-dimensional coordinates of the feature point corresponding to each video frame in the video frame group in the target scene are calculated, and used as the initial coordinates of the feature point corresponding to each video frame in the video frame group in the target scene.

7. The method according to claim 6, characterized in that, For each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, the three-dimensional coordinates of the feature point corresponding to each video frame in the video frame group are calculated in the target scene, and these coordinates are used as the initial coordinates of the feature point corresponding to each video frame in the video frame group in the target scene, including: For each feature point of the target object, based on the two-dimensional coordinates of the feature point in every two video frames in the video frame group, and the transformation relationships between the image coordinate system of the two video frames and the three-dimensional coordinate system of the target scene, the three-dimensional coordinates of the feature point in the target scene are calculated as the three-dimensional coordinates to be processed; the average value of each three-dimensional coordinate to be processed is calculated as the initial coordinates of the feature point in the target scene corresponding to each video frame in the video frame group.

8. The method according to claim 6, characterized in that, For each video frame in the video frame group, pose recognition is performed on that video frame to obtain the two-dimensional coordinates of each feature point of the target object in that video frame, including: For each video frame in the video frame group, the video frame is input into a pre-trained pose recognition model to obtain the two-dimensional coordinates and corresponding confidence scores of each feature point of the target object in the video frame; wherein, the confidence score corresponding to the two-dimensional coordinates of a feature point represents the probability that the feature point is located at the position represented by the two-dimensional coordinates in the video frame; The step of obtaining the second confidence score corresponding to the initial coordinates of the sub-feature point in the target scene of the video frame includes: The mean of the confidence scores corresponding to the two-dimensional coordinates of the sub-feature point in each video frame of the video frame group to which the video frame belongs is calculated and used as the second confidence score. The step of obtaining the first confidence score corresponding to the initial coordinates of the center feature point of the video frame includes: The mean confidence score of the two-dimensional coordinates of the central feature point in each video frame of the video frame group to which the video frame belongs is calculated and used as the first confidence score.

9. The method according to claim 6, characterized in that, Before calculating the three-dimensional coordinates of the feature point in the target scene for each feature point of the target object, based on the two-dimensional coordinates of the feature point in each video frame of the video frame group and the transformation relationships between the image coordinate system of each video frame in the video frame group and the three-dimensional coordinate system of the target scene, and using these coordinates as the initial coordinates of the feature point in the target scene for each video frame of the video frame group, the method further includes: Based on the two-dimensional coordinates of feature points of multiple target objects in each video frame of the video frame group, the same target objects in each video frame of the video frame group are determined.

10. The method according to claim 9, characterized in that, The method of determining the same target object in each video frame of the video frame group based on the two-dimensional coordinates of feature points of multiple target objects in each video frame includes: For each target object, the epipolar distance from each feature point of the target object to the corresponding epipolar plane is calculated, and the mean value of each epipolar distance corresponding to each feature point of the target object is calculated to obtain the mean distance corresponding to the target object; wherein, the epipolar plane corresponding to a feature point represents the plane in which the feature point is located in the target scene; For every two video frames in the video frame group, the similarity between the two target objects is calculated based on the average distance between them, resulting in a first similarity matrix. An element in the first similarity matrix represents the probability that the two target objects in the two video frames are the same. The two video frames are input into a pre-trained object matching model to obtain the similarity between every two target objects in the two video frames, resulting in a second similarity matrix; where each element in the second similarity matrix represents the probability that the two corresponding target objects in the two video frames are the same. The first similarity matrix and the second similarity matrix are fused to obtain the target similarity matrix; Based on the target similarity matrix, the same target objects in each video frame of the video frame group are identified.

11. An image generation apparatus, characterized in that, The device includes: The selection module is used to select one feature point from the feature points of the target object in the original video as the parent feature point; The first acquisition module is used to acquire the target coordinates of the parent feature point in the target scene for each video frame in the original video. The first determining module is used to calculate, for each child feature point of the parent feature point, the offset value of the target coordinate of the child feature point corresponding to the video frame to be processed relative to the target coordinate of the parent feature point corresponding to the video frame to be processed, as the offset value to be processed; wherein, the child feature points of the parent feature point include: feature points connected to the parent feature point among the feature points of the target object; the video frame to be processed includes: a first video frame in the original video whose timestamp is less than the timestamp of the video frame, and / or a second video frame whose timestamp is greater than the timestamp of the video frame; The second determining module is used to calculate the target coordinates of the child feature point corresponding to the video frame in the target scene based on the offset value to be processed and the target coordinates of the parent feature point corresponding to the video frame. The generation module is used to adjust the coordinates of each feature point of the virtual object in each preset image according to the target coordinates of each feature point of the target object in the target scene corresponding to each video frame, so as to obtain a target video in which the virtual object has the same action as the target object; The selection module is specifically used to determine a specified central feature point from the feature points of the target object in the original video, and use it as the parent feature point. The first acquisition module is specifically used to acquire the initial coordinates and the corresponding first confidence level of the central feature point corresponding to the video frame; wherein, the first confidence level represents the probability that the central feature point is located at the position represented by the initial coordinates in the target scene; the initial coordinates of the parent feature point corresponding to the video frame are determined based on the two-dimensional coordinates of the parent feature point in the video frame; If the first confidence level is less than a preset threshold, the average of the target coordinates of the center feature point corresponding to the third video frame and the target coordinates of the center feature point corresponding to the fourth video frame is calculated to obtain the target coordinates of the center feature point corresponding to the video frame in the target scene, which is used as the target coordinates of the parent feature point corresponding to the video frame in the target scene; wherein, the third video frame includes: video frames in the original video with timestamps less than the timestamp of the video frame; the fourth video frame includes: video frames in the original video with timestamps greater than the timestamp of the video frame; If the first confidence level is not less than the preset threshold, the initial coordinates of the center feature point corresponding to the video frame are used as the target coordinates of the parent feature point corresponding to the video frame in the target scene.

12. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-10.