Image Processing Method, Apparatus, Storage Medium, and Terminal

By identifying the two-dimensional position and depth of the key points of the object in the target image, and combining camera parameters, the three-dimensional position is directly determined from an image, the problem of complex and inaccurate three-dimensional human posture detection in the prior art is solved, and efficient and accurate three-dimensional position determination is achieved.

CN113470112BActive Publication Date: 2025-07-08GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110748267.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-07-08
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

In the prior art, the three-dimensional human posture detection method requires multiple images, and the detection process is complicated and inaccurate.

Method used

By identifying the two-dimensional position information and relative depth of the object key points in the target image in the image two-dimensional coordinate system, combined with preset camera internal and external parameters, the position information of the target object and model key points in the world three-dimensional coordinate system is directly determined from an image.

Benefits of technology

The complexity of three-dimensional position determination is reduced, the accuracy of detection is improved, and the three-dimensional position of the target model corresponding to the target object in space can be accurately determined through an image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113470112B_ABST
    Figure CN113470112B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, storage medium, and terminal, relating to the technical field of image processing. Determine the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image and the relative depth relative to the reference point in the target object; according to the two-dimensional position information and the relative depth, determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system; and then determine the second three-dimensional position information of each model key point in the target model corresponding to the target object. Since the first three-dimensional position information of each object key point in the target object is first determined, and then the second three-dimensional position information of each model key point in the target model corresponding to the target object is determined according to the two-dimensional position information and the first three-dimensional position information of each object key point in the target object, it is possible to determine the three-dimensional position information of the target model corresponding to the target object through one image, reducing the complexity of determining the three-dimensional position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular, to an image processing method, apparatus, storage medium, and terminal. Background Art

[0002] With the development of science and technology, the uses of images are becoming more and more extensive. For example, three-dimensional human postures can be recognized in images, which are widely used in fields such as security, gaming, and entertainment.

[0003] In related technologies, the three-dimensional human pose detection method usually obtains multiple two-dimensional position information from multiple images with different angles, and then converts the two-dimensional position information into three-dimensional position information according to the pre-determined relationship between the two-dimensional position information and the three-dimensional position information.

[0004] In related technologies, the current three-dimensional human pose detection method requires multiple images, and the detection method is relatively complex. Summary of the Invention

[0005] The present application provides an image processing method, apparatus, storage medium, and terminal, which can solve the technical problem that the three-dimensional human pose detection in related technologies is relatively complex.

[0006] In a first aspect, an embodiment of the present application provides an image processing method, which includes:

[0007] Identify a target object in a target image, and determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object;

[0008] Determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth;

[0009] Determine the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the first three-dimensional position information.

[0010] In a second aspect, an embodiment of the present application provides an image processing apparatus, which includes:

[0011] An object two-dimensional position determination module, configured to identify a target object in a target image, and determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object;

[0012] The object three-dimensional position determination module is used to determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth;

[0013] The model three-dimensional position determination module is used to determine the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the first three-dimensional position information.

[0014] In a third aspect, an embodiment of the present application provides a computer storage medium, which stores multiple operations, and the operations are adapted to be loaded and executed by a processor to perform the steps of the above method.

[0015] In a fourth aspect, an embodiment of the present application provides a terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor.

[0016] The beneficial effects brought by the technical solutions provided by some embodiments of the present application at least include:

[0017] The present application provides an image processing method, which identifies a target object in a target image, determines the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object; according to the two-dimensional position information and the relative depth, determines the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system; according to the two-dimensional position information and the first three-dimensional position information, determines the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system. Since the two-dimensional position information and the relative depth of each object key point in the target object can be determined based on a single target image, the first three-dimensional position information of each object key point in the target object can be determined first, and then, according to the two-dimensional position information and the first three-dimensional position information of each object key point in the target object, the second three-dimensional position information of each model key point in the target model corresponding to the target object can be determined, so that the three-dimensional position information of the target model corresponding to the target object can be determined through a single image, reducing the complexity of determining the three-dimensional position and increasing the accuracy of determining the three-dimensional position. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0019] Figure 1Exemplary system architecture diagram of an image processing method provided by an embodiment of this application;

[0020] Figure 2 System interaction diagram of an image processing method provided by an embodiment of this application;

[0021] Figure 3 Schematic flow diagram of an image processing method provided by another embodiment of this application;

[0022] Figure 4 Schematic flow diagram of an image processing method provided by another embodiment of this application;

[0023] Figure 5 Schematic diagram of a coordinate system reference provided by another embodiment of this application;

[0024] Figure 6 Schematic diagram of obtaining a coordinate probability distribution provided by another embodiment of this application;

[0025] Figure 7 Schematic diagram when the target object is in a standard posture provided by another embodiment of this application;

[0026] Figure 8 Schematic diagram of target model driving provided by another embodiment of this application;

[0027] Figure 9 Schematic diagram of target model driving provided by another embodiment of this application;

[0028] Figure 10 Schematic diagram of target model driving provided by another embodiment of this application;

[0029] Figure 11 Schematic diagram of target model driving provided by another embodiment of this application

[0030] Figure 12 Schematic diagram of the structure of an image processing device provided by another embodiment of this application;

[0031] Figure 13 Schematic diagram of the structure of an image processing device provided by another embodiment of this application;

[0032] Figure 14 Schematic diagram of the data flow of an image processing device provided by another embodiment of this application;

[0033] Figure 15 Schematic diagram of the structure of a terminal provided by an embodiment of this application. Detailed implementation manners

[0034] To make the features and advantages of this application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of this application.

[0035] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0036] Figure 1 It is an exemplary system architecture diagram of an image processing method provided for the embodiments of this application.

[0037] As Figure 1 shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a medium for the communication link between the terminal 101 and the server 103. The network 102 may include various types of wired communication links or wireless communication links. For example, the wired communication links include optical fibers, twisted pairs, or coaxial cables, and the wireless communication links include Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.

[0038] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various electronic devices, including but not limited to smart watches, smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal 101 is software, it can be installed in the above-listed electronic devices, which can be implemented as multiple software or software modules (for example, used to provide distributed services), or can be implemented as a single software or software module, and no specific limitation is made here.

[0039] The server 103 can be a business server that provides various services. It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or can be implemented as a single software or software module, and no specific limitation is made here.

[0040] It should be understood that Figure 1 the numbers of terminals, networks, and servers in [[ ]] are only illustrative. According to the implementation requirements, there can be any number of terminals, networks, and servers.

[0041] Please refer to Figure 2 , Figure 2 which is a system interaction diagram of an image processing method provided by an embodiment of the present application. It can be understood that in the embodiment of the present application, the execution entity can be a terminal or a processor in the terminal, or it can also be a related service that executes the image processing method in the terminal. For the convenience of description, the following takes the execution entity as the processor in the terminal as an example, and combines Figure 1 and Figure 2 to introduce the system interaction process in an image processing method.

[0042] S201. The processor identifies the target object in the target image, and determines the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to the reference point in the target object.

[0043] Optionally, determining the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to the reference point in the target object includes: obtaining the two-dimensional coordinate probability distribution map of each object key point in the target object in the image two-dimensional coordinate system, and obtaining the one-dimensional coordinate probability distribution map of each object key point in the target object in the one-dimensional coordinate system perpendicular to the image two-dimensional coordinate system; determining the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system according to the two-dimensional coordinate probability distribution map, and determining the relative depth of each object key point in the target object relative to the reference point in the target object according to the one-dimensional coordinate probability distribution map.

[0044] S202. The processor determines the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth.

[0045] Optionally, determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth includes: determining the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system according to the two-dimensional position information, the relative depth, and the preset camera internal parameters; determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the preset camera external parameters.

[0046] S203. The processor determines the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model corresponding to the target object according to the two-dimensional position information and the first three-dimensional position information.

[0047] Optionally, determining the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model corresponding to the target object according to the two-dimensional position information and the first three-dimensional position information includes: determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the first three-dimensional position information; driving the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle so that the posture of the target model is the same as the posture of the target object in the target image; obtaining the model three-dimensional position information of each model key point in the world three-dimensional coordinate system in the moved target model, and determining the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model according to the two-dimensional position information and the model three-dimensional position information.

[0048] Optionally, determining the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model according to the two-dimensional position information and the model three-dimensional position information includes: determining the true camera extrinsic parameters corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information; determining the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true camera extrinsic parameters.

[0049] Optionally, determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle includes: determining the normal vector corresponding to each object key point in the target object; determining the initial relationship between the parent node and the child node of each object key point in the target object according to the position information of the parent node, the position information of the child node, and the normal vector of each object key point in the target object; determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the position information of the parent node, the position information of the child node, and the initial relationship between the parent node and the child node of each object key point in the target object.

[0050] Optionally, driving the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle includes: driving the movement of each model key point in the target model corresponding to the target object through a forward driving method based on the driving rotation angle.

[0051] In an embodiment of the present application, first, a target object in a target image is recognized, and two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object are determined; according to the two-dimensional position information and the relative depth, first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system is determined; according to the two-dimensional position information and the first three-dimensional position information, second three-dimensional position information of each model key point in a target model corresponding to the target object in the world three-dimensional coordinate system is determined. Since the two-dimensional position information of each object key point in the target object and the relative depth can be determined based on a single target image, the first three-dimensional position information of each object key point in the target object can be determined first, and then, according to the two-dimensional position information and the first three-dimensional position information of each object key point in the target object, the second three-dimensional position information of each model key point in the target model corresponding to the target object can be determined, so that the three-dimensional position information of the target model corresponding to the target object can be determined through a single image, reducing the complexity of determining the three-dimensional position and increasing the accuracy of determining the three-dimensional position.

[0052] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an image processing method provided in another embodiment of the present application.

[0053] As Figure 3 shown, the method includes:

[0054] S301. Recognize a target object in the target image, and determine two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object.

[0055] It can be understood that an image processing method provided in an embodiment of the present application is mainly used to recognize or detect an object in an image and determine the three-dimensional position information of a model corresponding to the object, where the image can be an image of any type or source. Therefore, the image to be recognized or detected can be obtained first, and the image is determined as the target image.

[0056] After obtaining the target image, the object in the target image can be recognized to determine the target object in the target image. For example, an object box detection method can be used to determine the target object in the target image. When there are more than one target object in the target image, the categories of different target objects can be the same or different; for example, multiple target objects are all people; or multiple target objects are all vehicles. Another example is that the target objects in the target image include: people and animals; or the target objects in the target image include people and vehicles, and the target object category is specifically determined according to the actual application scenario requirements. For the sake of description, the target object in the target image is taken as an example of a person for introduction below.

[0057] After determining the target object in the target image, the object key points in the target object can also be determined. The object key points refer to the core points used to represent or constitute the target object. For example, when the target object in the target image is a person, the object key points of the target object can be determined as the joint points of the human body. Determining the object key points in the target object can be a key point detection method, etc. For example, when the target object in the target image is a person, the key point detection method can be a method implemented based on the Cascaded Pyramid Network (CPN) network, or a method implemented based on the Simple Baselines network, etc.

[0058] After determining the object key points of the target object, a two-dimensional coordinate system can be established based on the target image, and this two-dimensional coordinate system can be determined as the image two-dimensional coordinate system. The specific process of establishing the image two-dimensional coordinate system can be as follows: First, determine the area corresponding to the target image, where this area includes the target object. Then, select a point in this area as the origin. Finally, determine the two-dimensional coordinate system based on the origin and the plane where this area is located.

[0059] After determining the object key points of the target object and the image two-dimensional coordinate system, since both the object key points of the target object and the image two-dimensional coordinate system are determined based on the target image, that is, the object key points of the target object and the image two-dimensional coordinate system can be in the same dimension. Therefore, the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system can be determined.

[0060] Furthermore, a reference point can also be determined based on each object key point in the target object. This reference point coincides with the origin of the above image two-dimensional coordinate system. For example, when the object key points of the target object are determined as the joint points in the human body, then the reference point can be the pelvic joint point among the joint points. The purpose of determining the reference point in the target object is to determine the relative distance of each object key point in the target object relative to this reference point in the direction perpendicular to the image two-dimensional coordinate system. This relative distance can also be considered as the relative depth of each object key point in the target image relative to the reference point in the target object.

[0061] S302. Determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth.

[0062] Optionally, based on the above establishment of the image two-dimensional coordinate system, a world three-dimensional coordinate system of the target object in the real world can also be established. Then the position of the target object in the world three-dimensional coordinate system is also the position of the target object in the real world.

[0063] After determining the two-dimensional position information of each object key point in the target object in the two-dimensional image coordinate system and the relative depth relative to the reference point in the target object, since the two-dimensional position information of each object key point in the target object represents the position information of the target object on a certain two-dimensional plane, and the relative depth of each object key point in the target object represents the position information of the target object on a certain one-dimensional plane, therefore, based on the two-dimensional position information and the relative depth of each object key point in the target object, the three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system can be converted, and this three-dimensional position information is determined as the first three-dimensional position information.

[0064] S303. Determine the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the first three-dimensional position information.

[0065] After determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system, in order to obtain the three-dimensional position information of the target model corresponding to the target object in the real world, the two-dimensional position information of each object key point in the target object and the first three-dimensional position information can be combined to describe the position of the target model corresponding to the target object from the perspective of the camera, and determine the distance from each model key point in the target model to the camera. This distance can be regarded as the absolute depth of each model key point in the target model to the camera. Furthermore, according to this absolute depth, the absolute coordinates of each model key point in the target model in the world three-dimensional coordinate system, that is, the second three-dimensional position information, are determined, and this second three-dimensional position information is used as the final three-dimensional position information of the target model in the real world.

[0066] In the embodiment of the present application, first, the target object in the target image is recognized, and the two-dimensional position information of each object key point in the target object in the two-dimensional image coordinate system and the relative depth relative to the reference point in the target object are determined; according to the two-dimensional position information and the relative depth, the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system is determined; according to the two-dimensional position information and the first three-dimensional position information, the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system is determined. Since the two-dimensional position information and the relative depth of each object key point in the target object can be determined according to a target image, the first three-dimensional position information of each object key point in the target object can be determined first, and then, according to the two-dimensional position information and the first three-dimensional position information of each object key point in the target object, the second three-dimensional position information of each model key point in the target model corresponding to the target object is determined, so that the three-dimensional position information of the target model corresponding to the target object can be determined through one image, reducing the complexity of determining the three-dimensional position and increasing the accuracy of determining the three-dimensional position.

[0067] Please refer toFigure 4 , Figure 4 It is a schematic flowchart of an image processing method provided by another embodiment of this application.

[0068] As Figure 4 shown, the method includes:

[0069] S401. Identify the target object in the target image, obtain the two-dimensional coordinate probability distribution map of each object key point in the target object in the two-dimensional coordinate system of the image, and obtain the one-dimensional coordinate probability distribution map of each object key point in the target object in the one-dimensional coordinate system perpendicular to the two-dimensional coordinate system of the image.

[0070] In the embodiment of this application, a pixel coordinate system, a two-dimensional coordinate system of the image, a one-dimensional coordinate system, a three-dimensional coordinate system of the camera, and a three-dimensional coordinate system of the world can be respectively established based on the target image. Please refer to Figure 5 , Figure 5 It is a schematic diagram of the coordinate system reference provided by another embodiment of this application. As Figure 5 shown, where the coordinate system uv is the pixel coordinate system. The pixel coordinate system is a two-dimensional coordinate system obtained by taking a vertex of the rectangular area corresponding to the target image as the origin and taking the two borders near the origin in the rectangular border as the u coordinate axis and the v coordinate axis respectively; the coordinate system xy is the two-dimensional coordinate system of the image. The two-dimensional coordinate system of the image is a coordinate system formed based on the plane where the target image is located, and its origin O is the position where the reference point in the target object is located; the coordinate system Zc is the one-dimensional coordinate system. The one-dimensional coordinate system is a coordinate system perpendicular to the two-dimensional coordinate system of the image, that is, the coordinate axis of the one-dimensional coordinate system is perpendicular to the two-dimensional coordinate system xy, and the coordinate axis of the one-dimensional coordinate system passes through the origin O of the two-dimensional coordinate system xy; the coordinate system XcYcZc is the three-dimensional coordinate system of the camera. The three-dimensional coordinate system of the camera is established from the camera angle corresponding to the target image. The plane formed by the Xc coordinate axis and the Yc coordinate axis of the three-dimensional coordinate system of the camera is parallel to the plane where the two-dimensional coordinate system xy of the image is located, and the Zc coordinate axis of the three-dimensional coordinate system of the camera is parallel to the coordinate axis of the one-dimensional coordinate system. The origin of the three-dimensional coordinate system of the camera coincides with the origin of the one-dimensional coordinate system, both are Oc, and the Zc coordinate axis in the three-dimensional coordinate system of the camera is parallel to the coordinate axis of the one-dimensional coordinate system; the coordinate system ZwXwYw is the three-dimensional coordinate system of the world. The three-dimensional coordinate system of the world can be considered as the coordinate system corresponding to the target object in the real world. The position of the coordinate axis of the three-dimensional coordinate system of the world can be not limited. In binocular vision, generally, the origin of the three-dimensional coordinate system of the world is set at the midpoint of the left camera, the right camera, or the X-axis direction of both.

[0071] Further, in Figure 5The middle pixel coordinate system can be used to represent the position information of pixels in the target image; the image two-dimensional coordinate system can be used to represent the position information of image points in the target image; the one-dimensional coordinate system can be used to represent the relative depth of an image point in the target image relative to a certain image point; the camera three-dimensional coordinate system can be used to represent the three-dimensional position information of an object in the target image or the model corresponding to the object in the target image from the perspective of the camera; the world three-dimensional coordinate system is used to represent the three-dimensional position information of an object in the target image or the model corresponding to the object in the target image in the real world. Among them, P(Xw, Yw, Zw) is the coordinate point in the world three-dimensional coordinate system, p(x, y) is the coordinate point in the image two-dimensional coordinate system, and the pixel coordinates corresponding to p(x, y) in the pixel coordinate system are (u, v), and f is the camera focal length, which is equal to the distance from Oc to O.

[0072] After determining the coordinate system, it is possible to obtain the two-dimensional coordinate probability distribution diagram of each object key point in the target object in the image two-dimensional coordinate system, and obtain the one-dimensional coordinate probability distribution diagram of each object key point in the target object in the one-dimensional coordinate system perpendicular to the image two-dimensional coordinate system.

[0073] Specifically, please refer to Figure 6 , Figure 6 which is a schematic diagram of obtaining the coordinate probability distribution provided by another embodiment of this application. As Figure 6 shown, Figure 6 it can be regarded as a network structure for pose recognition. After determining the target object in the target image 610, features in the target image 610 can be extracted based on the EfficientNet 620, and then the features are sampled based on at least one Upsample 630. Among them, the Upsample 630 can be implemented based on the Conv2D 631 and PixelShuffle 632. Finally, based on the above image two-dimensional coordinate system and one-dimensional coordinate system, the sampled data can be divided into two processing parts. One part of the sampled data obtains the two-dimensional coordinate probability distribution diagram 660 of each object key point in the target object in the image two-dimensional coordinate system based on the Conv2D 631. The two-dimensional coordinate probability distribution diagram 660 represents the distribution of the two-dimensional coordinate points (coordinate points composed of the x coordinate and y coordinate) of each object key point in the target object in the image two-dimensional coordinate system; the other part of the sampled data obtains the one-dimensional coordinate probability distribution diagram 670 of each object key point in the target object in the one-dimensional coordinate system based on the Avgpool 640 and fully connected (FC) 650. The one-dimensional coordinate probability distribution diagram 670 represents the distribution of the one-dimensional coordinate points (coordinate points composed of the Zc coordinate) of each object key point in the target object in the one-dimensional coordinate system.

[0074] Optionally, the network structure for pose recognition includes Figure 6 the structure shown in the figure but is not limited to the above method, and any network that can learn two-dimensional pose points and relative depth through the network can be included.

[0075] S402. Determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system according to the two-dimensional coordinate probability distribution diagram, and determine the relative depth of each object key point in the target object relative to the reference point in the target object according to the one-dimensional coordinate probability distribution diagram.

[0076] After determining the two-dimensional coordinate probability distribution diagram of each object key point in the target object in the image two-dimensional coordinate system and the one-dimensional coordinate probability distribution diagram in the one-dimensional coordinate system perpendicular to the image two-dimensional coordinate system, the two-dimensional coordinate probability distribution diagram can be converted into the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system. The specific conversion method may not be limited. For example, the data in the two-dimensional coordinate probability distribution diagram is averaged to obtain the x coordinate and y coordinate of each object key point in the target object in the image two-dimensional coordinate system respectively; the coordinates in the dataset of the two-dimensional coordinate probability distribution diagram can also be obtained as the x coordinate and y coordinate of each object key point in the target object in the image two-dimensional coordinate system.

[0077] Similarly, the one-dimensional coordinate probability distribution diagram of each object key point in the target object in the one-dimensional coordinate system can also be converted into the one-dimensional position information of each object key point in the target object in the one-dimensional coordinate system. Since there is only one coordinate system in the one-dimensional coordinate system and the coordinate axis of the one-dimensional coordinate system passes through the position of the reference point in the target object, the coordinate value of each object key point in the one-dimensional coordinate system is the relative distance from each object key point to the reference point, and this distance can also be regarded as the relative depth of each object key point in the target object relative to the reference point in the target object.

[0078] S403. Determine the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system according to the two-dimensional position information, relative depth, and preset camera internal parameters.

[0079] In the process of determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and relative depth of each object key point in the target object, it can be divided into two steps. The first step is to convert the two-dimensional position information and relative depth of each object key point in the target object into camera three-dimensional position information. Specifically, it can be calculated through the following formula:

[0080]

[0081] Among them, Zc is the coordinate value corresponding to the relative depth of the object key point, x and y are the coordinate values corresponding to the two-dimensional position information of the object key point respectively, Xc, Yc, and Zc are the coordinate values corresponding to the three-dimensional position information of the object key point in the camera three-dimensional coordinate system respectively, fx, fy, u, and v are preset camera internal parameters, where the preset camera internal parameters are standard data, and the preset camera internal parameters can be obtained through Zhang Zhengyou calibration.

[0082] S404. Determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the preset camera external parameters.

[0083] In the process of determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and relative depth of each object key point in the target object, the second step is to convert the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system into the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system. Specifically, it can be calculated by the following formula:

[0084]

[0085] Among them, Xc, Yc, and Zc are the coordinate values corresponding to the three-dimensional position information of the object key point in the camera three-dimensional coordinate system respectively, Xw, Yw, and Zw are the coordinate values corresponding to the first three-dimensional position information of the object key point in the world three-dimensional coordinate system respectively, R and T are the preset camera external parameters, which are the rotation matrix and the translation matrix respectively. The R matrix is obtained by multiplying three rotation matrices. If the three rotation matrices are R1, R2, and R3, then R1, R2, and R3 are the rotation matrices around the Xc coordinate axis, Yc coordinate axis, and Zc coordinate axis respectively. The T matrix is obtained according to the set default depth, and the depth is also the distance from the reference point in the target object to the origin in the camera three-dimensional coordinate system. For example, when the set default depth is 2, the unit corresponding to the default depth can be set according to needs, and can be units such as meters, centimeters, or millimeters. When the unit corresponding to the default depth is meters, then the T matrix can be obtained as

[0086] S405. Determine the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the first three-dimensional position information.

[0087] After determining the first three-dimensional position information of each object key point of the target object in the world three-dimensional coordinate system, the three-dimensional position information of the target model corresponding to the target object in the world three-dimensional coordinate system can also be determined. Specifically, according to the first three-dimensional position information of each object key point of the target object obtained in the above steps, the driving rotation angle of each object key point of the target object relative to the initial rotation angle can be determined.

[0088] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the target object in the standard pose provided by another embodiment of the present application. The initial rotation angle can be considered as the target object being in the standard pose. In Figure 7 , when the target object 710 is a person, the standard pose is a T-shaped pose with feet together and arms outstretched. The angles of each object key point of the target object 710 in the standard pose are the initial angles. Therefore, the driving rotation angle is related to the current pose of the target object in the target image.

[0089] The method for determining the driving rotation angle of each object key point of the target object relative to the initial rotation angle may include: first, determining the normal vectors corresponding to each object key point of the target object. Specifically, at least three object key points can be determined. For example, when the target object is a person, the hip point position, the left hip position, and the right hip position can be used as the object key points for determining the normal vector, and then the corresponding normal vector is calculated through the normal vector calculation formula. The normal vector calculation formula is:

[0090] Normal vector = TriangleNormal(hip point position, left hip position, right hip position), where TriangleNormal is the standard triangle function. Based on the standard triangle function of the hip point position, the left hip position, and the right hip position, the standard triangle function can be determined, and the normal vector corresponding to the hip point position and the left hip position can be determined.

[0091] Furthermore, then, according to the position information of the parent node, the position information of the child node, and the normal vector among each object key point of the target object, the initial relationship between the parent node and the child node among each object key point of the target object is determined, and the initial relationship is also the node initial inverse matrix.

[0092] Node initial inverse matrix = Quaternion.Inverse(Quaternion.LookRotation(parent node position point - child node position point, normal vector)), where the Quaternion.LookRotation function is the look rotation function, and Quaternion.Inverse is the inverse function.

[0093] Finally, according to the position information of the parent nodes, the position information of the child nodes, and the initial relationship between the parent nodes and the child nodes among the object key points in the target object, the driving rotation angles corresponding to the object key points in the target object are determined. For a node, if this node has child nodes, then this node is the parent node of the child nodes, but this node may also be the child node of other nodes. Therefore, the parent node and the child node are a relative relationship. Then, the driving rotation angles corresponding to the object key points are also the driving rotation angles of the parent node bones.

[0094] The driving rotation angle of the parent node bone = Quaternion.LookRotation(parent node position – child node position, normal vector) * initial inverse matrix of the node * initial rotation matrix of the node, where the Quaternion.LookRotation function is the look-at rotation function.

[0095] S406. Drive the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle, so that the pose of the target model is the same as the pose of the target object in the target image.

[0096] Optionally, the target model corresponding to the target object can also be preset in advance. The shape of the target model is the same as or similar to that of the target object, but the target model is in a standard pose (not in any motion state). Therefore, after determining the driving rotation angles corresponding to the object key points in the target object, the movement of each model key point in the target model corresponding to the target object can be driven. Among them, the model key points are similar to the object key points, and the model key points are the core points that represent or constitute the target model, until the pose of the target model is the same as the pose of the target object in the target image.

[0097] The method of driving the movement of each object key point in the target model corresponding to the target object may not be limited. For example, the Forward Kinematics (FK) method can be used to drive the movement of each model key point in the target model corresponding to the target object based on the driving rotation angle. Specifically, when using the forward driving method to drive the movement of each model key point in the target model corresponding to the target object based on the driving rotation angle, first calculate the orientation of the target object according to the positions of the spine point and the left and right hip joint points, then calculate the driving rotation angle of each child node point based on the parent node through the bone relationship, and finally drive the movement of the target model by setting the angle to rotate the model components.

[0098] Further, on the basis of driving the movement of each model key point in the target model corresponding to the target object using the Forward Kinematics (FK) method based on the driving rotation angle, the inverse kinematics method can be continued to drive the movement of each model key point in the target model corresponding to the target object, so that when driving the movement of each model key point in the target model corresponding to the overall driving target object, it is more accurate. When driving the movement of each model key point in the target model corresponding to the target object using the inverse kinematics method based on the driving rotation angle, first, based on the driving rotation angle of each child node, and then based on the calculated orientation of the target object and the bone relationship according to the positions of the spine point, left and right hip joint points, the driving rotation angle of the child node is transformed, and the driving rotation angle of the parent node based on the child node is inversely solved level by level. Finally, by setting the angle rotation model components, the target model is driven to move.

[0099] Please refer to Figure 8 and Figure 9 , Figure 8 which is a schematic diagram of target model driving provided by another embodiment of the present application, Figure 9 which is a schematic diagram of target model driving provided by another embodiment of the present application.

[0100] As Figure 8 shown, in Figure 8 there is a target object 810 in the target image. Based on the target object 810, the target model 820 corresponding to the target object can be determined in Figure 9 , and based on the position information of each object key point in the target object, the target model 820 is driven to run, so that the posture 820 of the target model is the same as the posture of the target object 810 in the target image.

[0101] S407. Obtain the model three-dimensional position information of each model key point in the target model after movement, and determine the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the two-dimensional position information and the model three-dimensional position information.

[0102] Since the posture of the target model is the same as the posture of the target object in the target image, the position information of each model key point in the target model can be considered as the position information of each object key point in the target object in the real world. Then, the model three-dimensional position information of each model key point in the target model after movement can be obtained, and according to the two-dimensional position information of the target object in the two-dimensional coordinate system and the model three-dimensional position information, the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system is determined.

[0103] The method for determining the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model may be as follows: First, determine the true camera extrinsic parameters corresponding to the target model based on the two-dimensional position information of the target object in the two-dimensional coordinate system and the model three-dimensional position information. Since in the above embodiment, during the process of calculating the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system, the preset camera extrinsic parameters used are those set by oneself or the default camera extrinsic parameters, in order to obtain more realistic three-dimensional position information of the target model in the world three-dimensional coordinate system, the true camera extrinsic parameters of the target model can also be calculated. Optionally, the true camera extrinsic parameters of the target object can be determined according to the two-dimensional position information of each object key point in the target object and the model three-dimensional position information. Specifically, the Perspective-n-Point (PnP) algorithm can be solved based on the method of minimizing the reprojection error, and the two-dimensional position information of each object key point in the target object is matched with the model three-dimensional position information of the target model to obtain the true camera extrinsic parameters of the target model; the true camera extrinsic parameters of the target model can also be solved through the EPnP (Efficient Perspective-n-Point Camera Pose Estimation) algorithm. Optionally, the true camera extrinsic parameters obtained after solving can be further corrected according to the model three-dimensional position information of the target model to make the true camera extrinsic parameters more accurate.

[0104] Furthermore, the true camera extrinsic parameters of the target model mainly include a rotation matrix and a translation matrix. In the embodiments of the present application, the translation matrix in the true camera extrinsic parameters is mainly obtained. The translation matrix represents the distance from the reference point in the object to the origin in the camera three-dimensional coordinate system. Then, the translation matrix in the true camera extrinsic parameters represents the true distance from the reference point in the target model to the origin in the camera three-dimensional coordinate system, that is, the absolute depth from the reference point in the target model to the origin in the camera three-dimensional coordinate system.

[0105] Furthermore, determine the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true camera extrinsic parameters. After calculating the true camera extrinsic parameters of the target model, similar to the above step of "determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the preset camera extrinsic parameters", directly according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true camera extrinsic parameters, the second three-dimensional position information of each model key point in the world three-dimensional coordinate system in the target model can be determined, which will not be elaborated here.

[0106] In an embodiment of the present application, an image processing method is proposed, which solves the problems that the current three-dimensional human pose detection method requires multiple images, the detection method is relatively complex, and the depth of each model key point in the target model is directly predicted inaccurately. By using the knowledge of graphics to solve the problems of image science, the true position of the target model corresponding to the target object in space can be obtained. Based on the true position of the target model in space, the target model corresponding to the target object can be accurately moved forward and backward in space, or the target model can be fitted into the video, or applied to virtual fitting, etc.

[0107] Please refer to Figure 10 and Figure 11 , Figure 10 which is a schematic diagram of target model driving provided by another embodiment of the present application. Figure 11 which is a schematic diagram of target model driving provided by another embodiment of the present application.

[0108] As Figure 10 shown, when the image processing method in the embodiment of the present application is applied to virtual fitting, after determining the true three-dimensional position information of each model key point in the target model 1010 corresponding to the target object in the world three-dimensional coordinate system, the target model 1010 can be accurately moved forward and backward in space, and the target model 1010 can be fitted into the icon image. As Figure 11 shown, the virtual clothes 1020 can also be controlled to fit into the target image where the target object is located based on the target model 1010, realizing virtual fitting from the user's perspective.

[0109] In the embodiment of the present application, since the two-dimensional position information and relative depth of each object key point in the target object can be determined according to a target image, the first three-dimensional position information of each object key point in the target object can be determined first, and then the second three-dimensional position information of each model key point in the corresponding target model in the target object can be determined according to the two-dimensional position information and the first three-dimensional position information of each object key point in the target object. The three-dimensional position information of the target model can be determined through one image, reducing the complexity of determining the three-dimensional position and increasing the accuracy of determining the three-dimensional position.

[0110] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of an image processing device provided by another embodiment of the present application.

[0111] As Figure 12 shown, the image processing device 1200 includes:

[0112] The object two-dimensional position determination module 1210 is configured to identify a target object in a target image, and determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth with respect to a reference point in the target object.

[0113] The object three-dimensional position determination module 1220 is configured to determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth.

[0114] The model three-dimensional position determination module 1230 is configured to determine the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the first three-dimensional position information.

[0115] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of an image processing apparatus provided in another embodiment of the present application.

[0116] As Figure 13 shown, the image processing apparatus 1300 includes:

[0117] The probability distribution map acquisition module 1310 is configured to acquire the two-dimensional coordinate probability distribution map of each object key point in the target object in the image two-dimensional coordinate system, and acquire the one-dimensional coordinate probability distribution map of each object key point in the target object in the one-dimensional coordinate system perpendicular to the image two-dimensional coordinate system.

[0118] The two-dimensional and depth calculation module 1320 is configured to determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system according to the two-dimensional coordinate probability distribution map, and determine the relative depth of each object key point in the target object with respect to a reference point in the target object according to the one-dimensional coordinate probability distribution map.

[0119] The camera position calculation module 1330 is configured to determine the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system according to the two-dimensional position information, the relative depth, and a preset camera internal parameter.

[0120] The first three-dimensional position determination module 1340 is configured to determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and a preset camera external parameter.

[0121] The driving rotation angle determination module 1350 is configured to determine the driving rotation angle of each object key point in the target object with respect to the initial rotation angle according to the first three-dimensional position information.

[0122] Among them, determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle includes: determining the normal vector corresponding to each object key point in the target object; determining the initial relationship between the parent node and the child node among the object key points in the target object according to the position information of the parent node, the position information of the child node, and the normal vector among the object key points in the target object; determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the position information of the parent node, the position information of the child node, and the initial relationship between the parent node and the child node among the object key points in the target object.

[0123] The driving module 1360 is configured to drive the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle, so that the pose of the target model is the same as the pose of the target object in the target image.

[0124] Driving the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle includes: driving the movement of each model key point in the target model corresponding to the target object through a forward driving method based on the driving rotation angle.

[0125] The second three-dimensional position determination module 1370 is configured to obtain the model three-dimensional position information of each model key point in the target model after movement in the world three-dimensional coordinate system, and determine the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the two-dimensional position information and the model three-dimensional position information.

[0126] Among them, determining the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the two-dimensional position information and the model three-dimensional position information includes: determining the true camera external parameters corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information; determining the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true camera external parameters.

[0127] Please refer to Figure 14 , Figure 14 which is a schematic data flow diagram of an image processing device provided in another embodiment of the present application.

[0128] As Figure 14 shown, taking the image processing device 1200 as an example, in the object two-dimensional position determination module, the target object is obtained through object recognition from the target image, and then the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to the reference point in the target object are obtained through pose estimation. The two-dimensional position information of each object key point in the target object and the relative depth relative to the reference point in the target object can constitute the 2.5D pose of the target object.

[0129] In the object three-dimensional position determination module, by setting the camera parameters, based on the back-projection method, the first three-dimensional position information of each object key point of the target object in the world three-dimensional coordinate system can be determined, that is, the 3D pose of the target object.

[0130] In the model three-dimensional position determination module, first, based on forward driving, the model three-dimensional position information of the target model corresponding to the target object is determined, that is, the 3D pose of the target model. Then, based on the two-dimensional position information of the target object in the two-dimensional coordinate system, the model three-dimensional position information, and the solution algorithm, the true camera parameters corresponding to the target model are determined. Finally, based on the camera parameters and the back-projection method, the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system is determined, that is, the final 3D pose of the target model.

[0131] The embodiment of the present application also provides a computer storage medium. The computer storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the steps of the method in any one of the above embodiments.

[0132] Further, please refer to Figure 15 , Figure 15 which is a schematic structural diagram of a terminal provided by the embodiment of the present application. As Figure 15 shown, the terminal 1500 may include: at least one central processing unit 1501, at least one network interface 1504, a user interface 1503, a memory 1505, and at least one communication bus 1502.

[0133] Among them, the communication bus 1502 is used to realize the connection and communication between these components.

[0134] Among them, the user interface 1503 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1503 may further include a standard wired interface and a wireless interface.

[0135] Among them, the network interface 1504 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0136] Among them, the central processing unit 1501 may include one or more processing cores. The central processing unit 1501 connects various parts within the entire terminal 1500 through various interfaces and circuits, and executes various functions of the terminal 1500 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1505, and by calling the data stored in the memory 1505. Optionally, the central processing unit 1501 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The central processing unit 1501 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the central processing unit 1501 and may be implemented separately by a single chip.

[0137] Among them, the memory 1505 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1505 includes a non-transitory computer-readable storage medium. The memory 1505 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1505 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1505 may also be at least one storage device located far from the aforementioned central processing unit 1501. As Figure 15 shown, the memory 1505, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an image processing program.

[0138] In Figure 15In the terminal 1500 shown, the user interface 1503 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the central processing unit 1501 can be used to call the image processing program stored in the memory 1505 and specifically perform the following operations:

[0139] Identify the target objects in the target image, determine the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image and the relative depth relative to the reference point in the target object; according to the two-dimensional position information and the relative depth, determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system; according to the two-dimensional position information and the first three-dimensional position information, determine the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system.

[0140] Optionally, determining the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image and the relative depth relative to the reference point in the target object includes: obtaining the two-dimensional coordinate probability distribution map of each object key point in the target object in the two-dimensional coordinate system of the image, and obtaining the one-dimensional coordinate probability distribution map of each object key point in the target object in the one-dimensional coordinate system perpendicular to the two-dimensional coordinate system of the image; determining the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image according to the two-dimensional coordinate probability distribution map, and determining the relative depth of each object key point in the target object relative to the reference point in the target object according to the one-dimensional coordinate probability distribution map.

[0141] Optionally, according to the two-dimensional position information and the relative depth, determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system includes: determining the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system according to the two-dimensional position information, the relative depth, and the preset camera internal parameters; determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the preset camera external parameters.

[0142] Optionally, according to the two-dimensional position information and the first three-dimensional position information, determining the second three-dimensional position information of each model key point in the target model corresponding to the target object in the world three-dimensional coordinate system includes: determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the first three-dimensional position information; driving the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle so that the posture of the target model is the same as the posture of the target object in the target image; obtaining the model three-dimensional position information of each model key point in the moved target model in the world three-dimensional coordinate system, and determining the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the two-dimensional position information and the model three-dimensional position information.

[0143] Optionally, determining the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the two-dimensional position information and the model three-dimensional position information includes: determining the true external camera parameters corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information; determining the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true external camera parameters.

[0144] Optionally, determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle includes: determining the normal vector corresponding to each object key point in the target object; determining the initial relationship between the parent node and the child node of each object key point in the target object according to the position information of the parent node, the position information of the child node and the normal vector of each object key point in the target object; determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the position information of the parent node, the position information of the child node and the initial relationship between the parent node and the child node of each object key point in the target object.

[0145] Optionally, driving the movement of each model key point in the target model corresponding to the target object according to the driving rotation angle includes: driving the movement of each model key point in the target model corresponding to the target object through a forward driving method based on the driving rotation angle.

[0146] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.

[0147] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, in each embodiment of the present application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module.

[0149] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0150] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present application.

[0151] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0152] The above is the description of an image processing method, device, storage medium, and terminal provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An image processing method, characterized in that, The method includes: Identifying a target object in a target image, and determining two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image and a relative depth relative to a reference point in the target object; wherein, the relative depth is used to represent the relative distance of the object key point in the direction perpendicular to the two-dimensional coordinate system of the image relative to the reference point; Determining first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth; Determining a driving rotation angle of each object key point in the target object relative to an initial rotation angle according to the first three-dimensional position information; wherein, the initial rotation angle is used to represent that the target object is in a standard pose; in the standard pose, the angles of the object key points of the target object are initial angles; Driving each model key point in the target model corresponding to the target object to move according to the driving rotation angle, so that the pose of the target model is the same as the pose of the target object in the target image; Obtaining model three-dimensional position information of each model key point in the target model after movement in the world three-dimensional coordinate system; Determining a true camera extrinsic parameter corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information; Determining second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and the true camera extrinsic parameter; wherein, the position information of each model key point in the target model is used to represent the position information of each object key point in the target object in the real world; Driving the target model to move in space based on the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system; The determining the true camera extrinsic parameter corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information includes: Solving the n-point perspective algorithm based on the method of minimizing the reprojection error, and matching the two-dimensional position information with the model three-dimensional position information to obtain the true camera extrinsic parameter corresponding to the target model.

2. The method according to claim 1, wherein The determining the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image and the relative depth relative to the reference point in the target object includes: Obtaining a two-dimensional coordinate probability distribution map of each object key point in the target object in the two-dimensional coordinate system of the image, and obtaining a one-dimensional coordinate probability distribution map of each object key point in the target object in the one-dimensional coordinate system perpendicular to the two-dimensional coordinate system of the image; Determining the two-dimensional position information of each object key point in the target object in the two-dimensional coordinate system of the image according to the two-dimensional coordinate probability distribution map, and determining the relative depth of each object key point in the target object relative to the reference point in the target object according to the one-dimensional coordinate probability distribution map.

3. The method according to claim 1 or 2, characterized in that, Determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth includes: Determining the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system according to the two-dimensional position information, the relative depth, and a preset camera internal parameter; Determining the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the camera three-dimensional coordinate system and a preset camera external parameter.

4. The method according to claim 1, wherein Determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle includes: Determining the normal vector corresponding to each object key point in the target object; Determining the initial relationship between the parent node and the child node among each object key point in the target object according to the position information of the parent node, the position information of the child node, and the normal vector among each object key point in the target object; Determining the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the position information of the parent node, the position information of the child node, and the initial relationship between the parent node and the child node among each object key point in the target object.

5. The method according to claim 1, wherein Driving each model key point in the target model corresponding to the target object to move according to the driving rotation angle includes: Driving each model key point in the target model corresponding to the target object to move through a forward driving method based on the driving rotation angle.

6. An image processing apparatus, characterized in that, The device includes: An object two-dimensional position determination module, configured to identify a target object in a target image, and determine the two-dimensional position information of each object key point in the target object in the image two-dimensional coordinate system and the relative depth relative to a reference point in the target object; wherein, the relative depth is used to represent the relative distance of the object key point relative to the reference point in the direction perpendicular to the image two-dimensional coordinate system; An object three-dimensional position determination module, configured to determine the first three-dimensional position information of each object key point in the target object in the world three-dimensional coordinate system according to the two-dimensional position information and the relative depth; A driving rotation angle determination module, configured to determine the driving rotation angle of each object key point in the target object relative to the initial rotation angle according to the first three-dimensional position information; A driving module, configured to drive each model key point in the target model corresponding to the target object to move according to the driving rotation angle, so that the posture of the target model is the same as the posture of the target object in the target image; wherein, the shape of the target model is the same as the shape of the target object; the target model before driving is not in a moving state; A second three-dimensional position determination module, configured to obtain the model three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system after movement; The second three-dimensional position determination module is specifically configured to determine the true camera external parameter corresponding to the target model according to the two-dimensional position information and the model three-dimensional position information. Determine the second three-dimensional position information of each model key point in the world three-dimensional coordinate system according to the three-dimensional position information of each object key point in the target object in the three-dimensional coordinate system of the camera and the true camera extrinsic parameters; wherein, the position information of each model key point in the target model is used to represent the position information of each object key point in the target object in the real world. The second three-dimensional position determination module is specifically configured to: solve the n-point perspective algorithm based on the method of minimizing the reprojection error, match the two-dimensional position information with the model three-dimensional position information, and obtain the true camera extrinsic parameters corresponding to the target model. The image processing device is further configured to drive the target model to move in space based on the second three-dimensional position information of each model key point in the target model in the world three-dimensional coordinate system.

7. A computer storage medium, characterized in that, The computer storage medium stores a plurality of operations, and the operations are adapted to be loaded and executed by a processor to perform the steps of the method according to any one of claims 1 to 5.

8. A terminal, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, electronic device and storage medium

    CN109448090A

  • Image processing method and device, electronic equipment and storage medium

    CN111582207A

  • Method and device for determining three-dimensional position of target object and roadside equipment

    CN112184914A