A method, apparatus, device, and storage medium for determining pose.
By using a compensation model in a 3D tracking algorithm to determine the pose deviation of an object in a 2D image, the problem of pose estimation error caused by unclear contour texture is solved, the accuracy of pose estimation is improved, and it can be applied to autonomous driving and AR navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2026-04-03
AI Technical Summary
Existing 3D tracking algorithms suffer from significant errors in estimating the pose of objects in 2D images due to the indistinct contours and textures of some objects.
By acquiring the target object in the two-dimensional image, the initial pose and the position information of the specified region image are determined using a preset tracking algorithm. These are then input into a pre-trained compensation model to determine the pose deviation and perform compensation to obtain the actual pose.
It improves the accuracy of 3D tracking algorithms in estimating the pose of objects in 2D images, and enhances the route planning accuracy of autonomous driving equipment and the AR navigation accuracy of mobile devices.
Smart Images

Figure CN116309823B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer vision technology, and in particular to a method, apparatus, device and storage medium for determining pose. Background Technology
[0002] With the development of computer technology, estimating the pose of objects in two-dimensional images based on three-dimensional tracking algorithms has become a fundamental task in computer vision technology.
[0003] However, because the contour texture of some objects is not obvious, the 3D tracking algorithm cannot accurately extract the contour texture features of these objects, which in turn leads to a large error in the pose estimated by the 3D tracking algorithm.
[0004] Therefore, how to improve the accuracy of the pose of objects in two-dimensional images estimated by three-dimensional tracking algorithms is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a method, apparatus, device, and storage medium for determining pose, in order to partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This specification provides a method for determining pose, the method comprising:
[0008] Acquire a two-dimensional image and identify the target objects involved in the two-dimensional image;
[0009] The initial pose of the target object is determined by a preset tracking algorithm, and a specified region of the image containing the target object is determined in the two-dimensional image. The image position information of the specified region of the image in the two-dimensional image is also determined.
[0010] The initial pose, the specified region image, and the image position information are input into a pre-trained compensation model so that the compensation model determines the pose deviation between the initial pose and the actual pose of the target object when the two-dimensional image is acquired, based on the initial pose, the specified region image, and the image position information.
[0011] Based on the pose deviation, the initial pose is compensated to obtain the actual pose of the target object.
[0012] Optionally, based on the initial pose, the specified region image, and the image position information of the specified region image in the two-dimensional image, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined, specifically including:
[0013] Based on the initial pose, determine the pose characteristics of the target object;
[0014] Based on the image of the specified region, determine the contour features of the target object;
[0015] The positional features of the target object are determined based on the image position information of the specified region image in the two-dimensional image;
[0016] Based on the pose features, contour features, and position features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined.
[0017] Optionally, based on the pose features, contour features, and position features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined, specifically including:
[0018] Based on the contour features and position features of the target object, the target pose features of the target object are determined;
[0019] Based on the target pose features and the pose features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined.
[0020] Optionally, based on the pose features, contour features, and position features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined, specifically including:
[0021] Through each feature extraction layer of the compensation model, the target pose features of the target object are determined based on the contour features and position features of the target object;
[0022] Based on the target pose features and the pose features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined.
[0023] Optionally, the target pose features of the target object are determined through each feature extraction layer of the compensation model based on the contour features and position features of the target object, specifically including:
[0024] For each feature extraction layer of the compensation model, the similarity matrix between the contour features and the position features of the target object is determined by the feature extraction layer based on the contour features and the position features of the target object, and the comprehensive features of the target object are determined based on the similarity matrix.
[0025] The comprehensive features output from each feature extraction layer are fused to obtain the target pose features of the target object.
[0026] Optionally, training the compensation model specifically includes:
[0027] Identify the target object in the two-dimensional image of the sample;
[0028] The initial pose of the sample target object, the image of the specified region, and the image position information are input into the compensation model so that the compensation model can determine the pose deviation between the initial pose of the sample target object and the actual pose of the sample target object when acquiring the sample two-dimensional image based on the initial pose of the sample target object, the image of the specified region, and the image position information.
[0029] The compensation model is trained with the optimization objective of minimizing the difference between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample 2D image, and the difference between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image.
[0030] This specification provides a pose determination device, comprising:
[0031] An acquisition module is used to acquire a two-dimensional image and identify the target objects involved in the two-dimensional image;
[0032] The determination module is used to determine the initial pose of the target object through a preset tracking algorithm, and to determine a specified region image containing the target object in the two-dimensional image, and to determine the image position information of the specified region image in the two-dimensional image;
[0033] The compensation module is used to input the initial pose, the specified region image, and the image position information into a pre-trained compensation model, so that the compensation model determines the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the initial pose, the specified region image, and the image position information.
[0034] The execution module is used to compensate the initial pose based on the pose deviation to obtain the actual pose of the target object.
[0035] Optionally, the compensation module is specifically used to: determine the pose features of the target object based on the initial pose; determine the contour features of the target object based on the specified region image; determine the position features of the target object based on the image position information of the specified region image in the two-dimensional image; and determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the pose features, the contour features, and the position features of the target object.
[0036] Optionally, the compensation module is specifically used to: determine the target pose features of the target object based on the contour features and the position features of the target object; and determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
[0037] Optionally, the compensation module is specifically used to determine the target pose features of the target object based on the contour features and the position features of the target object through each feature extraction layer of the compensation model; and to determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
[0038] Optionally, the compensation module is specifically used to, for each feature extraction layer of the compensation model, determine the similarity matrix between the contour features and the position features of the target object based on the contour features and the position features of the target object, and determine the comprehensive features of the target object based on the similarity matrix; and fuse the comprehensive features output by each feature extraction layer to obtain the target pose features of the target object.
[0039] Optionally, the device further includes: a training module;
[0040] The training module is specifically used to: determine the target object in the sample 2D image; input the initial pose, specified region image, and image position information of the target object into the compensation model, so that the compensation model determines the pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image, based on the initial pose of the target object, the specified region image, and the image position information; and train the compensation model with the optimization objective of minimizing the difference between the pose deviation between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample 2D image, and the true pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image.
[0041] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for determining the pose.
[0042] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for determining the pose.
[0043] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0044] The pose determination method provided in this specification first acquires a two-dimensional image and identifies the target object involved in the two-dimensional image. Through a preset tracking algorithm, the initial pose corresponding to the target object is determined, as well as the specified region image containing the target object in the two-dimensional image is determined, and the image position information of the specified region image in the two-dimensional image is determined. The initial pose, the specified region image, and the image position information are input into a pre-trained compensation model, so that the compensation model determines the pose deviation between the initial pose and the actual pose of the target object when the two-dimensional image is acquired, based on the initial pose, the specified region image, and the image position information. Based on the pose deviation, the initial pose is compensated to obtain the actual pose of the target object.
[0045] As can be seen from the above method, the contour features of the target object can be determined by the compensation model based on the image information of the specified area of the target object in the two-dimensional image captured by the camera. Furthermore, the image position information of the specified area of the target object in the entire two-dimensional image can be used to determine the corresponding image of the target object in the two-dimensional image, as well as its two-dimensional position features. Based on the determined contour features, two-dimensional position features, and the initial pose of the target object determined by the preset tracking algorithm, the pose deviation between the initial pose determined by the preset tracking algorithm and the actual pose of the target object when the two-dimensional image was captured can be predicted. Then, the predicted pose deviation can be used to compensate for the initial pose of the target object determined by the preset tracking algorithm, thereby improving the accuracy of the pose of the object in the two-dimensional image estimated by the three-dimensional tracking algorithm. Attached Figure Description
[0046] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0047] Figure 1 This is a flowchart illustrating a method for determining a pose provided in this specification.
[0048] Figure 2 This is a schematic diagram showing the image location information of a specified area image provided in this specification within a two-dimensional image.
[0049] Figure 3 This is a schematic diagram illustrating the process of determining the pose deviation provided in this specification.
[0050] Figure 4 This is a schematic diagram of a pose determination device provided in this specification;
[0051] Figure 5 This specification provides a corresponding Figure 1 A schematic diagram of an electronic device. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0053] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0054] Figure 1 This is a flowchart illustrating a pose determination method provided in this specification, including the following steps:
[0055] S101: Acquire a two-dimensional image and identify the target objects involved in the two-dimensional image.
[0056] In this specification, the specified device can identify the pose of the target object in the two-dimensional image acquired by the sensor set on the specified device, and perform the corresponding task based on the identified pose of the target object.
[0057] The aforementioned pose can refer to the 6D pose of the target object. Here, the 6D pose refers to the translation and rotation transformations of the sensor's camera coordinate system relative to the world coordinate system at the moment of capturing the 2D image corresponding to the target object. The 6D here refers to three degrees of freedom of translation transformation (i.e., translation along the X-axis, translation along the Y-axis, and translation along the Z-axis) and three degrees of freedom of rotation transformation (i.e., rotation around the X-axis, rotation around the Y-axis, and rotation around the Z-axis). After determining the 6D pose of the target object, the pose of the target object in 3D space can be determined based on the 6D pose of the target object, and then the corresponding task can be performed based on the pose of the target object in 3D space.
[0058] The aforementioned tasks could include, for example, planning the driving routes of autonomous vehicles to enable them to avoid collisions with target objects, or providing AR navigation for users in environments such as shopping malls using their mobile devices.
[0059] In this specification, the execution subject used to implement the pose determination method can refer to a specified device set on the business platform, such as a server, or a specified device such as a desktop computer, laptop computer, or mobile phone. For ease of description, the following description will only use the server as the execution subject to illustrate the pose determination method provided in this specification.
[0060] S102: Using a preset tracking algorithm, determine the initial pose of the target object based on the image information of the two-dimensional image, determine a specified region image in the two-dimensional image that contains the target object, and determine the image position information of the specified region image in the two-dimensional image.
[0061] After identifying the target object from the 2D image acquired by the sensor, the server can determine the initial pose of the target object based on the 2D image using a preset tracking algorithm. This tracking algorithm can refer to a 3D object tracking algorithm.
[0062] Furthermore, the server can determine a specified region image containing the target object from a two-dimensional image. This can be understood as cropping the specified region containing the target object from the two-dimensional image to obtain the specified region image.
[0063] Furthermore, since the specified area image only contains two-dimensional image information corresponding to the target object, it is not possible to obtain the image position information of the specified area image in the two-dimensional image based on the specified area image. In other words, after cropping the specified area image from the two-dimensional image, if only the cropped image is viewed, it is impossible to determine which position of the original two-dimensional image the cropped image belongs to. That is to say, the two-dimensional position information of the specified area image is lost during the process of cropping the specified area image from the two-dimensional image. Therefore, in order to compensate for the loss of the two-dimensional position information of the specified area image during the process of cropping the specified area image from the two-dimensional image, the server can also determine the image position information of the specified area image in the two-dimensional image.
[0064] It should be noted that the image location information of the specified region image in the two-dimensional image can include two matrices: a first matrix and a second matrix. The first matrix contains the x-coordinate value of each pixel in the specified region image within the two-dimensional image, and the second matrix contains the y-coordinate value of each pixel in the specified region image within the two-dimensional image. For example... Figure 2 As shown.
[0065] Figure 2 This is a schematic diagram showing the image location information of the specified area image provided in this specification within a two-dimensional image.
[0066] exist Figure 2 In a two-dimensional image, the coordinates of four pixels belonging to a specified region are (1, 1), (1, 2), (2, 1), and (2, 2). The first matrix determined by the server is then... The second matrix is...
[0067] S103: Input the initial pose, the specified region image, and the image position information into a pre-trained compensation model, so that the compensation model determines the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image, based on the initial pose, the specified region image, and the image position information.
[0068] Furthermore, the server can input the determined initial pose of the target object, the image of a specified region containing the target object in the 2D image, and the image position information of the specified region in the 2D image into a pre-trained compensation model. This allows the compensation model to determine the pose deviation between the initial pose of the target object and its actual pose when the 2D image was acquired, based on these parameters. Specifically, as follows... Figure 3 As shown.
[0069] Figure 3 This is a schematic diagram illustrating the process of determining the pose deviation provided in this specification.
[0070] Combination Figure 3 It can be seen that the server can determine the pose features of the target object based on the initial pose of the target object through the compensation model, determine the contour features of the target object based on the specified area image containing the target object in the two-dimensional image, and determine the position features of the target object based on the image position information of the specified area image in the two-dimensional image.
[0071] Furthermore, the server can determine the target pose features of the target object by compensating each feature extraction layer of the model, based on the contour and position features of the target object, and determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and pose features of the target object.
[0072] Specifically, for each feature extraction layer of the compensation model, the similarity matrix between the contour features and position features of the target object is determined by the feature extraction layer. Based on the similarity matrix, the comprehensive features of the target object are determined. The comprehensive features output by each feature extraction layer are fused to obtain the target pose features of the target object. The specific formula can be found in the following formula.
[0073]
[0074] in, The S here o For contour features, P 2d For location features, P 3d As a pose feature, S o P 2d and P 3d The learnable projection matrix, σ(·), represents the softmax normalization function, which is used to normalize the similarity values between the contour features and positional features of the target object. Y is the scaling parameter required for normalization processing, set according to actual needs. (m) This represents the comprehensive features extracted through the m-th feature extraction layer.
[0075] The similarity value between the outline features and positional features of the target object is determined based on the similarity matrix between the outline features and positional features of the target object.
[0076] Furthermore, the server can fuse the comprehensive features output by each feature extraction layer to obtain the target pose features of the target object. This can be achieved by the server initially fusing the comprehensive features output by each feature extraction layer in a specified manner to obtain the initial target pose features. The specified manner can refer to methods such as splicing and fusion. Then, the initial target pose features can be input into the preset feedforward network layer in the compensation model to determine the target pose features of the target object through the feedforward network layer.
[0077] In addition, in practical applications, the compensation model needs to be trained in advance before it can be deployed on the server to determine the pose deviation between the initial pose obtained by the tracking algorithm and the actual pose of the target object when acquiring the two-dimensional image.
[0078] Specifically, when training the compensation model, the server can determine the target object in the sample two-dimensional image, and input the initial pose of the target object, the specified region image, and the image position information into the compensation model so that the compensation model can determine the pose deviation between the initial pose of the target object and the actual pose of the target object when the sample two-dimensional image is acquired, based on the initial pose of the target object, the specified region image, and the image position information.
[0079] The compensation model can then be trained with the optimization objective of minimizing the difference between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample 2D image, and the difference between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image.
[0080] S104: Based on the pose deviation, compensate for the initial pose to obtain the actual pose of the target object.
[0081] Furthermore, after determining the pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the two-dimensional image, the server can compensate for the initial pose determined by the tracking algorithm based on the pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the two-dimensional image, so as to obtain the actual pose of the target object, and perform task execution based on the actual pose of the target object.
[0082] The above describes a pose determination method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding pose determination device, such as... Figure 4 As shown.
[0083] Figure 4 A schematic diagram of a pose determination device provided in this specification includes:
[0084] The acquisition module 401 is used to acquire a two-dimensional image and identify the target objects involved in the two-dimensional image;
[0085] The determining module 402 is used to determine the initial pose corresponding to the target object through a preset tracking algorithm, and to determine a specified area image containing the target object in the two-dimensional image, and to determine the image position information of the specified area image in the two-dimensional image;
[0086] The compensation module 403 is used to input the initial pose, the specified region image, and the image position information into a pre-trained compensation model, so that the compensation model determines the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the initial pose, the specified region image, and the image position information.
[0087] The execution module 404 is used to compensate the initial pose based on the pose deviation to obtain the actual pose of the target object.
[0088] Optionally, the compensation module 403 is specifically configured to: determine the pose features of the target object based on the initial pose; determine the contour features of the target object based on the specified region image; determine the position features of the target object based on the image position information of the specified region image in the two-dimensional image; and determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the pose features, the contour features, and the position features of the target object.
[0089] Optionally, the compensation module 403 is specifically used to: determine the target pose features of the target object based on the contour features and the position features of the target object; and determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
[0090] Optionally, the compensation module 403 is specifically used to determine the target pose features of the target object based on the contour features and the position features of the target object through each feature extraction layer of the compensation model; and to determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
[0091] Optionally, the compensation module 403 is specifically used to, for each feature extraction layer of the compensation model, determine the similarity matrix between the contour features and the position features of the target object based on the contour features and the position features of the target object, and determine the comprehensive features of the target object based on the similarity matrix; and fuse the comprehensive features output by each feature extraction layer to obtain the target pose features of the target object.
[0092] Optionally, the device further includes: a training module 405;
[0093] The training module 405 is specifically used to: determine the target object in the sample two-dimensional image; input the initial pose, specified region image, and image position information of the target object into the compensation model, so that the compensation model determines the pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample two-dimensional image, based on the initial pose of the target object, the specified region image, and the image position information; and train the compensation model with the optimization objective of minimizing the difference between the pose deviation between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample two-dimensional image, and the true pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample two-dimensional image.
[0094] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 One method provided.
[0095] This instruction manual also provides Figure 5 One of the corresponding Figure 1 A schematic diagram of the structure of an electronic device. (e.g.) Figure 5 As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The method described.
[0096] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0097] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0098] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0099] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0100] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0101] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0102] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0103] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0104] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0105] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0106] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0107] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0108] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0109] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0111] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0112] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for determining pose, characterized in that, The method includes: Acquire a two-dimensional image and identify the target objects involved in the two-dimensional image; The initial pose of the target object is determined by a preset tracking algorithm, and a specified region of the image containing the target object is determined in the two-dimensional image. The image position information of the specified region of the image in the two-dimensional image is also determined. The image position information includes a first matrix and a second matrix. The first matrix contains the horizontal coordinate value of each pixel in the specified region of the image in the two-dimensional image, and the second matrix contains the vertical coordinate value of each pixel in the specified region of the image in the two-dimensional image. The initial pose, the specified region image, and the image position information are input into a pre-trained compensation model, so that the compensation model determines the pose features of the target object based on the initial pose, determines the contour features of the target object based on the specified region image, determines the position features of the target object based on the image position information of the specified region image in the two-dimensional image, and determines the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the pose features, contour features, and position features of the target object. Based on the pose deviation, the initial pose is compensated to obtain the actual pose of the target object.
2. The method as described in claim 1, characterized in that, Based on the pose features, contour features, and position features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined, specifically including: Based on the contour features and position features of the target object, the target pose features of the target object are determined; Based on the target pose features and the pose features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined.
3. The method as described in claim 1, characterized in that, Based on the pose features, contour features, and position features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined, specifically including: Through each feature extraction layer of the compensation model, the target pose features of the target object are determined based on the contour features and position features of the target object; Based on the target pose features and the pose features of the target object, the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image is determined.
4. The method as described in claim 3, characterized in that, Through each feature extraction layer of the compensation model, the target pose features of the target object are determined based on the contour features and position features of the target object, specifically including: For each feature extraction layer of the compensation model, the similarity matrix between the contour features and the position features of the target object is determined by the feature extraction layer based on the contour features and the position features of the target object, and the comprehensive features of the target object are determined based on the similarity matrix. The comprehensive features output from each feature extraction layer are fused to obtain the target pose features of the target object.
5. The method as described in claim 1, characterized in that, Training the compensation model specifically includes: Identify the target object in the two-dimensional image of the sample; The initial pose of the sample target object, the image of the specified region, and the image position information are input into the compensation model so that the compensation model can determine the pose deviation between the initial pose of the sample target object and the actual pose of the sample target object when acquiring the sample two-dimensional image based on the initial pose of the sample target object, the image of the specified region, and the image position information. The compensation model is trained with the optimization objective of minimizing the difference between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample 2D image, and the difference between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image.
6. A pose determination device, characterized in that, include: An acquisition module is used to acquire a two-dimensional image and identify the target objects involved in the two-dimensional image; The determination module is used to determine the initial pose of the target object through a preset tracking algorithm, and to determine a specified region image containing the target object in the two-dimensional image, and to determine the image position information of the specified region image in the two-dimensional image; The image location information includes: a first matrix and a second matrix; the first matrix contains the horizontal coordinate value of each pixel in the specified region image in the two-dimensional image, and the second matrix contains the vertical coordinate value of each pixel in the specified region image in the two-dimensional image; The compensation module is used to input the initial pose, the specified region image, and the image position information into a pre-trained compensation model, so that the compensation model determines the pose features of the target object based on the initial pose, determines the contour features of the target object based on the specified region image, determines the position features of the target object based on the image position information of the specified region image in the two-dimensional image, and determines the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the pose features, contour features, and position features of the target object. The execution module is used to compensate the initial pose based on the pose deviation to obtain the actual pose of the target object.
7. The apparatus as claimed in claim 6, characterized in that, The compensation module is specifically used to determine the target pose features of the target object based on the contour features and the position features of the target object; and to determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
8. The apparatus as claimed in claim 6, characterized in that, The compensation module is specifically used to determine the target pose features of the target object based on the contour features and position features of the target object through each feature extraction layer of the compensation model; and to determine the pose deviation between the initial pose and the actual pose of the target object when acquiring the two-dimensional image based on the target pose features and the pose features of the target object.
9. The apparatus as claimed in claim 8, characterized in that, The compensation module is specifically used to, for each feature extraction layer of the compensation model, determine the similarity matrix between the contour features and the position features of the target object based on the contour features and the position features of the target object, and determine the comprehensive features of the target object based on the similarity matrix; and fuse the comprehensive features output by each feature extraction layer to obtain the target pose features of the target object.
10. The apparatus as claimed in claim 6, characterized in that, The device further includes: a training module; The training module is specifically used to: determine the target object in the sample 2D image; input the initial pose, specified region image, and image position information of the target object into the compensation model, so that the compensation model determines the pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image, based on the initial pose of the target object, the specified region image, and the image position information; and train the compensation model with the optimization objective of minimizing the difference between the pose deviation between the initial pose of the target object output by the compensation model and the actual pose of the target object when acquiring the sample 2D image, and the true pose deviation between the initial pose of the target object and the actual pose of the target object when acquiring the sample 2D image.
11. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 5.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Target detection method and target detection device
CN111127551A
Gait recognition method based on skeleton and contour feature fusion
CN115050101A