Computer visual representation learning method and device based on transformation isomorphic constraint

By constructing a representation learning method of transformed isomorphic constraints in hidden space, the problem that the existing technology cannot be applied to regression tasks is solved, the accuracy and training efficiency of three-dimensional manual pose estimation are improved, and better generalization and robustness are achieved.

CN120374722APending Publication Date: 2025-07-25INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432363.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing representation learning methods cannot be effectively applied to regression tasks, especially object detection tasks, such as three-dimensional hand posture estimation. The existing methods fail to effectively combine the characteristics of the regression tasks, resulting in poor performance and performance.

Method used

Using a computer vision representation learning method based on transform isomorphic constraints, we train the backbone network in hidden space by constructing transform isomorphic constraints, and combine reconstruction and discriminant constraints to improve the performance and robustness of the regression task.

Benefits of technology

The task accuracy and training convergence efficiency of the regression task are improved, and the network's performance in the regression task is enhanced, especially in the three-dimensional hand position estimation task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374722A_ABST
    Figure CN120374722A_ABST
Patent Text Reader

Abstract

The invention provides a computer vision representation learning method and device based on transformation isomorphic constraint, and the method comprises the steps: transforming each training image, obtaining a transformed image processed by an image transformation matrix, obtaining a pose transformation matrix of a target in the transformed image according to the image transformation matrix, and obtaining a pose representation of the target in the transformed image; the image transformation matrix and the pose transformation matrix form a transformation pair; gathering all transformation pairs of all training images to obtain an isomorphic transformation set; inputting the transformed image into a reconstruction type backbone network to obtain a reconstruction result and a transformation set in a hidden space of the backbone network, constructing a first-order constraint, and constructing a reconstruction constraint according to the reconstruction result and the transformed image corresponding to the reconstruction result; and extracting visual features of a training image by adopting the trained backbone network, decoding the visual features into a predicted pose through a decoder, constructing a loss function to train the decoder according to the predicted pose and a target pose, and combining the trained decoder and the backbone network to obtain a pose extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of representation learning in the field of computer vision, and also relates to the technical field of three-dimensional human hand pose estimation in the field of object detection, and particularly relates to a representation learning method, device, electronic device, computer-readable storage medium and computer program product based on transformation isomorphism constraint. Background Art

[0002] The paradigm of representation learning methods in the field of computer vision is usually: (1) pre-training on a large scale of unlabeled or easily obtainable labeled image data; (2) extracting the pre-trained backbone network and combining it with a task head network with a relatively simple structure and relatively few parameters for fine-tuning under a small amount of target task data, so that the learning of the target task is easier and the effect is better. The backbone networks widely used at present are basically obtained through representation learning, and the representation learning methods are mainly divided into two types: discriminative representation learning methods (extracting semantic information of classification tasks through supervised learning) and reconstructive representation learning methods (learning detailed features through image reconstruction tasks). The main representatives of the former are CNN (Convolutional Neural Network), ResNet (Residual Network), GAN (Generative Adversarial Network), etc., and the main representatives of the latter are ViT (Vision Transformer), BERT (Bidirectional Encoder Representations from Transformers), GPT (General Purpose Transformer), etc.

[0003] Discriminative-based representation learning is to send a large amount of simple task data into the network for supervised training. The reconstructive-based representation learning scheme is to randomly damage the image, and the damaged image is reconstructed by the backbone network to complete its representation learning. Both of these methods have an unsolved problem: they do not consider the scenario of regression tasks. A regression task is a task where the target value is a continuous variable, as opposed to a classification task where the target value is a discrete variable.

[0004] Regression tasks, such as prediction, are for continuous target values. For example, in an object detection task, the position of an object in an image is represented by image coordinates (x, y), where x and y are continuous variables (i.e., the value range is the intervals [0, H] and [0, W]). In contrast, classification tasks have discrete prediction targets, such as image classification and OCR. That is, the difference between regression tasks and classification tasks lies in whether the target value is continuous or discrete.

[0005] The data used in discriminative representation learning schemes is basically prepared for classification tasks. The backbone network trained by such a scheme can efficiently extract semantic information such as categories and entities from images, but this information has little relevance to regression tasks. Reconstruction-based representation learning focuses on learning information that can reconstruct image details, which is too redundant for a regression task and instead interferes with the network's learning and reasoning. In summary, existing representation learning methods cannot well adapt to the learning of regression tasks. To solve this problem, the present invention proposes a representation learning training framework based on transform isomorphism. Summary of the Invention

[0006] The object of the present invention is to solve the problem that the above-mentioned prior art cannot be applied to the representation learning of regression tasks, and proposes a representation learning framework suitable for regression tasks. Compared with traditional representation learning frameworks, the present invention pays more attention to the performance and performance of the backbone network in regression tasks. On top of the framework, the present invention also provides a practice of applying this framework to the hand pose estimation task. However, the present invention is not limited to this. The present invention can be used for all object detection tasks belonging to regression tasks, such as pose-related estimation tasks, including human pose estimation, head orientation estimation, object pose estimation, etc.

[0007] Aiming at the deficiencies of the prior art, as Figure 3 shown, the present invention proposes a computer vision representation learning method based on transform isomorphism constraints, which includes:

[0008] Initial step, obtaining multiple first training images, performing rotation, flipping, and flip-rotation transformations on the first training images to obtain transformed images processed by an image transformation matrix, modeling a transformation model in the latent space of a reconstruction-based backbone network according to the image transformation matrix, and the transformation of the image and its corresponding transformation model in the latent space form an isomorphic transformation pair. Collecting the transformation pairs of all training images to obtain an isomorphic transformation set;

[0009] Training step: Randomly damage the transformed image and input it into the reconstruction backbone network, which generates image features; the reconstruction network restores the image based on the image features to obtain a reconstructed image, and forms a reconstruction constraint by comparing it with the transformed image; according to the set of isomorphic transformations, use the transformation model to construct a first-order constraint; use the first-order constraint and the reconstruction constraint to perform representation learning training on the reconstruction backbone network, save the trained reconstruction backbone network, and discard the reconstruction network and the transformation model to obtain a feature extraction model;

[0010] Fine-tuning step: Obtain multiple second training images with target pose annotations. After the feature extraction model extracts the features of the second training images, input them into the pose decoder to obtain predicted poses, and construct a supervision constraint based on the target pose and the predicted poses to fine-tune the pose decoder;

[0011] Prediction step: Based on the pose decoder and the feature extraction model after fine-tuning, construct a pose extraction model, and input the target image into the pose extraction model to obtain the pose information of the target object in the target image.

[0012] The computer vision representation learning method based on transformation isomorphism constraints, wherein the training step further includes: using a second-order constraint to perform representation learning training on the reconstruction backbone network; the construction process of the second-order constraint includes:

[0013] According to a preset operation rule table of transformation isomorphism relations, the operation rule table includes multiple transformation schemes. Combine the transformation pairs corresponding to the transformation schemes in the set of isomorphic transformations in pairs, and perform rotation, flipping, and flip-rotation transformations on the transformed image again to obtain a re-transformed image and its transformation pair. Construct the second-order constraint based on the re-transformed image and its transformation pair, and the transformation model.

[0014] The computer vision representation learning method based on transformation isomorphism constraints, wherein the initial step includes:

[0015] Use the geometric transformation matrix of the image To describe the rotation and flipping transformations performed on the training image, and the pose transformation matrix of the target object in the corresponding training image is Where R is an orthogonal matrix used to represent the rotation orientation of the target object, and t is a three-dimensional vector used to represent the spatial position of the target object;

[0016] The pose of the target object in the transformed training image can then be described by the matrix diag(A,1)×[R,t].

[0017] Where:

[0018]

[0019] diag is a diagonal matrix function, which means arranging the incoming square matrix or scalar parameter as the diagonal elements of the matrix to form a new matrix; if A = diag(A, 1), then the transformation A in the image space and the transformation in the pose space There is a corresponding relationship: for the image I containing the root joint pose θ, the transformation isomorphism relationship has the following expression:

[0020]

[0021] Among them, is the pose decoder, is the target pose estimation network. Through the above formula process, a pair of transformation pairs collects N transformation pairs to form a pairwise isomorphic transformation set

[0022] The described computer vision representation learning method based on transformation isomorphism constraints, where the operation rule table is used to record the binary combination relationship of transformations; the operation rule table is used for the transformation of training images and the poses of objects in the training images.

[0023] The described computer vision representation learning method based on transformation isomorphism constraints, where the training step includes:

[0024] The backbone network with learning parameter Θ is Model the transformations in the latent space of the backbone network respectively h ν ,r ξ ,k λ are learnable parameters, and construct the first-order constraint:

[0025]

[0026] H,R α ,K β respectively represent the horizontal flip of the image I, the rotation α around the center, and the rotation β around the center after the horizontal flip.

[0027] The described computer vision representation learning method based on transformation isomorphism constraints, where the training step includes:

[0028] Based on the operation rule table, combine the transformation pairs corresponding to the transformation schemes in the isomorphic transformation set pairwise to construct multiple second-order constraints:

[0029]

[0030] As Figure 4 shown, the present invention also proposes a computer vision representation learning device based on transformation isomorphism constraints, which includes:

[0031] The initial module obtains multiple first training images, performs rotation, flipping, and flip-rotation transformation on the first training images to obtain transformed images processed by an image transformation matrix. According to the image transformation matrix, a transformation model in the latent space of the reconstruction backbone network is modeled. The transformation of the image and its corresponding transformation model in the latent space form an isomorphic transformation pair. The transformation pairs of all training images are collected to obtain an isomorphic transformation set;

[0032] The training module randomly damages the transformed images and inputs them into the reconstruction backbone network, and the backbone network generates image features; the reconstruction network restores the image based on the image features to obtain a reconstructed image, and forms a reconstruction constraint by comparing it with the transformed image; according to the isomorphic transformation set, a first-order constraint is constructed using the transformation model; using the first-order constraint and the reconstruction constraint, the reconstruction backbone network is trained for representation learning, and the trained reconstruction backbone network is saved and the reconstruction network and the transformation model are discarded to obtain a feature extraction model;

[0033] The fine-tuning module obtains multiple second training images with target pose annotations. After the feature extraction model extracts the features of the second training images, it inputs them into a pose decoder to obtain a predicted pose. A supervision constraint is constructed based on the target pose and the predicted pose to fine-tune the pose decoder;

[0034] The prediction module constructs a pose extraction model based on the pose decoder and the feature extraction model after fine-tuning, and inputs the target image into the pose extraction model to obtain the pose information of the target object in the target image.

[0035] The present invention proposes an electronic device, which includes the computer vision representation learning device based on transformation isomorphism constraint described above. The electronic device is connected to an information display device, and the information display device is used to display the pose information of the target object with display parameters, attributes set by the user, or through an artificial intelligence model.

[0036] The present invention proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the computer vision representation learning method based on transformation isomorphism constraint are implemented.

[0037] The present invention proposes a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the computer vision representation learning methods based on transformation isomorphism constraint are implemented.

[0038] As can be seen from the above solutions, the advantages of the present invention are:

[0039] The present invention proposes a representation learning framework for regression learning. Specifically, transformation isomorphism constraints are discovered and imposed on the latent space. The present invention can improve the performance of representation learning methods in specific regression tasks, such as task accuracy and training convergence efficiency. At the same time, it has better robustness and generalization compared to traditional single-supervised learning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the concept of transformation isomorphism;

[0041] Figure 2 It is a schematic diagram of representation learning;

[0042] Figure 3 It is a flowchart of the method of the present invention;

[0043] Figure 4 It is a block diagram of the device of the present invention;

[0044] Figure 5 It is a schematic diagram of the structure of the first electronic device of the present invention;

[0045] Figure 6 It is a schematic diagram of the application environment structure of the first electronic device of the present invention;

[0046] Figure 7 It is a schematic diagram of the structure of the second electronic device of the present invention.

[0047] REFERENCE MARKS:

[0048] A - The first electronic device;

[0049] B - Computer vision representation learning device based on transformation isomorphism constraints;

[0050] C - Data acquisition device;

[0051] D - Information display device;

[0052] 1000 - The second electronic device;

[0053] Ⅰ - Computing unit;

[0054] Ⅱ - ROM;

[0055] Ⅲ - RAM;

[0056] Ⅳ - Bus;

[0057] Ⅴ - Interface;

[0058] Ⅵ - Input unit;

[0059] Ⅶ - Output unit;

[0060] Ⅷ - Storage medium;

[0061] IX - Communication Unit. Detailed Implementation Manner

[0062] It should be noted that, in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0063] Without more limitations, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0064] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0065] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0066] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single - CPU or a multi - CPU. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smart phones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0067] The memory is used to store the software program for implementing the solution of the present invention and is controlled by a processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated herein.

[0068] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine certain components, or have different component arrangements.

[0069] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other arbitrary combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0070] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0071] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0072] It should also be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0073] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0074] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0075] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0076] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.

[0077] When conducting research on representation learning in the field of computer vision, the inventor found that the defect in the prior art was caused by the fact that the current representation learning framework did not consider the unique nature of the regression task. Through research on the underlying features of the regression task and the similarity with the target space and input space, the inventor found that this defect could be solved by constructing isomorphic constraints on the transformations of the input space, feature space, and output space.

[0078] Isomorphic transformation means that there exists or can be constructed a family of transformations with the same semantics and algorithms in the input space, feature space, and output space of a specific task. By forcing the network to learn and extract a family of transformations that conform to such algorithms, the present invention can train a backbone network that is more effective for the regression task, because the family of transformations can more precisely depict the structure of the latent space required for the regression task. Specifically, in order to achieve the above technical effects, the present invention proposes the following key technical points:

[0079] Key point 1, a representation learning training system based on isomorphic transformation; realizing the entire process from the discovery and selection of isomorphic transformation, constructing isomorphic transformation constraints, and conducting training, thus realizing a training method for the backbone network for representation learning;

[0080] Key point 2, a hand pose estimation technology for training a single RGB image based on the representation learning system of the isomorphic transformation in Key point 1; realizing the selection of specific image transformations, the modeling of latent space transformations, and the fine-tuning of the head network for the hand pose estimation task.

[0081] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. The specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0082] Before elaborating on the content of the present invention in detail, it is necessary to first clarify the concept of isomorphic transformation. The hand pose estimation task is to estimate the three-dimensional pose of the hand from a single RGB color image containing the hand (described by the three-dimensional spatial positions of the hand joint points or described by the rotation orientations of the hand joint points). Figure 1Taking the hand pose estimation task as an example, when the transformation of the hand in the input space changes, the hand pose in the pose space also changes accordingly. This idea can also be applied to the latent space. It explains the isomorphic transformation relationship that should exist among the input space (i.e., the image), the target space (i.e., the pose), and the latent space in the hand pose estimation task. Obviously, in the input space and the target space, horizontal flipping, rotation, and the combined transformation of flipping and rotation of the image correspond to horizontal flipping, rotation, and the combined transformation of flipping and rotation of the resulting pose. Any nested transformation in the image space is also exactly the same as the nested corresponding transformation in the pose space. For example, after performing "horizontal flipping + rotation by 30° + combined transformation of flipping and rotation by 15°" on the image, the pose in the image is the same as the result of performing "horizontal flipping + rotation by 30° + combined transformation of flipping and rotation by 15°" in the pose space on the pose in the original image without transformation.

[0083] The object of the present invention is achieved by the following technical solutions:

[0084] Construction stage:

[0085] Step 1, find the isomorphic transformation relationship between the input space and the output space of the target task. The isomorphic transformation relationship needs to be obtained by studying the relationship between the target space and the input space. Taking object pose detection estimation as an example, such a strict relationship can be observed: the transformation of the image that can be described by a partial matrix can be equivalently applied to the coordinates of the target in the camera space. When the target is hand pose, the transformation of the image that can be described by a partial matrix can be equivalently applied to the coordinates of the hand in the camera space.

[0086] Taking the hand pose estimation task as an example, a certain rotation and flipping transformation of the image can be described by the matrix For the rotation and flipping transformation of the image, that is, each pixel of the image is rotated and flipped to a specified position. The key lies in calculating the position of the target pixel reached after the rotation and flipping transformation of the current pixel. Since the pixel position can be represented by a two-dimensional vector (x, y), then using the matrix the position of the target pixel A·(x, y) can be expressed, where A is the geometric transformation matrix of the image. Suppose the pose of the root joint of the hand to be estimated in the original image is described as the matrix T , where is the rotation transformation matrix and displacement vector describing the position of the root joint of the hand. Among them, R is an orthonormal matrix used to represent the rotation orientation of the root joint, and t is a three-dimensional vector used to represent the spatial position of the root joint. Then the pose of the root joint of the hand in the image after the transformation can be described by the matrix diag(A, 1)×[R, t]. Among them: is the rotation transformation matrix and displacement vector describing the position of the root joint of the hand. Among them, R is an orthonormal matrix used to represent the rotation orientation of the root joint, and t is a three-dimensional vector used to represent the spatial position of the root joint. Then the pose of the root joint of the hand in the image after the transformation can be described by the matrix diag(A, 1)×[R, t]. Among them:

[0087]

[0088] The diag is a diagonal matrix function, which means arranging the incoming square matrix or scalar parameter as the diagonal elements of a matrix to form a new matrix. Denote A = diag(A, 1), then there is actually a corresponding relationship between the transformation A in the image space and the transformation A in the hand root joint pose space: Assume an ideal hand root joint pose estimation network Then for an image I containing a root joint pose of θ, the transformation isomorphism relationship has the following expression:

[0089]

[0090] Through the above process, a pair of isomorphic transformations in the form of Formula 1 can be found This process can be repeated multiple times to find N groups of different pairs of isomorphic transformations Note that it can be carefully selected so that the transformations included satisfy the following properties with as few transformations as possible: (1) Completeness: For any two randomly selected transformations (both from the input space or the target space) for the composite operation, the result is still a transformation in. (2) Invertibility: For any transformation pair in in there must exist a transformation pair such that the composite transformations Y°X and are both identity transformations. For the transformation space, the existence of the identity transformation is natural. Therefore, if satisfies the above properties (1) and (2), then it can be considered that the selected transformation sets on the input space and the target space also form a group under the transformation isomorphism relationship. And the construction of the group is beneficial to the training of the representation learning model

[0091] Step 2, construct an operation rule table for the transformation isomorphism relationship. Based on the found transformation set , an operation rule table for the binary associative relationship of the transformation can be constructed. Still taking hand pose estimation as an example, select a transformation set that satisfies the above properties where H, R α , K βThey represent "horizontal flipping", "rotation by α around the center", and "rotation by β around the center after horizontal flipping" of the image respectively, where α and β are both angles. The binary relation combination operation rule table 1 can be constructed. The operation rule table should be constructed not only for the transformation of the input space, but also for the transformation of the target space. However, due to the existence of transformation isomorphism, such a construction is easy in one case. Only need to replace H, R, K with That's all.

[0092] x\y H <![CDATA[R α > <![CDATA[K β > H <![CDATA[R0]]> <![CDATA[K α > <![CDATA[R β > <![CDATA[R γ > <![CDATA[K -γ > <![CDATA[R α+γ > <![CDATA[K -γ+β > <![CDATA[K ω > <![CDATA[R -ω > <![CDATA[K α+ω > <![CDATA[R -ω+β >

[0093] Table 1

[0094] Table 1 is the binary relation operation rule table, and the combination order is y°x, that is, first perform transformation x and then perform transformation y. H, R γ , K ω represent "horizontal flipping", "rotation by γ around the center", and "rotation by ω around the center after horizontal flipping" respectively.

[0095] Feature learning stage:

[0096] Step 3, construct transformation isomorphism constraints based on the operation rules. After completing Steps 1 and 2, the transformation set is a group composed of minimized isomorphic transformations. And such isomorphic transformations exist between the input space and the target space. Although this relationship is artificially discovered, it exists objectively and naturally. The present invention believes that there should also be a set of transformations in the latent space of feature learning that have a transformation isomorphism relationship with the transformations found in the input space through the above steps. It is easy to prove that the transformation isomorphism relationship has transitivity. Therefore, when the transformation isomorphism relationship with the transformations in the input space holds, it can be asserted that there are transformation isomorphism relationships between any two of the input space, the latent space, and the target space.

[0097] The transformation set on the latent space cannot be found through Step 1 because the latent space has the problem of low interpretability, and geometric intuition and algebraic intuition cannot directly act on the latent space. Here, the method of feature learning is introduced to make it possible to construct the latent space and learn isomorphic transformations simultaneously.

[0098] Continuing with the above example, denote the backbone network with learnable parameters Θ as and use the learnable model to model the transformations h ν , r ξ , k λ in the latent space as learnable parameters. Based on Formula 1, a first-order constraint can be constructed:

[0099]

[0100] Note that since the two transformations of central rotation transformation and horizontal flipping followed by central rotation require a rotation angle description, so r ξ , k λ There is an input of the rotation angle.

[0101] Similarly, based on Table 1, 9 second-order constraints can be constructed, corresponding to each item in the table. For example, for the item in the third row and the third column of the table, the following constraint can be constructed (Equation 3):

[0102]

[0103] h ν , r ξ , k λ Model the "flipping", "rotation" and "rotation after flipping" transformations on the latent space, corresponding to the transformations on the latent space, that is, the script style Equation Three The meaning of is to first apply "rotate α" and then apply "rotate β" in the image space. According to Table 1, the result of these two transformations is "rotate α + β", that is, R α+β , which is isomorphic to r on the latent space ξ (·, α + β).

[0104] In this way, 12 constraints can be obtained from the transformation isomorphism. These 12 constraints belong to the discriminative representation learning constraints. In addition to the reconstruction constraints for representation learning, these 13 constraints are the constraint forms adopted by the present invention in the practice of hand pose estimation tasks. Reconstruction-based representation learning is to recover the original image from the features extracted from the image, while discriminative representation learning is based on the method of pulling positive samples closer and pushing negative samples away. There are fundamental differences between the two. The training method of the present invention combines "reconstruction constraints" and multiple "discriminative constraints", improving the training efficiency and model accuracy. It should be noted that a relatively simple model should be selected to model the transformation on the latent space. If the complexity of the latent space transformation is too high, it will lead to overfitting to the structure of the latent space itself and unable to generate strong enough transformation isomorphism constraints.

[0105] Step 1, find the transformation isomorphism relationship between the image space and the hand pose space in the hand pose estimation task. Follow the three transformations H, R α , K β , that is, "horizontal flipping", "central rotation α" and "central rotation β after horizontal flipping". These three transformations meet the properties (1) and (2) proposed in Step 1.

[0106] Step 2, construct an operation rule table for the transformation isomorphism relationship. The result is shown in Table 1.

[0107] Step 3: Construct transformation isomorphism constraints based on the operation rules. In the specific implementation of this step, it can be divided into two sub-steps, namely, determining the form of the loss function of the constraints and determining the specific form of the transformation on the latent space.

[0108] Step 3-1: Determine the form of the loss function of the constraints. The construction of the isomorphism constraints is shown in Formulas 2 and 3. By introducing an auxiliary reconstruction network The reconstruction constraint is in the form of

[0109]

[0110] According to the constraints described in Formulas 2, 3, and 4, the following loss functions can be constructed: first-order isomorphism loss, second-order isomorphism loss, and reconstruction loss.

[0111] The construction of the first-order isomorphism loss is the sum of the absolute differences between the two terms on both sides of the equal sign for each term in Formula 2:

[0112]

[0113] The construction of the second-order isomorphism loss is the loss obtained by describing the constraints for each term in Table 1 in the form of Formula 3 and then constructing it in the same way:

[0114]

[0115] The construction of the reconstruction loss follows the mainstream:

[0116]

[0117] The final loss function is the weighted sum of the three:

[0118]

[0119] In this example, w1 = w2 = 10 -3 , w3 = 1.

[0120] Step 3-2: Determine the specific form of the transformation on the latent space. The transformation on the latent space whose specific form is to be determined is They respectively correspond to "horizontal flip", "central rotation", and "central rotation after horizontal flip" in the image space / pose space. In this example, the most common and simple modeling method is adopted for their modeling: using matrices to model the three transformations respectively.

[0121] The "horizontal flip" transformation is modeled as:

[0122] h v = E + AB,

[0123] where n is the dimension of the latent space and E is the n×n identity matrix, is a learnable parameter matrix, and d is a non - negative integer less than n. Since the horizontal flip transformation is a full - rank transformation, the identity matrix E is added to make h v is likely to be full - rank with high probability. Decomposing AB avoids having too many learnable transformation parameters, which can not only control the complexity of the matrix but also reduce the hardware requirements of the training system.

[0124] The modeling methods of the "central rotation" transformation and the "central rotation after horizontal flip" transformation are the same. Here, the "central rotation" transformation is taken as an example to illustrate the modeling method of the latent space transformation r ξ (·; α):

[0125] e = MLP ν ([cosα, sinα] T )

[0126] r ξ (v; α)=Ev + CD·cat(v, e)

[0127] where MLP is a simple multi - layer perceptron with output dimension k, cat represents the concatenation of vectors.

[0128] Step 4, optimize the parameters Θ, v, ξ, λ, ζ according to the loss function shown in Equation 5. The image data is from ImageNet - 1k. The training architecture is as Figure 2 shown. Different transformations are performed on the original image respectively. The latent vectors of the transformed images can be used as supervision signals to adjust the transformation model on the latent space. At the same time, the reconstruction loss avoids information loss of the latent vectors. Among them, the network parameters Θ, ν, ξ, λ, ζ are initialized using Kaiming initialization, the training optimizer uses the AdamW optimizer, the learning rate is set to 10 -4 , and the number of training epochs is set to 20.

[0129] Step 5, fine - tune the hand pose estimation. For the trained backbone network add a decoder to it, such as a pose estimation task head network (consisting of a five - layer multi - layer perceptron with an output dimension of 42×3, representing the rotation orientations of 42 hand joints), and fine - tune it on the DexYCB dataset. Among them, the task head network is initialized using Kaiming initialization, the training optimizer uses the AdamW optimizer, the learning rate is set to 10 -4 , and the number of training epochs is set to 15.

[0130] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0131] As Figure 4 shown, the present invention also proposes a computer vision representation learning device based on transformation isomorphism constraint, which includes:

[0132] An initial module that acquires multiple first training images, performs rotation, flipping, and flip-rotation transformation on the first training images to obtain transformed images processed by an image transformation matrix, models a transformation model in the latent space of a reconstruction backbone network according to the image transformation matrix, and the transformation of the image and its corresponding transformation model in the latent space form an isomorphism transformation pair. All transformation pairs of the training images are collected to obtain an isomorphism transformation set;

[0133] A training module that randomly damages the transformed images and inputs them into the reconstruction backbone network, and the backbone network generates image features; a reconstruction network restores the image based on the image features to obtain a reconstructed image, and forms a reconstruction constraint by comparing it with the transformed image; according to the isomorphism transformation set, a first-order constraint is constructed using the transformation model; using the first-order constraint and the reconstruction constraint, the reconstruction backbone network is trained for representation learning, and the trained reconstruction backbone network is saved and the reconstruction network and the transformation model are discarded to obtain a feature extraction model;

[0134] A fine-tuning module that obtains multiple second training images with target pose annotations, extracts the features of the second training images by the feature extraction model and inputs them into a pose decoder to obtain a predicted pose, constructs a supervision constraint according to the target pose and the predicted pose, and fine-tunes the pose decoder;

[0135] A prediction module that constructs a pose extraction model based on the pose decoder and the feature extraction model after fine-tuning, and inputs a target image into the pose extraction model to obtain the pose information of the target object in the target image.

[0136] As Figure 5 shown, in another embodiment of the present invention, a first electronic device A is also proposed, which includes the above-mentioned computer vision representation learning device based on transformation isomorphism constraint.

[0137] As Figure 6As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission solution. The data acquisition device C is used to acquire a target image, such as the image with hand pose described in the embodiments of the present invention. The information display device D is used to display the pose information of the target object analyzed by the present invention.

[0138] Among them, the information display device D can process and organize the data output by the first electronic device A based on an information display mechanism to improve the readability of the data output by the first electronic device A. This information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed. It can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user, enabling the user to understand this information more timely without having to access a secondary page or scroll the page, saving the user's operations. Or this information display mechanism can be an artificial intelligence AI display model. It can learn the key information of the user according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.

[0139] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the computer vision representation learning method based on transformation isomorphism constraints provided by the above-mentioned various methods.

[0140] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the computer vision representation learning method based on transform isomorphism constraints. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0141] Figure 7 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement the embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0142] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other through a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0143] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0144] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S3. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).

[0145] Although the embodiments of the present invention have been disclosed as above, it is not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrated and described examples here.

Claims

1. A computer vision representation learning method based on transformation isomorphism constraints, characterized in that, Including: Initial step: Obtain multiple first training images, perform rotation, flipping, and flip-rotation transformation on the first training images to obtain transformed images processed by an image transformation matrix. According to the image transformation matrix, model a transformation model in the latent space of a reconstruction backbone network. The transformation of the image and its corresponding transformation model in the latent space form an isomorphic transformation pair. Collect the transformation pairs of all training images to obtain an isomorphic transformation set; Training step: Randomly damage the transformed images and input them into the reconstruction backbone network, and the backbone network generates image features; The reconstruction network restores the image based on the image features to obtain a reconstructed image, and forms a reconstruction constraint by comparing it with the transformed image; According to the isomorphic transformation set, use the transformation model to construct a first-order constraint; Use the first-order constraint and the reconstruction constraint to perform representation learning training on the reconstruction backbone network, save the reconstruction backbone network after training is completed, and discard the reconstruction network and the transformation model to obtain a feature extraction model; Fine-tuning step: Obtain multiple second training images with target pose annotations. After the feature extraction model extracts the features of the second training images, input them into a pose decoder to obtain a predicted pose. Construct a supervision constraint based on the target pose and the predicted pose to fine-tune the pose decoder; Prediction step: Based on the pose decoder and the feature extraction model after fine-tuning, construct a pose extraction model, input a target image into the pose extraction model, and obtain the pose information of the target object in the target image.

2. The method for computer vision representation learning based on transformation isomorphism constraint according to claim 1, wherein The training step further includes: Using a second-order constraint to perform representation learning training on the reconstruction backbone network; The construction process of the second-order constraint includes: According to a pre-set operation rule table for transformation isomorphism relations, the operation rule table includes multiple transformation schemes. Combine the transformation pairs corresponding to the transformation schemes in the isomorphic transformation set in pairs, perform rotation, flipping, and flip-rotation transformation on the transformed images again to obtain re-transformed images and their transformation pairs. Construct the second-order constraint based on the re-transformed images and their transformation pairs, and the transformation model.

3. The computer vision representation learning method based on transformational isomorphism constraints according to claim 1, characterized in that, The initial step includes: Using the geometric transformation matrix of the image Describe the transformation performed on the first training image, and the pose transformation matrix of the target object in the corresponding training image is where R is an identity orthogonal matrix used to express the rotational orientation of the target object, and t is a three-dimensional vector used to express the spatial position of the target object; The pose of the target object in the transformed training image can be described by the matrix diag(A,1)×[R,t]. Where: diag is the diagonal matrix function, which means arranging the incoming square matrix or scalar parameter as the diagonal elements of a matrix to form a new matrix; if A = diag(A, 1), then there is a correspondence between the transformation A in the image space and the transformation in the pose space: for an image I containing a root joint pose of θ, the transformation isomorphism relation has the following expression: There is a correspondence: for an image I containing a root joint pose of θ, the transformation isomorphism relation has the following expression: Among them, is a pose decoder, is an object pose estimation network, and through the above formula process, a pair of transformation pairs collect N transformation pairs into a pairwise isomorphic transformation set 4. The computer vision representation learning method based on transformation isomorphism constraints according to claim 2, characterized in that The operation rule table is used to record the binary combination relationship of the transformation; The operation rule table is used for the transformation of the training images and the poses of the target objects in the training images.

5. The computer vision representation learning method based on transformation isomorphism constraint according to claim 3, characterized in that The training step includes: The backbone network with learning parameter Θ is Model the transformation model in the latent space of the backbone network respectively h ν ,r ξ ,k λ Are learnable parameters, construct the first-order constraint: H, R α , K β respectively represent the horizontal flipping, rotation by α around the center, and rotation by β around the center after horizontal flipping of the image I.

6. The computer vision representation learning method based on transformation isomorphism constraint according to any one of claims 1 to 5, characterized in that, The training step includes: Based on the operation rule table, combine the transformation pairs corresponding to the transformation schemes in the isomorphic transformation set in pairs to construct multiple second-order constraints:

7. A computer vision representation learning device based on transformation isomorphism constraints, characterized in that, Including: Initial module: Obtain multiple first training images, perform rotation, flipping, and flip-rotation transformation on the first training images to obtain transformed images processed by an image transformation matrix. According to the image transformation matrix, model a transformation model in the latent space of a reconstruction backbone network. The transformation of the image and its corresponding transformation model in the latent space form an isomorphic transformation pair. Collect the transformation pairs of all training images to obtain an isomorphic transformation set; A training module that randomly damages the transformed image and inputs it into the reconstruction backbone network, which generates image features; a reconstruction network restores the image based on the image features to obtain a reconstructed image, and compares it with the transformed image to form a reconstruction constraint; according to the set of isomorphic transformations, a first-order constraint is constructed using the transformation model; using the first-order constraint and the reconstruction constraint, the reconstruction backbone network is trained for representation learning, and the trained reconstruction backbone network is saved while discarding the reconstruction network and the transformation model to obtain a feature extraction model; A fine-tuning module that obtains multiple second training images with target pose annotations. After the feature extraction model extracts the features of the second training images, it inputs them into a pose decoder to obtain predicted poses. A supervision constraint is constructed based on the target pose and the predicted poses to fine-tune the pose decoder; A prediction module that constructs a pose extraction model based on the pose decoder and the feature extraction model after fine-tuning, and inputs a target image into the pose extraction model to obtain the pose information of the target object in the target image.

8. An electronic device, characterized in that, It includes the computer vision representation learning device based on transformation isomorphism constraint described in claim 7. The electronic device is connected to an information display device, and the information display device is used to display the pose information of the target object with display parameters, attributes set by the user, or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the steps of the computer vision representation learning method based on transformation isomorphism constraint described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the computer vision representation learning method based on transformation isomorphism constraint described in any one of claims 1-6.