Key point detection method and device, electronic equipment and storage medium
By using structural feature extractor, codebook and decoder in the image detection model, combined with preset feature vectors, the problem that the prior art cannot predict the location of the occluded key points is solved, and accurate prediction of the location of the key points in the image is achieved.
Patent Information
- Application Number
- CN202311541020.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art cannot accurately predict the location of occluded key points in an image.
By inputting the image to be detected, the model includes a structural feature extractor, a codebook and a decoder. It uses the preset feature vector of the structural features between key points to extract the structural feature vector between key points, and determines the closest target vector from the preset feature vector, and obtains the position coordinates of the key points through the decoder.
Accurate prediction of the locations of unoccluded and obstructed key points in the image is achieved, and the accuracy of key point detection is improved.
Smart Images

Figure CN120020892A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image detection technology, and in particular, to a key point detection method, apparatus, electronic device, and storage medium. Background Art
[0002] In the prior art, the detection of full-body key points in an image is usually implemented based on heat maps or regression. The prior art cannot predict the positions of occluded key points in an image. Summary of the Invention
[0003] Embodiments of the present disclosure provide a key point detection method, apparatus, electronic device, and storage medium, which can more accurately predict the positions of occluded key points.
[0004] In a first aspect, embodiments of the present disclosure provide a key point detection method, including:
[0005] Inputting an image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points;
[0006] Extracting, by the structural feature extractor, a first structural feature vector between the key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include the unoccluded key points and the occluded key points in the image to be detected;
[0007] Determining a first target vector adjacent to the first structural feature vector from the preset feature vectors;
[0008] Decoding, by the decoder, the first target vector into the position coordinates of the key points corresponding to the image to be detected.
[0009] In a second aspect, embodiments of the present disclosure further provide a key point detection apparatus, including:
[0010] An input module, configured to input an image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points;
[0011] A structural feature extraction module, configured to extract, by the structural feature extractor, a first structural feature vector between the key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include the unoccluded key points and the occluded key points in the image to be detected;
[0012] A vector determination module, configured to determine a first target vector adjacent to the first structural feature vector from the preset feature vectors;
[0013] A position prediction module, configured to decode, by means of the decoder, the first target vector into position coordinates of key points corresponding to the image to be detected.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:
[0015] One or more processors;
[0016] A storage device, configured to store one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the key point detection method according to any one of the embodiments of the present disclosure.
[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium including computer-executable instructions, where the computer-executable instructions are used to execute the key point detection method according to any one of the embodiments of the present disclosure when executed by a computer processor.
[0019] In the technical solution of the embodiment of the present disclosure, an image to be detected is input into a detection model; wherein, the detection model includes a structure feature extractor, a codebook, and a decoder, and the codebook includes preset feature vectors of structural features between preset full-scale key points; through the structure feature extractor, a first structural feature vector between key points corresponding to the image to be detected is extracted; wherein, the key points corresponding to the image to be detected include unoccluded key points and occluded key points in the image to be detected; a first target vector adjacent to the first structural feature vector is determined from the preset feature vectors; and by means of the decoder, the first target vector is decoded into position coordinates of key points corresponding to the image to be detected.
[0020] In the technical solution of the present disclosure, first, a first structural feature vector between occluded and unoccluded key points in the image to be detected can be extracted by means of a structure feature extractor; then, a first target vector closest to the first structure can be determined from the preset feature vectors in the codebook; and finally, based on the first target vector, the position coordinates of the occluded and unoccluded key points in the image to be detected can be decoded by means of a decoder. Since the preset feature vector can relatively accurately represent the structural features between the preset full-scale feature points, by replacing the first structural feature vector with the first target vector for predicting the position coordinates of the key points, not only can the position of the unoccluded key points be predicted, but also the position of the occluded key points can be predicted relatively accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original and elements are not necessarily drawn to scale.
[0022] Figure 1 A flowchart of a key point detection method provided by an embodiment of the present disclosure;
[0023] Figure 2 A schematic diagram of a codebook in a key point detection method provided by an embodiment of the present disclosure;
[0024] Figure 3 A schematic block diagram of the process of constructing a detection model in a key point detection method provided by an embodiment of the present disclosure;
[0025] Figure 4 A flowchart of the first construction stage of a detection model in a key point detection method provided by an embodiment of the present disclosure;
[0026] Figure 5 A flowchart of the second construction stage of a detection model in a key point detection method provided by an embodiment of the present disclosure;
[0027] Figure 6 A schematic diagram of the structure of a key point detection device provided by an embodiment of the present disclosure;
[0028] Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Specific Embodiments
[0029] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0030] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0031] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0032] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0034] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0035] Figure 1 The flowchart of a key point detection method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of key point detection, especially applicable to the situation of detecting occluded and unoccluded key points when there are occluded key points. This method can be executed by a key point detection device, which can be implemented in the form of software and / or hardware, and the device can be configured in an electronic device, such as configured in a computer.
[0036] As Figure 1 shown, the key point detection method provided in this embodiment may include:
[0037] S110. Input the image to be detected into the detection model; wherein, the detection model includes a structure feature extractor, a codebook and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points.
[0038] In the image key point detection scenario, at least one key point to be detected can be predefined in advance. In the embodiments of the present disclosure, the preset full-scale key points can represent all the preset number of key points to be detected. Exemplarily, the preset full-scale key points can include the whole body key points of a human body, such as 133 whole body key points of a human body: 68 face key points, 42 hand key points, 17 bone key points and 6 foot key points.
[0039] Exemplarily, Figure 2 The schematic diagram of the codebook in a key point detection method provided by an embodiment of the present disclosure. SeeFigure 2 , the codebook can be regarded as a matrix with dimensions of M×N, and the codebook can contain m (m≤M) preset feature vectors ( Figure 2 which can be represented by e1-em in ). The dimension of each preset feature vector can be N-dimensional. Among them, the preset feature vector can characterize the structural features between the preset full set of key points. It can be understood that through at least one preset feature vector in the codebook, the structural features between the preset full set of key points can be characterized. Among them, the structural features can be understood as the relative position relationship features between the key points.
[0040] In the embodiments of the present disclosure, the display range in the image to be detected can include at least some of the preset full set of key points, and there can be occluded key points among at least some of the key points in the image to be detected. Exemplarily, when the preset full set of key points is 133 full-body human key points, the image to be detected can be a full-body image of a human or a half-body image, and there can be occluded human key points in the image to be detected. Among them, the image to be detected can be input into a pre-constructed detection model to perform the processing operations of the following steps.
[0041] S120. Through a structural feature extractor, extract the first structural feature vector between the key points corresponding to the image to be detected.
[0042] In the embodiments of the present disclosure, the key points corresponding to the image to be detected can include the unoccluded key points and the occluded key points in the image to be detected.
[0043] Among them, the image to be detected input into the detection model can first pass through a size change network structure (resize) to change the size of the image to be detected to the preset size of the sample image used in the model construction process to ensure prediction accuracy. The detected image after size change can then pass through the structural feature extractor part in the detection model. Among them, the structural feature extractor can include an existing feature extraction network structure, for example, it can include a backbone network for extracting image features and a classification head network (class head, cls head) for adjusting the extracted feature representation, etc.
[0044] Since even if there are occluded key points in the image to be detected, the structural features between the occluded key points and the unoccluded key points still objectively exist, the structural feature extractor can extract all the structural features between the unoccluded key points and the occluded key points in the image to be detected to obtain at least one first structural feature vector. Exemplarily, when the preset full set of key points is 133 full-body human key points, 64 first structural feature vectors can be extracted to characterize the structural features of the key points corresponding to the image to be detected.
[0045] S130. Determine a first target vector that is adjacent to the first structural feature vector from the preset feature vectors.
[0046] In the embodiments of the present disclosure, when there are occluded key points in the image to be detected, the first structural feature vectors extracted may be deviated to a certain extent. Since the preset feature vectors in the constructed codebook can more accurately represent the structural relationship between key points, the first target vectors adjacent to the first structural feature vectors can be respectively selected from each preset feature vector.
[0047] Among them, the first target vector adjacent to each first structure feature can be determined based on existing vector similarity calculation methods (such as cosine similarity calculation methods, etc.). Exemplarily, when 64 first structural feature vectors are extracted, for each first structural feature vector, the most similar feature vector can be determined from each preset feature vector as the corresponding first target vector.
[0048] S140. Through the decoder, decode the first target vector into the position coordinates of the key points corresponding to the image to be detected.
[0049] After determining the first target vectors corresponding to the first structural feature vectors, the corresponding first structural feature vectors can be replaced with the first target vectors, and the first target vectors can be input into the decoder. Among them, the decoder may include an existing decoding network structure, such as a transformer decoder, etc.
[0050] By inputting the more accurate first target vector into the decoder to decode the position coordinates of the key points corresponding to the image to be detected, not only can the position of the unoccluded key points be predicted, but also the position of the occluded key points can be predicted more accurately. Among them, the position coordinates may refer to the pixel coordinates in the image to be detected.
[0051] The technical solution of the embodiments of the present disclosure can first extract the first structural feature vectors between the occluded and unoccluded key points in the image to be detected through the structural feature extractor; then determine the first target vector that is closest to the first structural feature from the preset feature vectors in the codebook; finally, through the decoder, decode the position coordinates of the occluded and unoccluded key points in the image to be detected based on the first target vector. Since the preset feature vectors can more accurately represent the structural features between the preset full-scale feature points, by replacing the first structural feature vectors with the first target vectors to predict the position coordinates of the key points, not only can the position of the unoccluded key points be predicted, but also the position of the occluded key points can be predicted more accurately.
[0052] The optional solutions in the key point detection method provided in the embodiments of the present disclosure can be combined with those in the above embodiments. The key point detection method provided in this embodiment describes the model construction process in detail. The model construction process can be divided into two construction stages: through the first construction stage, the codebook can construct a preset feature vector that more accurately represents the structural features between the preset full amount of feature points, enabling the encoder to extract a feature vector consistent with the preset feature vector based on the true position coordinates of the key points, and enabling the decoder to decode a position coordinate consistent with the true position coordinates of the key points based on the preset feature vector in the codebook; in the second construction stage, with the codebook and decoder parameters fixed, the structural feature extractor can extract the structural features between the key points from the image.
[0053] In the embodiments of the present disclosure, taking 133 full-body human key points as the preset full amount of key points as an example, the first construction stage and the second construction stage will be described. For the first construction stage and the second construction stage in the case of other types of preset full amounts of key points, reference can also be made to the relevant descriptions in the embodiments of the present disclosure, and specific limitations are not made here.
[0054] Exemplarily, Figure 3 is a schematic block diagram of the detection model construction process in a key point detection method provided in an embodiment of the present disclosure. As Figure 3 shown, the detection model construction process in the key point detection method provided in this embodiment may include:
[0055] The first construction stage: constructing the encoder, the codebook, and the decoder based on the actual position coordinates of the preset full amount of key points in the sample image;
[0056] The second construction stage: constructing the structural feature extractor based on the sample image, the constructed codebook, and the constructed decoder.
[0057] Among them, the size of the sample image can be pre-processed to be a preset size. The first construction stage is the stage of constructing the encoder, the codebook, and the decoder based on the actual position coordinates of the preset full amount of key points in the sample image (that is, the actual pixel coordinates of 133 full-body human key points).
[0058] Among them, the construction objectives of the first construction stage may include: enabling the codebook to construct a preset feature vector that can accurately represent the structural features between preset full-scale feature points; enabling the encoder to extract a feature vector consistent with the preset feature vector based on the true position coordinates of the key points; enabling the decoder to decode a position coordinate consistent with the true position coordinates of the preset full-scale key points (i.e., the predicted pixel coordinates of 133 full-body human key points) based on the preset feature vector in the codebook. Among them, based on the existing loss function, the first loss can be constructed according to the construction objectives of the first construction stage to model the structural relationship between the preset full-scale key points and realize the reconstruction of the position coordinates of the key points.
[0059] Among them, the second construction stage is the stage of constructing a structure feature extractor according to the sample image while fixing the constructed codebook and decoder parameters. Among them, the objectives of the second construction stage may include enabling the structure feature extractor to extract the structural features that can represent the key points from the image. Among them, based on the existing loss function, the second loss can be constructed according to the construction objectives of the second construction stage to realize the extraction of image structure features.
[0060] Exemplarily, Figure 4 is a schematic flowchart of the first construction stage of the detection model in a key point detection method provided by an embodiment of the present disclosure. Refer to Figure 4 , in some implementation manners, constructing the encoder, codebook, and decoder based on the actual position coordinates of the preset full-scale key points in the sample image may include:
[0061] S410. Input the actual position coordinates of the preset full-scale key points in the sample image into the encoder, and output the second structure feature vector of the preset full-scale key points in the sample image through the encoder.
[0062] Refer to Figure 3 , the two-dimensional actual position coordinates [x, y] of 133 full-body human key points in the sample image can be input into the encoder. The encoder may include a full connection layer (fc) and a transformer encoding network, etc. First, the input 133 two-dimensional coordinates can pass through a full connection layer (fc) to represent each two-dimensional coordinate as an embedding vector respectively; then, the 133 embeddings can be input into the transformer encoder to encode the structural features between the embeddings; finally, the structural features can be input into another full connection layer (fc) to raise the dimensionality of the structural information to the feature dimension of the codebook, and at least one second structure feature vector (for example, 64 second structure feature vectors) can be obtained.
[0063] S420. Determine a second target vector adjacent to the second structural feature vector from the current feature vectors included in the codebook. The preset feature vectors included in the codebook are constructed based on historical second structural feature vectors.
[0064] See Figure 3 , for each second structural feature vector, the closest second target vector can be selected from the current feature vectors in the codebook for replacement (for example, 64 second target vectors are obtained). The process of selecting the closest second target vector from the current feature vectors can refer to the process of determining the first target vector adjacent to the first structural feature vector from the preset feature vectors, which will not be elaborated here.
[0065] See Figure 2 , before constructing the detection model, the preset feature vectors in the codebook can be initialized with e1-em based on experience or experimental values. Since the feature vectors in the codebook can essentially be considered as a clustered representation of the structural features of the preset full set of key points. During the construction process of the detection model, ei can be clustered and updated according to the historical second structural feature vectors adjacent to each ei (i ∈ [1, m]). When the construction of the detection model is completed, the current feature vectors in the codebook can be used as the final preset feature vectors.
[0066] In some alternative implementation manners, the process of constructing the preset feature vectors included in the codebook may include: after determining the second target vector, obtaining each historical adjacent second structural feature vector corresponding to the vector position of the second target vector; setting update weights for each historical adjacent second structural feature vector according to the chronological order before and after the construction process of the detection model; and updating the second target vector according to each historical adjacent second structural feature vector and the update weights of each historical adjacent second structural feature vector.
[0067] In these alternative implementation manners, the vector position can be understood as the sorting of the feature vectors in the codebook. See Figure 2 , after determining the second target vector ei from the current feature vectors e1-em each time, each historical adjacent second structural feature vector corresponding to the vector position i of ei can be obtained (which may include the second structural feature vector corresponding to the current feature vector ei). And, according to the chronological order before and after the construction of the detection model, larger update weights can be set for the later historical adjacent second structural feature vectors, and smaller update weights can be set for the earlier historical adjacent second structural feature vectors, so as to set larger weights for the vectors that have a greater impact on ei. Furthermore, each historical adjacent second structural feature vector can be weighted and summed according to its respective update weight, and the weighted sum result can be used to update ei, so that the structural features between the key points represented by ei are more accurate.
[0068] Exemplarily, weights can be set for each of the historical neighboring second structural feature vectors corresponding to the second target vector ei based on the Exponential Moving Average (EMA), and weighted summation can be performed. In addition, other methods for updating the feature vectors in the codebook can also be applied here. For example, the k-means clustering algorithm (k-means) can be used to set weights for each of the historical neighboring second structural feature vectors corresponding to the second target vector ei and perform weighted summation. This is not an exhaustive list here.
[0069] S430. Through the decoder, decode the second target vector into the first predicted position coordinates of the preset full set of key points in the sample image.
[0070] See Figure 3 , each second target vector can be input into the decoder. Among them, the decoder can be considered as an inverse encoder and can include a fully connected layer (fc) and a transformer decoder network, etc. Through the decoder that is inverse to the encoder, the first predicted position coordinates of the preset full set of key points (such as the predicted pixel coordinates of 133 full-body human key points) can be output based on the second target vector.
[0071] S440. Determine the first loss based on the actual position coordinates, the first predicted position coordinates, the second structural feature vectors, and the second target vectors, and construct the encoder and decoder based on the first loss.
[0072] Among them, the key point position reconstruction loss of the encoder and decoder can be determined based on the first predicted position coordinates and the actual position coordinates; the encoding loss of the encoder can be determined based on the second structural feature vectors and the second target vectors. Furthermore, the encoder and decoder can be constructed based on the key point position reconstruction loss and the encoding loss.
[0073] Exemplarily, in some ways, determining the first loss based on the actual position coordinates, the first predicted position coordinates, the second structural feature vectors, and the second target vectors may include: determining the first position loss based on the first predicted position coordinates and the actual position coordinates; determining the first structural representation loss based on the second structural feature vectors and the second target vectors; and determining the first loss based on the first position loss and the first structural representation loss.
[0074] Among them, the loss function in the first construction stage can be designed as follows:
[0075] L = smooth L 1(G1,G2)+a×∑|ti - sg(ci)|;
[0076] Among them, G1 represents the actual position coordinates, G2 represents the first predicted position coordinates, ti represents the second structural feature vector output by the encoder, ci represents the second target vector, a is an adjustable parameter, and sg represents stopping the gradient backpropagation. Since the feature vectors in the codebook are discretized and the gradient cannot be backpropagated, the encoder and decoder can be updated using the first loss.
[0077] Through the first construction stage, the codebook can construct a preset feature vector that more accurately represents the structural features between the preset full amount of feature points, enabling the encoder to extract a feature vector consistent with the preset feature vector based on the true position coordinates of the key points, and enabling the decoder to decode the position coordinates consistent with the true position coordinates of the key points based on the preset feature vector in the codebook.
[0078] Exemplarily, Figure 5 is a schematic flowchart of the second construction stage of the detection model in a key point detection method provided by an embodiment of the present disclosure. Refer to Figure 5 , in some implementation manners, based on the sample image, the constructed codebook, and the constructed decoder, constructing the structural feature extractor may include:
[0079] S510. Extract the third structural feature vector between the preset full amount of key points in the sample image through the structural feature extractor.
[0080] Refer to Figure 3 , the structural feature extractor may include a backbone network and a classification head network (classhead, cls head), etc. Among them, the sample image is the same as the sample image in the first construction stage. The sample image passes through the backbone of the structural feature extractor to extract the image features of the sample image; the extracted sample image features may pass through the cls head of the structural feature extractor to represent the extracted features in the form of vectors in the codebook.
[0081] S520. Determine the third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook.
[0082] Among them, the process of determining the third target vector adjacent to the third structural feature vector from the preset feature vectors can refer to the process of determining the first target vector adjacent to the first structural feature vector from the preset feature vectors, which will not be elaborated here.
[0083] S530. Decode the third target vector into the second predicted position coordinates through the constructed decoder.
[0084] In this embodiment, the process of decoding the third target vector into the second predicted position coordinates can refer to the process of decoding the first target vector into the position coordinates of the key points corresponding to the image to be detected, which will not be elaborated here.
[0085] S540. Determine a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and construct a structural feature extractor based on the second loss.
[0086] Among them, the second structural feature vector includes the structural feature vectors of the preset full amount of key points in the sample image output by the constructed encoder.
[0087] In this embodiment, the codebook and decoder parameters are fixed and will not be updated anymore. The second construction stage mainly updates the parameters of the structural feature extractor. Exemplarily, in some implementation manners, determining the second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector may include: determining a second position loss according to the second predicted position coordinates and the actual position coordinates; determining a second structural representation loss according to the third structural feature vector and the second structural feature vector; and determining the second loss according to the second position loss and the second structural representation loss.
[0088] Among them, the loss function in the second construction stage can be designed as follows:
[0089] L = smooth L 1(G1, G2)+CE(L1, L2);
[0090] Among them, G1 represents the actual coordinate position, G2 represents the second predicted coordinate position, CE(.) represents the cross-entropy loss function, L1 represents the third structural feature vector, and L2 represents the second structural feature vector. Thus, the parameter update of the structural feature extractor can be realized.
[0091] The technical solution of the embodiment of the present disclosure describes the model construction process in detail. The model construction process can be divided into two construction stages: through the first construction stage, the codebook can construct a preset feature vector that more accurately represents the structural features between the preset full-scale feature points, enabling the encoder to extract a feature vector consistent with the preset feature vector based on the true position coordinates of the key points, and enabling the decoder to decode a position coordinate consistent with the true position coordinates of the key points based on the preset feature vector in the codebook; in the second construction stage, with the codebook and decoder parameters fixed, the structural feature extractor can extract the structural features between the key points from the image. The key point detection method provided by the embodiment of the present disclosure belongs to the same general concept as the key point detection method provided by the above embodiment. Technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0092] Figure 6 FIG. 4 is a schematic structural diagram of a key point detection device provided by an embodiment of the present disclosure. The key point detection device provided by this embodiment is applicable to the situation of key point detection, especially applicable to the situation of detecting occluded and unoccluded key points when there are occluded key points.
[0093] As Figure 6 shown, the key point detection device provided by the embodiment of the present disclosure may include:
[0094] An input module 610, configured to input an image to be detected into a detection model; wherein, the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points;
[0095] A structural feature extraction module 620, configured to extract a first structural feature vector between key points corresponding to the image to be detected through the structural feature extractor; wherein, the key points corresponding to the image to be detected include unoccluded key points and occluded key points in the image to be detected;
[0096] A vector determination module 630, configured to determine a first target vector adjacent to the first structural feature vector from the preset feature vectors;
[0097] A position prediction module 640, configured to decode the first target vector into the position coordinates of the key points corresponding to the image to be detected through the decoder.
[0098] In some optional implementation manners, the key point detection device may further include: a model construction module, configured to construct a detection model;
[0099] wherein, the model construction module may include a first construction unit and a second construction unit;
[0100] The first construction unit can be used to construct an encoder, a codebook, and a decoder based on the actual position coordinates of preset full-scale key points in a sample image;
[0101] The second construction unit can be used to construct a structural feature extractor based on the sample image, the constructed codebook, and the constructed decoder.
[0102] In some optional implementation manners, the first construction unit can be used to:
[0103] Input the actual position coordinates of preset full-scale key points in the sample image into the encoder, and output a second structural feature vector of the preset full-scale key points in the sample image through the encoder;
[0104] Determine a second target vector adjacent to the second structural feature vector from the current feature vectors included in the codebook; wherein, the preset feature vectors included in the codebook are constructed based on historical second structural feature vectors;
[0105] Decode the second target vector into a first predicted position coordinate of preset full-scale key points in the sample image through the decoder;
[0106] Determine a first loss according to the actual position coordinate, the first predicted position coordinate, the second structural feature vector, and the second target vector, and construct the encoder and the decoder based on the first loss.
[0107] In some optional implementation manners, the first construction unit can be used to construct the preset feature vectors included in the codebook based on the following construction process:
[0108] After determining the second target vector, obtain each historical adjacent second structural feature vector corresponding to the vector position of the second target vector;
[0109] Set update weights for each historical adjacent second structural feature vector according to the front-back time sequence of the construction process of the detection model;
[0110] Update the second target vector according to each historical adjacent second structural feature vector and the update weights of each historical adjacent second structural feature vector.
[0111] In some optional implementation manners, the first construction unit can be used to:
[0112] Determine a first position loss according to the first predicted position coordinate and the actual position coordinate;
[0113] Determine a first structural representation loss according to the second structural feature vector and the second target vector;
[0114] Determine a first loss according to the first position loss and the first structural representation loss.
[0115] In some alternative implementations, the second construction unit can be used for:
[0116] Extracting a third structural feature vector between preset full - volume key points in a sample image through a structural feature extractor;
[0117] Determining a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook;
[0118] Decoding the third target vector into second predicted position coordinates through the constructed decoder;
[0119] Determining a second loss based on the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and constructing the structural feature extractor based on the second loss;
[0120] Wherein, the second structural feature vector includes the structural feature vectors of preset full - volume key points in the sample image based on the output of the constructed encoder.
[0121] In some alternative implementations, the preset full - volume key points include the key points of the whole human body.
[0122] The key - point detection device provided by the embodiments of the present disclosure can execute the key - point detection method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0123] It should be noted that the various units and modules included in the above - mentioned device are only divided according to functional logic, but are not limited to the above division as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.
[0124] Next, with reference to Figure 7 , which shows a schematic structural diagram of an electronic device (such as Figure 7 the terminal device or the server in) 700 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in - vehicle terminals (such as in - vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0125] As Figure 7As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0126] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0127] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the key point detection method of the embodiment of the present disclosure are executed.
[0128] The electronic device provided by the embodiment of the present disclosure and the key point detection method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0129] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the key point detection method provided by the above embodiment is implemented.
[0130] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0131] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0132] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0133] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:
[0134] Input the image to be detected into a detection model; wherein, the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points; extract a first structural feature vector between the key points corresponding to the image to be detected through the structural feature extractor; wherein, the key points corresponding to the image to be detected include unoccluded key points and occluded key points in the image to be detected; determine a first target vector adjacent to the first structural feature vector from the preset feature vectors; decode the first target vector into the position coordinates of the key points corresponding to the image to be detected through the decoder.
[0135] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet service provider using the Internet).
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0137] The units involved in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the names of the units and modules do not constitute a limitation to the units and modules themselves in some cases.
[0138] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0139] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] According to one or more embodiments of the present disclosure, a key point detection method is provided, the method comprising:
[0141] Inputting an image to be detected into a detection model; wherein the detection model includes a structure feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structural features between preset full-scale key points;
[0142] Extracting, by the structure feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unoccluded key points and occluded key points in the image to be detected;
[0143] Determining, from the preset feature vectors, a first target vector adjacent to the first structural feature vector;
[0144] Decoding, by the decoder, the first target vector into position coordinates of key points corresponding to the image to be detected.
[0145] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0146] In some alternative implementation manners, the process of constructing the detection model includes:
[0147] Constructing an encoder, the codebook, and the decoder based on the actual position coordinates of preset full-scale key points in a sample image;
[0148] Constructing the structure feature extractor based on the sample image, the constructed codebook, and the constructed decoder.
[0149] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0150] In some alternative implementation manners, the constructing the encoder, the codebook, and the decoder based on the actual position coordinates of preset full-scale key points in a sample image includes:
[0151] Inputting the actual position coordinates of the preset full-scale key points in the sample image into the encoder, and outputting a second structure feature vector of the preset full-scale key points in the sample image through the encoder;
[0152] Determining a second target vector adjacent to the second structure feature vector from the current feature vectors included in the codebook; wherein, the preset feature vectors included in the codebook are constructed based on historical second structure feature vectors;
[0153] Decoding the second target vector into a first predicted position coordinate of the preset full-scale key points in the sample image through the decoder;
[0154] Determining a first loss according to the actual position coordinate, the first predicted position coordinate, the second structure feature vector, and the second target vector, and constructing the encoder and the decoder based on the first loss.
[0155] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0156] In some alternative implementation manners, the process of constructing the preset feature vectors included in the codebook includes:
[0157] After determining the second target vector, obtaining respective historical adjacent second structure feature vectors corresponding to the vector position of the second target vector;
[0158] Setting update weights for the respective historical adjacent second structure feature vectors according to the front-to-back time sequence of the process of constructing the detection model.
[0159] Update the second target vector according to the historical neighboring second structural feature vectors and the update weights of the historical neighboring second structural feature vectors.
[0160] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0161] In some optional implementation manners, the determining the first loss according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector, and the second target vector includes:
[0162] Determine a first position loss according to the first predicted position coordinates and the actual position coordinates;
[0163] Determine a first structural characterization loss according to the second structural feature vector and the second target vector;
[0164] Determine the first loss according to the first position loss and the first structural characterization loss.
[0165] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0166] In some optional implementation manners, the constructing the structural feature extractor based on the sample image, the constructed codebook, and the constructed decoder includes:
[0167] Extract a third structural feature vector between preset full-scale key points in the sample image through the structural feature extractor;
[0168] Determine a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook;
[0169] Decode the third target vector into second predicted position coordinates through the constructed decoder;
[0170] Determine a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and construct the structural feature extractor based on the second loss;
[0171] Wherein, the second structural feature vector includes the structural feature vectors of preset full-scale key points in the sample image output by the constructed encoder.
[0172] According to one or more embodiments of the present disclosure, a key point detection method is provided, further including:
[0173] In some alternative implementations, the preset full set of key points includes the key points of the whole human body.
[0174] According to one or more embodiments of the present disclosure, a key point detection device is provided, and the device includes:
[0175] An input module, configured to input an image to be detected into a detection model; wherein, the detection model includes a structure feature extractor, a codebook, and a decoder, and the codebook contains preset feature vectors of the structure features between preset full set of key points;
[0176] A structure feature extraction module, configured to extract a first structure feature vector between the key points corresponding to the image to be detected through the structure feature extractor; wherein, the key points corresponding to the image to be detected include the unoccluded key points and the occluded key points in the image to be detected;
[0177] A vector determination module, configured to determine a first target vector adjacent to the first structure feature vector from the preset feature vectors;
[0178] A position prediction module, configured to decode the first target vector into the position coordinates of the key points corresponding to the image to be detected through the decoder.
[0179] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0180] In addition, although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0181] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A key point detection method, characterized in that: include: Inputting the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points; Extracting a first structural feature vector between key points corresponding to the image to be detected by the structural feature extractor; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected; Determine a first target vector adjacent to the first structural feature vector from the preset feature vectors; The first target vector is decoded into position coordinates of key points corresponding to the image to be detected by the decoder.
2. The method according to claim 1, characterized in that The construction process of the detection model includes: Based on the actual position coordinates of all preset key points in the sample image, the encoder, the codebook and the decoder are constructed; The structural feature extractor is constructed based on the sample image, the constructed codebook and the constructed decoder.
3. The method according to claim 2, characterized in that The encoder, the codebook and the decoder are constructed based on the actual position coordinates of all preset key points in the sample image, including: Inputting the actual position coordinates of all preset key points in the sample image into the encoder, and outputting the second structural feature vector of all preset key points in the sample image through the encoder; Determining a second target vector adjacent to the second structural feature vector from the current feature vector contained in the codebook; wherein the preset feature vector contained in the codebook is constructed based on the historical second structural feature vector; Decoding the second target vector into first predicted position coordinates of all preset key points in the sample image by the decoder; A first loss is determined according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector, and the encoder and the decoder are constructed based on the first loss.
4. The method according to claim 3, characterized in that The process of constructing the preset feature vector contained in the codebook includes: After determining the second target vector, obtaining each second structural feature vector of the historical neighborhood corresponding to the vector position of the second target vector; According to the time sequence of the construction process of the detection model, setting update weights for each of the historically adjacent second structural feature vectors; The second target vector is updated according to each of the historically adjacent second structural feature vectors and the update weights of each of the historically adjacent second structural feature vectors.
5. The method according to claim 3, characterized in that: The determining the first loss according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector comprises: Determining a first position loss according to the first predicted position coordinates and the actual position coordinates; Determining a first structural representation loss according to the second structural feature vector and the second target vector; A first loss is determined based on the first position loss and the first structural characterization loss.
6. The method according to claim 2, characterized in that The constructing of the structural feature extractor based on the sample image, the constructed codebook and the constructed decoder comprises: Extracting, by means of the structural feature extractor, a third structural feature vector between preset full-quantity key points in the sample image; Determine a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook; Decoding the third target vector into a second predicted position coordinate by the constructed decoder; Determine a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector and the second structural feature vector, and construct the structural feature extractor based on the second loss; The second structural feature vector includes a structural feature vector of all key points preset in the sample image output based on the constructed encoder.
7. The method according to any one of claims 1 to 6, characterized in that: The preset full amount of key points includes key points of the human body.
8. A key point detection device, characterized in that: include: An input module, used to input the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points; A structural feature extraction module, used to extract a first structural feature vector between key points corresponding to the image to be detected through the structural feature extractor; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected; A vector determination module, used to determine a first target vector adjacent to the first structural feature vector from the preset feature vector; The position prediction module is used to decode the first target vector into the position coordinates of the key point corresponding to the image to be detected through the decoder.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the key point detection method as described in any one of claims 1-7.
10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the key point detection method according to any one of claims 1 to 7 when executed by a computer processor.