Key point detection method and apparatus, electronic device and storage medium
By extracting the structural feature vectors of key points in the image and decoding with preset feature vectors, the problem of being unable to predict the location of the occluded key points in the prior art is solved, and the accurate position prediction of all key points in the image is achieved.
Patent Information
- Application Number
- PCT/CN2024/127884
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-22
AI Technical Summary
The prior art cannot accurately predict the location of occluded key points in an image.
The structural feature vector between the unoccluded and occluded key points in the image to be detected is extracted by the structural feature extractor, the target vector closest to it is determined using the preset feature vector, and the target vector is decoded by the decoder to obtain the position coordinates of the key points.
Accurate position prediction of unoccluded and obstructed key points in the image is achieved, and the accuracy of key point detection is improved.
Smart Images

Figure CN2024127884_22052025_PF_FP_ABST
Abstract
Description
Key point detection method, device, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 17, 2023, with application number 202311541020.5 and invention name “A key point detection method, device, electronic device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The embodiments of the present disclosure relate to the field of image detection technology, and in particular to a key point detection method, device, electronic device, and storage medium. Background Art
[0004] In the prior art, the detection of whole-body key points in an image is usually achieved based on a heat map or regression approach.
[0005] Summary of the Invention
[0006] In a first aspect, an embodiment of the present disclosure provides a key point detection method, comprising:
[0007] Inputting the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points;
[0008] Extracting, by the structural feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected;
[0009] Determine a first target vector adjacent to the first structural feature vector from the preset feature vectors;
[0010] The first target vector is decoded into position coordinates of key points corresponding to the image to be detected by the decoder.
[0011] In a second aspect, an embodiment of the present disclosure further provides a key point detection device, comprising:
[0012] An input module, configured to input an image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook includes a preset feature vector of a preset structural feature between all key points;
[0013] A structural feature extraction module, configured to extract, through the structural feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected;
[0014] a vector determination module, configured to determine a first target vector adjacent to the first structural feature vector from the preset feature vector;
[0015] The position prediction module is used to decode the first target vector into the position coordinates of the key point corresponding to the image to be detected through the decoder.
[0016] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0017] one or more processors;
[0018] a storage device for storing one or more programs,
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the key point detection method as described in any of the embodiments of the present disclosure.
[0020] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the key point detection method as described in any one of the embodiments of the present disclosure.
[0021] The technical solution of the embodiment of the present disclosure is to input the image to be detected into a detection model; wherein, the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between the full amount of key points; through the structural feature extractor, a first structural feature vector between the key points corresponding to the image to be detected is extracted; wherein, the key points corresponding to the image to be detected include unobstructed key points and obscured key points in the image to be detected; a first target vector adjacent to the first structural feature vector is determined from the preset feature vector; and through the decoder, the first target vector is decoded into the position coordinates of the key points corresponding to the image to be detected. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0023] FIG1 is a schematic diagram of a flow chart of a key point detection method provided by an embodiment of the present disclosure;
[0024] FIG2 is a schematic diagram of a codebook in a key point detection method provided by an embodiment of the present disclosure;
[0025] FIG3 is a schematic block diagram of a detection model building process in a key point detection method provided by an embodiment of the present disclosure;
[0026] FIG4 is a flow chart of the first phase of constructing a detection model in a key point detection method provided by an embodiment of the present disclosure;
[0027] FIG5 is a flow chart of the second construction phase of a detection model in a key point detection method provided by an embodiment of the present disclosure;
[0028] FIG6 is a schematic structural diagram of a key point detection device provided by an embodiment of the present disclosure;
[0029] FIG7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0031] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0035] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0036] In the related art, the detection of whole-body key points in an image is usually achieved based on a heat map or regression method. The existing technology cannot predict the position of occluded key points in an image. The embodiments of the present disclosure provide a key point detection method, device, electronic device and storage medium, which can more accurately predict the position of occluded key points. In the technical solution of the present disclosure, first, a structural feature extractor can be used to extract the first structural feature vector between the occluded and unoccluded key points in the image to be detected; then, the first target vector closest to the first structural feature can be determined from the preset feature vectors in the codebook; finally, a decoder can be used to decode the position coordinates of the occluded and unoccluded key points in the image to be detected based on the first target vector. Since the preset feature vector can more accurately characterize the structural features between the preset full feature points, the key point position coordinates are predicted by replacing the first structural feature vector with the first target vector. Not only can the position of the unoccluded key points be predicted, but the position of the occluded key points can also be predicted more accurately.
[0037] Figure 1 is a flow chart illustrating a key point detection method provided by an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to key point detection scenarios, and is particularly applicable to detecting both occluded and unoccluded key points in the presence of occluded key points. This method can be performed by a key point detection device, which can be implemented in software and / or hardware and can be configured in an electronic device, such as a computer.
[0038] As shown in FIG1 , the key point detection method provided in this embodiment may include:
[0039] S110. Input the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook includes a preset feature vector of the structural features between all preset key points.
[0040] In the image keypoint detection scenario, at least one keypoint to be detected can be predefined. In the disclosed embodiment, the preset full set of keypoints can represent a preset number of all keypoints to be detected. For example, the preset full set of keypoints can include full-body keypoints, such as 133 full-body keypoints: 68 facial keypoints, 42 hand keypoints, 17 skeletal keypoints, and 6 foot keypoints.
[0041] Exemplarily, FIG2 is a schematic diagram of a codebook in a key point detection method provided by an embodiment of the present disclosure. Referring to FIG2 , the codebook can be regarded as a matrix of M×N dimensions, and the codebook can contain m (m≤M) preset feature vectors (represented by e1-em in FIG2 ), and the dimension of each preset feature vector can be N dimensions. Among them, the preset feature vector can characterize the structural features between the preset full amount of key points. It can be understood that the structural features between the preset full amount of key points can be characterized by at least one preset feature vector in the codebook. Among them, the structural features can be understood as the relative position relationship features between the key points.
[0042] In the embodiment of the present disclosure, the display range of the image to be detected may include at least some of the preset full set of key points, and at least some of the key points in the image to be detected may include occluded key points. For example, when the preset full set of key points is 133 human body key points, the image to be detected may be a full-body image of the human body or a half-body image, and occluded human key points may exist in the image to be detected. The image to be detected may be input into a pre-built detection model to perform the processing operations of the following steps.
[0043] S120 , extracting a first structural feature vector between key points corresponding to the image to be detected by a structural feature extractor.
[0044] In the embodiment of the present disclosure, the key points corresponding to the image to be detected may include unobstructed key points and obstructed key points in the image to be detected.
[0045] The image to be tested, which is input into the detection model, can first be resized to the preset size of the sample image used in the model construction process to ensure prediction accuracy. The resized detection image can then be passed through the structural feature extractor part of the detection model. The structural feature extractor can include an existing feature extraction network structure, such as a backbone network for extracting image features and a classification head network (class head, cls head) for adjusting the extracted feature representation.
[0046] Even if there are occluded key points in the image to be detected, the structural features between the occluded key points and the unoccluded key points still exist objectively. Therefore, the structural feature extractor can extract all the structural features between the unoccluded key points and the occluded key points in the image to be detected, and obtain at least one first structural feature vector. For example, when the preset total number of key points is 133 key points of the human body, 64 first structural feature vectors can be extracted to represent the structural features of the key points corresponding to the image to be detected.
[0047] S130: Determine a first target vector adjacent to the first structural feature vector from the preset feature vectors.
[0048] In the disclosed embodiment, when there are occluded key points in the image to be detected, the extracted first structural feature vector may have a certain degree of deviation. Because the preset feature vectors in the constructed codebook can more accurately represent the structural relationship between key points, first target vectors adjacent to each first structural feature vector can be selected from each preset feature vector.
[0049] Here, a first target vector adjacent to each first structural feature can be determined based on an existing vector similarity calculation method (e.g., a cosine similarity calculation method). For example, when 64 first structural feature vectors are extracted, the most similar feature vector can be determined from the preset feature vectors for each first structural feature vector as the corresponding first target vector.
[0050] S140 . Decode the first target vector into position coordinates of key points corresponding to the image to be detected through a decoder.
[0051] After determining each first target vector corresponding to each first structural feature vector, each first target vector can be used to replace the corresponding first structural feature vector and input into a decoder. The decoder can include an existing decoding network structure, such as a transformer decoding network.
[0052] By inputting the more accurate first target vector into the decoder, the position coordinates of the key points corresponding to the image to be detected are decoded. This not only allows the position prediction of unobstructed key points, but also allows for more accurate prediction of the positions of occluded key points. The position coordinates can refer to pixel coordinates in the image to be detected.
[0053] The technical solution of the embodiment of the present disclosure is to first extract the first structural feature vector between the occluded and unoccluded key points in the image to be detected through a structural feature extractor; then, the first target vector closest to the first structural feature can be determined from the preset feature vectors in the codebook; finally, the position coordinates of the occluded and unoccluded key points in the image to be detected can be decoded by a decoder based on the first target vector. Since the preset feature vector can more accurately characterize the structural features between the preset full set of feature points, the position coordinates of the key points are predicted by replacing the first structural feature vector with the first target vector. This not only allows the position prediction of the unoccluded key points, but also allows the more accurate prediction of the positions of the occluded key points.
[0054] The various optional schemes in the key point detection method provided in the embodiment of the present disclosure and the above embodiment can be combined. The key point detection method provided in this embodiment describes the model construction process in detail. The model construction process can be divided into two construction stages: through the first construction stage, the code book can be constructed to more accurately characterize the preset feature vectors of the structural features between the preset full feature points, so that the encoder can extract the feature vectors consistent with the preset feature vectors based on the real position coordinates of the key points, and the decoder can decode the position coordinates consistent with the real position coordinates of the key points based on the preset feature vectors in the code book; in the second construction stage, the structural feature extractor can extract the structural features that can characterize the key points from the image while fixing the code book and decoder parameters.
[0055] In the embodiments of this disclosure, the first and second construction phases will be described using a method where all 133 key points are preset for the entire human body. For other types of methods where all key points are preset, the first and second construction phases can also be described in the relevant descriptions of the embodiments of this disclosure, and are not specifically limited here.
[0056] For example, FIG3 is a schematic block diagram of a detection model construction process in a key point detection method provided by an embodiment of the present disclosure. As shown in FIG3 , the detection model construction process in the key point detection method provided by this embodiment may include:
[0057] The first construction phase: Based on the actual position coordinates of all preset key points in the sample image, the encoder, codebook and decoder are constructed;
[0058] The second construction phase: Based on the sample image, the constructed codebook and the constructed decoder, the structural feature extractor is constructed.
[0059] The sample images can be resized in advance so that the sizes are all the same as the preset sizes. The first construction phase is to construct the encoder, codebook, and decoder based on the actual position coordinates of all preset key points in the sample images (i.e., the actual pixel coordinates of the 133 key points of the human body).
[0060] The construction objectives of the first construction phase may include: enabling the codebook to construct a preset feature vector that more accurately characterizes the structural features between the preset full set of feature points; enabling the encoder to extract a feature vector consistent with the preset feature vector based on the true position coordinates of the key points; and enabling the decoder to decode position coordinates consistent with the true position coordinates of the preset full set of key points based on the preset feature vectors in the codebook (i.e., the predicted pixel coordinates of the 133 key points of the human body). Based on the existing loss function, a first loss may be constructed according to the construction objectives of the first construction phase to model the structural relationship between the preset full set of key points and reconstruct the key point position coordinates.
[0061] The second construction phase involves constructing a structural feature extractor based on the sample image while maintaining the constructed codebook and decoder parameters. The goal of the second construction phase may include enabling the structural feature extractor to extract structural features that characterize key points from the image. A second loss function may be constructed based on the construction goal of the second construction phase to achieve image structural feature extraction.
[0062] For example, FIG4 is a flow chart of the first phase of constructing a detection model in a key point detection method provided by an embodiment of the present disclosure. Referring to FIG4 , in some implementations, constructing an encoder, a codebook, and a decoder based on the actual position coordinates of all preset key points in a sample image may include:
[0063] S410: Input the actual position coordinates of all preset key points in the sample image into an encoder, and output a second structural feature vector of all preset key points in the sample image through the encoder.
[0064] Referring to Figure 3, the two-dimensional actual position coordinates [x, y] of 133 key points of the human body in the sample image can be input into the encoder. The encoder may include a fully connected layer (fc) and a transformer encoder network. First, the 133 input two-dimensional coordinates can pass through a fully connected layer (fc) to represent each two-dimensional coordinate as an embedded vector (embedding); then, the 133 embeddings can be input into the transformer encoder to encode the structural features between the embeddings; finally, the structural features can be input into another fully connected layer (fc) to upgrade the structural information to the feature dimension of the codebook, thereby obtaining at least one second structural feature vector (for example, 64 second structural feature vectors).
[0065] S420: Determine a second target vector adjacent to the second structural feature vector from the current feature vectors included in the codebook; and construct a preset feature vector included in the codebook based on the historical second structural feature vector.
[0066] Referring to FIG3 , for each second structural eigenvector, the nearest second target vector can be selected from the current eigenvector in the codebook as a replacement (e.g., 64 second target vectors are obtained). The process of selecting the nearest second target vector from the current eigenvector can be referred to as the process of determining the first target vector adjacent to the first structural eigenvector from the preset eigenvectors, and is not further described here.
[0067] As shown in Figure 2, the preset feature vectors in the codebook can be initialized based on empirical or experimental values before the detection model is constructed. Because the feature vectors in the codebook can essentially be considered a clustered representation of the structural features of the preset full set of key points, during the detection model construction process, each ei (i∈[1,m]) can be clustered and updated based on the historical second structural feature vector adjacent to it. When the detection model is completed, the current feature vector in the codebook can be used as the final preset feature vector.
[0068] In some optional implementations, the process of constructing the preset feature vectors included in the codebook may include: after determining the second target vector, obtaining each historically adjacent second structural feature vector corresponding to the vector position of the second target vector; setting update weights for each historically adjacent second structural feature vector according to the sequence of the detection model construction process; and updating the second target vector based on each historically adjacent second structural feature vector and the update weights of each historically adjacent second structural feature vector.
[0069] In these optional implementations, the vector position can be understood as the ordering of the feature vectors in the codebook. Referring to Figure 2, each time the second target vector ei is determined from the current feature vector e1-em, the second structural feature vectors of the historical neighborhood corresponding to the vector position i of ei can be obtained (which may include the second structural feature vector corresponding to the current feature vector ei). In addition, according to the before-after sequence of the construction of the detection model, the second structural feature vectors of the later historical neighborhood can be set with a larger update weight, and the second structural feature vectors of the early historical neighborhood can be set with a smaller update weight, so as to set a larger weight for the vector that has a greater impact on ei. Furthermore, the second structural feature vectors of the historical neighborhood can be weighted and summed according to their respective update weights, and the weighted summation result can be used to update ei, so that the structural features between the key points represented by ei are more accurate.
[0070] For example, weights can be set for each of the second structural feature vectors in the historical neighborhood corresponding to the second target vector ei based on an exponential moving average (EMA) indicator, and a weighted sum can be performed. In addition, other methods for updating feature vectors in the codebook can also be applied here, such as using a k-means clustering algorithm (k-means) to set weights for each of the second structural feature vectors in the historical neighborhood corresponding to the second target vector ei, and a weighted sum can be performed. These methods are not exhaustive here.
[0071] S430: Decode the second target vector into first predicted position coordinates of all preset key points in the sample image through a decoder.
[0072] As shown in Figure 3, each second target vector can be input into a decoder. The decoder can be considered the reverse of the encoder and can include, for example, a fully connected layer (FC) and a transformer decoder network. The decoder, which is the reverse of the encoder, can output the first predicted position coordinates of all preset key points (e.g., the predicted pixel coordinates of 133 key points on the entire human body) based on the second target vector.
[0073] S440: Determine a first loss according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector, and the second target vector, and construct an encoder and a decoder based on the first loss.
[0074] The key point position reconstruction loss of the encoder and decoder can be determined based on the first predicted position coordinates and the actual position coordinates; the encoding loss of the encoder can be determined based on the second structural feature vector and the second target vector. Furthermore, the encoder and decoder can be constructed based on the key point position reconstruction loss and the encoding loss.
[0075] Exemplarily, in some embodiments, determining the first loss based on the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector may include: determining the first position loss based on the first predicted position coordinates and the actual position coordinates; determining the first structural characterization loss based on the second structural feature vector and the second target vector; determining the first loss based on the first position loss and the first structural characterization loss.
[0076] The loss function of the first construction stage can be designed as follows: L = smooth L 1(G1,G2)+a×∑|ti-sg(ci)|;
[0077] Where G1 represents the actual position coordinates, G2 represents the first predicted position coordinates, ti represents the second structural feature vector output by the encoder, ci represents the second target vector, a is an adjustable parameter, and sg indicates the stop gradient backpropagation. Because the feature vectors in the codebook are discretized, gradient backpropagation cannot be performed. Therefore, the first loss can be used to update the encoder and decoder.
[0078] Through the first construction stage, the codebook can construct a preset feature vector that more accurately characterizes the structural features between the preset full-quantity feature points, so that the encoder can extract a feature vector consistent with the preset feature vector based on the real position coordinates of the key point, and the decoder can decode the position coordinates consistent with the real position coordinates of the key point based on the preset feature vector in the codebook.
[0079] For example, FIG5 is a flow chart of the second construction phase of the detection model in a key point detection method provided by an embodiment of the present disclosure. Referring to FIG5 , in some implementations, constructing a structural feature extractor based on a sample image, a constructed codebook, and a constructed decoder may include:
[0080] S510 , extracting a third structural feature vector between all preset key points in the sample image through a structural feature extractor.
[0081] As shown in Figure 3, the structural feature extractor may include a backbone network and a classification head network (CLS head). The sample image is consistent with the sample image in the first construction phase. The sample image passes through the backbone of the structural feature extractor to extract image features of the sample image. The extracted sample image features can then pass through the CLS head of the structural feature extractor to represent the extracted features as vectors in the codebook.
[0082] S520: Determine a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook.
[0083] The process of determining the third target vector adjacent to the third structural feature vector from the preset feature vector can refer to the process of determining the first target vector adjacent to the first structural feature vector from the preset feature vector, and will not be repeated here.
[0084] S530: Decode the third target vector into the second predicted position coordinates through the constructed decoder.
[0085] In this embodiment, the process of decoding the third target vector into the second predicted position coordinates through the decoder can refer to the process of decoding the first target vector into the position coordinates of the key point corresponding to the image to be detected, and will not be repeated here.
[0086] S540: Determine a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and construct a structural feature extractor based on the second loss.
[0087] The second structural feature vector includes a structural feature vector of all key points preset in the sample image output by the constructed encoder.
[0088] In this embodiment, the codebook and decoder parameters are fixed and will not be updated. The second construction phase mainly updates the parameters of the structural feature extractor. Exemplarily, in some implementations, determining the second loss based on the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector may include: determining the second position loss based on the second predicted position coordinates and the actual position coordinates; determining the second structural representation loss based on the third structural feature vector and the second structural feature vector; determining the second loss based on the second position loss and the second structural representation loss.
[0089] The loss function of the second construction phase can be designed as follows: L = smooth L 1(G1,G2)+CE(L1,L2);
[0090] Where G1 represents the actual coordinate position, G2 represents the second predicted coordinate position, CE(.) represents the cross entropy loss function, L1 represents the third structural feature vector, and L2 represents the second structural feature vector. This allows the parameters of the structural feature extractor to be updated.
[0091] The technical solution of the embodiment of the present disclosure describes the model construction process in detail. The model construction process can be divided into two construction stages: through the first construction stage, the codebook can be constructed to more accurately characterize the preset feature vectors of the structural features between the preset full feature points, so that the encoder can extract the feature vectors consistent with the preset feature vectors based on the real position coordinates of the key points, and the decoder can decode the position coordinates consistent with the real position coordinates of the key points based on the preset feature vectors in the codebook; in the second construction stage, the structural feature extractor can extract the structural features that can characterize the key points from the image under the condition of fixing the codebook and decoder parameters. The key point detection method provided in the embodiment of the present disclosure and the key point detection method provided in the above embodiment belong to the same public concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0092] Figure 6 is a schematic diagram of the structure of a key point detection device provided by an embodiment of the present disclosure. The key point detection device provided by this embodiment is suitable for key point detection, and is particularly suitable for detecting both blocked and unblocked key points when there are blocked key points.
[0093] As shown in FIG6 , the key point detection device provided by the embodiment of the present disclosure may include:
[0094] Input module 610, configured to input the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook includes a preset feature vector of the structural features between all preset key points;
[0095] The structural feature extraction module 620 is configured to extract, through a structural feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected;
[0096] A vector determination module 630 is configured to determine a first target vector adjacent to the first structural feature vector from a preset feature vector;
[0097] The position prediction module 640 is configured to decode the first target vector into position coordinates of key points corresponding to the image to be detected through a decoder.
[0098] In some optional implementations, the key point detection device may further include: a model building module for building a detection model;
[0099] The model construction module may include a first construction unit and a second construction unit;
[0100] The first construction unit can be used to construct an encoder, a codebook, and a decoder based on the actual position coordinates of all preset key points in the sample image;
[0101] The second construction unit can be used to construct a structural feature extractor based on the sample image, the constructed codebook and the constructed decoder.
[0102] In some optional implementations, the first building block may be used to:
[0103] Input the actual position coordinates of all preset key points in the sample image into the encoder, and output the second structural feature vector of all preset key points in the sample image through the encoder;
[0104] Determining a second target vector adjacent to the second structural feature vector from the current feature vectors included in the codebook; wherein the preset feature vectors included in the codebook are constructed based on the historical second structural feature vectors;
[0105] Decoding the second target vector into first predicted position coordinates of all preset key points in the sample image through a decoder;
[0106] A first loss is determined according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector, and an encoder and a decoder are constructed based on the first loss.
[0107] In some optional implementations, the first constructing unit may be configured to construct a preset feature vector included in the codebook based on the following construction process:
[0108] After determining the second target vector, obtaining each of the historically adjacent second structural feature vectors corresponding to the vector position of the second target vector;
[0109] According to the time sequence of the construction process of the detection model, an update weight is set for each of the historically adjacent second structural feature vectors;
[0110] The second target vector is updated according to each of the historically adjacent second structural feature vectors and the update weights of each of the historically adjacent second structural feature vectors.
[0111] In some optional implementations, the first building block may be used to:
[0112] determining a first position loss based on the first predicted position coordinates and the actual position coordinates;
[0113] determining a first structural representation loss based on the second structural feature vector and the second target vector;
[0114] A first loss is determined based on the first position loss and the first structural characterization loss.
[0115] In some optional implementations, the second building block may be used to:
[0116] Extracting the third structural feature vector between all the preset key points in the sample image through the structural feature extractor;
[0117] Determining a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook;
[0118] Decoding the third target vector into the second predicted position coordinates through the constructed decoder;
[0119] Determining a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and constructing a structural feature extractor based on the second loss;
[0120] The second structural feature vector includes a structural feature vector of all key points preset in the sample image output by the constructed encoder.
[0121] In some optional implementations, the preset full set of key points includes key points of the entire human body.
[0122] The key point detection device provided in the embodiments of the present disclosure can execute the key point detection method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0123] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0124] Reference is now made to FIG7 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server in FIG7 ) 700 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG7 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0125] As shown in Figure 7, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0126] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.
[0127] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the key point detection method of the embodiment of the present disclosure are performed.
[0128] The electronic device provided by the embodiment of the present disclosure and the key point detection method provided by the above embodiment belong to the same disclosed concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0129] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the key point detection method provided in the above embodiment is implemented.
[0130] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0131] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0132] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0133] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0134] The image to be detected is input into the detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between the preset full set of key points; through the structural feature extractor, the first structural feature vector between the key points corresponding to the image to be detected is extracted; wherein the key points corresponding to the image to be detected include unobstructed key points and obscured key points in the image to be detected; a first target vector adjacent to the first structural feature vector is determined from the preset feature vector; through the decoder, the first target vector is decoded into the position coordinates of the key points corresponding to the image to be detected.
[0135] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the names of the units and modules do not, in certain circumstances, limit the units and modules themselves.
[0138] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.
[0139] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] According to one or more embodiments of the present disclosure, a key point detection method is provided, the method comprising:
[0141] Inputting the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points;
[0142] Extracting, by the structural feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected;
[0143] Determine a first target vector adjacent to the first structural feature vector from the preset feature vectors;
[0144] The first target vector is decoded into position coordinates of key points corresponding to the image to be detected by the decoder.
[0145] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0146] In some optional implementations, the process of constructing the detection model includes:
[0147] Constructing the encoder, the codebook, and the decoder based on the actual position coordinates of all preset key points in the sample image;
[0148] The structural feature extractor is constructed based on the sample image, the constructed codebook and the constructed decoder.
[0149] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0150] In some optional implementations, constructing the encoder, the codebook, and the decoder based on the actual position coordinates of all preset key points in the sample image includes:
[0151] Inputting the actual position coordinates of all preset key points in the sample image into the encoder, and outputting the second structural feature vector of all preset key points in the sample image through the encoder;
[0152] Determining a second target vector adjacent to the second structural feature vector from the current feature vector contained in the codebook; wherein the preset feature vector contained in the codebook is constructed based on the historical second structural feature vector;
[0153] Decoding the second target vector into first predicted position coordinates of all preset key points in the sample image by the decoder;
[0154] A first loss is determined according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector, and the second target vector, and the encoder and the decoder are constructed based on the first loss.
[0155] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0156] In some optional implementations, the process of constructing the preset feature vector contained in the codebook includes:
[0157] After determining the second target vector, obtaining each second structural feature vector of the historical neighborhood corresponding to the vector position of the second target vector;
[0158] Setting update weights for each of the historically adjacent second structural feature vectors according to a time sequence of the construction process of the detection model;
[0159] The second target vector is updated according to the historically adjacent second structural feature vectors and the update weights of the historically adjacent second structural feature vectors.
[0160] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0161] In some optional implementations, determining the first loss based on the actual position coordinates, the first predicted position coordinates, the second structural feature vector, and the second target vector includes:
[0162] determining a first position loss based on the first predicted position coordinates and the actual position coordinates;
[0163] determining a first structural representation loss based on the second structural feature vector and the second target vector;
[0164] A first loss is determined according to the first position loss and the first structural characterization loss.
[0165] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0166] In some optional implementations, constructing the structural feature extractor based on the sample image, the constructed codebook, and the constructed decoder includes:
[0167] Extracting, by the structural feature extractor, a third structural feature vector between all preset key points in the sample image;
[0168] Determining a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook;
[0169] Decoding the third target vector into a second predicted position coordinate by the constructed decoder;
[0170] Determining a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector, and the second structural feature vector, and constructing the structural feature extractor based on the second loss;
[0171] The second structural feature vector includes a structural feature vector of all key points preset in the sample image output based on the constructed encoder.
[0172] According to one or more embodiments of the present disclosure, a key point detection method is provided, further comprising:
[0173] In some optional implementations, the preset full set of key points includes key points of the entire human body.
[0174] According to one or more embodiments of the present disclosure, a key point detection device is provided, the device comprising:
[0175] An input module, configured to input an image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook, and a decoder, and the codebook includes a preset feature vector of a preset structural feature between all key points;
[0176] A structural feature extraction module, configured to extract, through the structural feature extractor, a first structural feature vector between key points corresponding to the image to be detected; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected;
[0177] a vector determination module, configured to determine a first target vector adjacent to the first structural feature vector from the preset feature vector;
[0178] The position prediction module is used to decode the first target vector into the position coordinates of the key point corresponding to the image to be detected through the decoder.
[0179] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0180] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0181] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A key point detection method, comprising: Inputting the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points; Extracting a first structural feature vector between key points corresponding to the image to be detected by the structural feature extractor; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected; Determine a first target vector adjacent to the first structural feature vector from the preset feature vectors; The first target vector is decoded into position coordinates of key points corresponding to the image to be detected by the decoder.
2. The method according to claim 1, wherein: The construction process of the detection model includes: Based on the actual position coordinates of all preset key points in the sample image, the encoder, the codebook and the decoder are constructed; The structural feature extractor is constructed based on the sample image, the constructed codebook and the constructed decoder.
3. The method according to claim 2, wherein: The encoder, the codebook and the decoder are constructed based on the actual position coordinates of all preset key points in the sample image, including: Inputting the actual position coordinates of all preset key points in the sample image into the encoder, and outputting the second structural feature vector of all preset key points in the sample image through the encoder; Determining a second target vector adjacent to the second structural feature vector from the current feature vector contained in the codebook; wherein the preset feature vector contained in the codebook is constructed based on the historical second structural feature vector; Decoding the second target vector into first predicted position coordinates of all preset key points in the sample image by the decoder; A first loss is determined according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector, and the encoder and the decoder are constructed based on the first loss.
4. The method according to claim 3, wherein: The process of constructing the preset feature vector contained in the codebook includes: After determining the second target vector, obtaining each second structural feature vector of the historical neighborhood corresponding to the vector position of the second target vector; According to the time sequence of the construction process of the detection model, setting update weights for each of the historically adjacent second structural feature vectors; According to each second structural feature vector of the historical neighborhood, and each second structural feature vector of the historical neighborhood The second target vector is updated by updating the weight of 5. The method according to claim 3, wherein: The determining the first loss according to the actual position coordinates, the first predicted position coordinates, the second structural feature vector and the second target vector comprises: Determining a first position loss according to the first predicted position coordinates and the actual position coordinates; Determining a first structural representation loss according to the second structural feature vector and the second target vector; A first loss is determined based on the first position loss and the first structural characterization loss.
6. The method according to claim 2, wherein: The constructing of the structural feature extractor based on the sample image, the constructed codebook and the constructed decoder comprises: Extracting, by means of the structural feature extractor, a third structural feature vector between preset full-quantity key points in the sample image; Determine a third target vector adjacent to the third structural feature vector from the preset feature vectors included in the constructed codebook; Decoding the third target vector into a second predicted position coordinate by the constructed decoder; Determine a second loss according to the actual position coordinates, the second predicted position coordinates, the third structural feature vector and the second structural feature vector, and construct the structural feature extractor based on the second loss; The second structural feature vector includes a structural feature vector of all key points preset in the sample image output based on the constructed encoder.
7. The method according to any one of claims 1 to 6, wherein: The preset full amount of key points includes key points of the human body.
8. A key point detection device, comprising: An input module, used to input the image to be detected into a detection model; wherein the detection model includes a structural feature extractor, a codebook and a decoder, and the codebook contains a preset feature vector of the structural features between all key points; A structural feature extraction module, used to extract a first structural feature vector between key points corresponding to the image to be detected through the structural feature extractor; wherein the key points corresponding to the image to be detected include unobstructed key points and obstructed key points in the image to be detected; A vector determination module, used to determine a first target vector adjacent to the first structural feature vector from the preset feature vector; The position prediction module is used to decode the first target vector into the position coordinates of the key point corresponding to the image to be detected through the decoder.
9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement A key point detection method as described in any one of claims 1-7 is provided.
10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the key point detection method according to any one of claims 1 to 7 when executed by a computer processor.
Citation Information
Patent Citations
A human body key point detection method based on a graph convolution network
CN109359568A
Human body key point detection method and system, electronic equipment and readable storage medium
CN114724183A
Template matching pose determination method and system and electronic equipment
CN116824181A
Method and Apparatus for uplink transmission in RRC_INACTIVE state
KR1020230158234A
Spinal-column curvature measurement method, apparatus, computer device, and storage medium
WO2021114622A1