Model training method, target integrity detection method and related apparatus

CN122737697APending Publication Date: 2026-09-11HEFEI IFLYTEK TOYCLOUD TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610887919.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,这种方式存在书页完整性判断误报率较高的问题

Benefits of technology

[0040]借由上述技术方案,本申请提供的模型训练方法、目标完整性检测方法、设备、存储介质和程序产品,通过构建差异化的可见性状态监督机制,实现了对关键点位置和置信度预测的有效解耦。具体的,该方法在获取标注关键点坐标及可见性状态(指示关键点存在但不可见的第一状态,或,指示关键点完全可见的第二状态)的训练样本后,在模型训练过程中,利用第一状态和第二状态的关键点共同计算位置损失,确保关键点检测模型能够学习到包括超出画面边界在内的关键点真实空间分布,维持了位置预测的准确性;而针对置信度损失的计算,方案采用了非对称的监督策略,即仅基于第二状态的关键点计算置信度损失,或者在对第一状态关键点施加的置信度监督强度低于第二状态关键点的前提下联合计算置信度损失。这种设计使得模型在优化过程中,一方面被引导去精确拟合所有存在关键点的几何位置,另一方面又被明确告知第一状态关键点不应具有高置信度响应,由此,训练得到的关键点检测模型在推理时,能够输出与关键点实际可见程度高度一致的置信度数值,即完全可见的关键点呈现高置信度,而存在但不可见的关键点呈现低置信度,避免了将关键点超出画面的不完整图像误判为完整的情况,提升了目标(比如,书页)完整性检测的准确率与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737697A_ABST
    Figure CN122737697A_ABST
Patent Text Reader

Abstract

This application discloses a model training method, a target integrity detection method, an apparatus, a storage medium, and a program product, relating to the field of artificial intelligence technology. The method includes: after acquiring training samples with labeled keypoint coordinates and visibility states (a first state indicating the existence but not visibility of keypoints, or a second state indicating the complete visibility of keypoints), during model training, the location loss is jointly calculated using keypoints from both the first and second states. This ensures that the keypoint detection model can learn the true spatial distribution of keypoints, including those beyond the image boundaries, maintaining the accuracy of location prediction. For the calculation of confidence loss, an asymmetric supervision strategy is adopted, i.e., the confidence loss is calculated only based on keypoints from the second state, or the confidence loss is jointly calculated under the premise that the confidence supervision intensity applied to keypoints from the first state is lower than that applied to keypoints from the second state. This application improves the accuracy and reliability of target integrity detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, a target integrity detection method, an apparatus, a storage medium, and a program product. Background Technology

[0002] In applications such as children's English education and smart reading, the system needs to perform a quality pre-check on the book page images taken by the user to determine whether the page is completely within the frame, that is, whether all four (or six) corners of the page are within the frame. If any corners are outside the boundaries, the user should be prompted to retake the photo to ensure the normal operation of subsequent OCR, image registration, and other functions.

[0003] The method for determining whether a book page is complete involves processing the page image using a keypoint detection model to detect the corner coordinates and their confidence levels, and then combining this with a threshold to determine the page's integrity. However, this method suffers from a high false alarm rate in determining page integrity. Summary of the Invention

[0004] In view of the above problems, this application provides a model training method, a target integrity detection method, a device, a storage medium, and a program product to reduce the false alarm rate of page integrity detection. The specific solution is as follows:

[0005] The first aspect of this application provides a method for training a keypoint detection model, including:

[0006] Acquire training samples, which include an image and annotation information of at least one key point of a target object in the image. The annotation information of each key point includes the coordinates and visibility state of the key point. The visibility state is one of at least two states including a first state and a second state. The first state indicates that the key point exists but is not visible, and the second state indicates that the key point is fully visible.

[0007] The image is input into the key point detection model to be trained to obtain the predicted coordinates and prediction confidence of the key points;

[0008] Based on the key points whose visibility states are the first state and the second state, calculate the position loss;

[0009] Calculate the confidence loss based on keypoints whose visibility state is the second state; or, calculate the confidence loss based on keypoints whose visibility state is the first state and the second state, in such a way that the confidence supervision intensity applied to keypoints whose visibility state is the first state is less than the confidence supervision intensity applied to keypoints whose visibility state is the second state.

[0010] The parameters of the keypoint detection model are updated based on the location loss and the confidence loss.

[0011] In one possible implementation, the confidence loss is calculated based on key points with visibility states of first and second states, including:

[0012] The first loss is calculated based on the prediction confidence of key points with visibility in the first state and the first confidence supervision target;

[0013] The second loss is calculated based on the prediction confidence of key points whose visibility state is the second state and the second confidence supervision target; the first confidence supervision target is less than the second confidence supervision target;

[0014] The confidence loss is obtained by weighted summation of the first loss and the second loss; the weight of the first loss is less than or equal to the weight of the second loss.

[0015] In one possible implementation, the confidence loss is calculated based on key points with visibility states of first and second states, including:

[0016] The first loss is calculated based on the prediction confidence of key points with visibility in the first state and the first confidence supervision target;

[0017] The second loss is calculated based on the prediction confidence of key points whose visibility state is the second state and the second confidence supervision target; the first confidence supervision target is equal to the second confidence supervision target;

[0018] The confidence loss is obtained by weighted summation of the first loss and the second loss; the weight of the first loss is less than the weight of the second loss.

[0019] In one possible implementation, before obtaining the training samples, the following is also included:

[0020] At the preset callback time point, the pre-registered callback function is invoked to configure the function used to calculate the confidence loss as the target loss function;

[0021] The target loss function is used to calculate the confidence loss based on keypoints with a visibility state of the second state; or, the confidence loss is calculated based on keypoints with visibility states of the first and second states in such a way that the confidence supervision intensity applied to keypoints with visibility states of the first state is less than the confidence supervision intensity applied to keypoints with visibility states of the second state.

[0022] One possible implementation also includes:

[0023] Before inputting the image into the key point detection model to be trained, it is randomly determined with a preset probability whether to perform online occlusion enhancement processing on the image;

[0024] If it is determined that the image will not be subjected to online occlusion enhancement processing, the image will be input into the key point detection model to be trained;

[0025] If it is determined that the image will be subjected to online occlusion enhancement processing, at least one key point is selected from the key points in the second visibility state of the image as the target key point for occlusion, and the visibility state of the target key point is updated to the first state, while the visibility states of other key points remain unchanged, thus obtaining the enhanced training sample; the image of the enhanced training sample is input into the key point detection model to be trained.

[0026] A second aspect of this application provides a target integrity detection method, comprising:

[0027] The keypoint detection model is used to detect keypoints of the target object in the image to be detected, and the predicted coordinates and prediction confidence of each keypoint of the target object are obtained; the keypoint detection model is trained by the keypoint detection model training method described above.

[0028] Based on the predicted coordinates and prediction confidence of each key point, it is determined whether the target object in the image to be detected is complete.

[0029] In one possible implementation, determining whether a target object in the image to be detected is complete based on the predicted coordinates and prediction confidence of each key point includes:

[0030] If the prediction confidence of each key point is greater than or equal to the confidence threshold, the predicted coordinates of each key point are all within the safe boundary, and the ratio of the area of ​​the polygon enclosed by each key point to the total area of ​​the image to be detected is greater than the ratio threshold, then the target object in the image to be detected is determined to be complete; otherwise, the target object in the image to be detected is determined to be incomplete.

[0031] The security boundary is the boundary formed by indenting the physical boundary of the image to be detected by a preset proportion.

[0032] One possible implementation also includes:

[0033] If the predicted coordinates and prediction confidence of each key point of at least two target objects are obtained, calculate the area of ​​the detection box for each target object;

[0034] The target object corresponding to the largest detection box is identified as the object that needs to be inspected for integrity.

[0035] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the key point detection model training method of the first aspect or any implementation thereof, and / or implement the target integrity detection method of the second aspect or any implementation thereof.

[0036] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0037] The memory is used to store computer programs;

[0038] The processor is used to execute the computer program so that the electronic device can implement the key point detection model training method of the first aspect or any implementation of the first aspect, and / or implement the target integrity detection method of the second aspect or any implementation of the second aspect.

[0039] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the key point detection model training method of the first aspect or any implementation thereof, and / or to implement the target integrity detection method of the second aspect or any implementation thereof.

[0040] By employing the aforementioned technical solutions, the model training method, target integrity detection method, device, storage medium, and program products provided in this application achieve effective decoupling of keypoint location and confidence prediction through the construction of a differentiated visibility state supervision mechanism. Specifically, after acquiring training samples with labeled keypoint coordinates and visibility states (a first state indicating the existence but not visibility of the keypoint, or a second state indicating the complete visibility of the keypoint), the method uses keypoints from both the first and second states to jointly calculate the location loss during model training. This ensures that the keypoint detection model can learn the true spatial distribution of keypoints, including those beyond the image boundaries, thus maintaining the accuracy of location prediction. For the calculation of confidence loss, the solution adopts an asymmetric supervision strategy, i.e., calculating the confidence loss only based on keypoints from the second state, or jointly calculating the confidence loss under the premise that the confidence supervision intensity applied to keypoints from the first state is lower than that applied to keypoints from the second state. This design guides the model during optimization to accurately fit the geometric positions of all keypoints while explicitly informing it that keypoints in the first state should not have high confidence responses. As a result, the trained keypoint detection model can output confidence values ​​that are highly consistent with the actual visibility of keypoints during inference. That is, fully visible keypoints have high confidence, while keypoints that exist but are not visible have low confidence. This avoids misjudging incomplete images where keypoints extend beyond the frame as complete images, thus improving the accuracy and reliability of target (e.g., book pages) integrity detection. Attached Figure Description

[0041] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0042] Figure 1 A flowchart illustrating an implementation of the keypoint detection model training method provided in this application;

[0043] Figure 2 A flowchart illustrating an implementation of the confidence loss calculation for key points based on visibility states as the first and second states provided in this application;

[0044] Figure 3 Another implementation flowchart for calculating confidence loss based on visibility state as the first and second states provided in this application;

[0045] Figure 4 A flowchart illustrating an implementation of the target integrity detection method provided in this application;

[0046] Figure 5A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0047] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0048] As will be known to those skilled in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0049] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0050] During training, the labeled data for image samples in a keypoint detection model includes: the coordinates (i.e., ground truth coordinates) of keypoints (such as the corners of a book) and the visibility (i.e., true visibility) of the keypoints. Keypoint visibility is typically categorized into three states: v=0 indicates the keypoint does not exist (the target itself does not have this keypoint), v=1 indicates the keypoint exists but is not visible (the keypoint is outside the frame boundary or occluded), and v=2 indicates the keypoint is fully visible. In the book page shooting scenario, "incomplete pages" precisely corresponds to the situation where the corners of the page extend beyond the frame (v=1), and this type of sample accounts for 30% to 60% of real-world datasets. The loss function during training includes positional loss and confidence loss, both calculated based on keypoints with v=1 and v=2.

[0051] The inventors of this application discovered that when calculating the confidence loss, the confidence supervision target for keypoints with v=1 and v=2 is the same (e.g., both are 1.0), aiming to make the model believe that as long as the keypoint physically exists, regardless of whether it is visible, its confidence should be high. This leads to the trained keypoint detection model still outputting high confidence values ​​for corner points (v=1 keypoints) that exceed the screen boundary during the inference phase, with no significant difference from the confidence of v=2 keypoints. Moreover, a condition for page integrity is that the confidence predicted by the keypoint detection model is greater than the confidence threshold. If the keypoint detection model still outputs high confidence values ​​for corner points (v=1 keypoints) that exceed the screen boundary during the inference phase, the logic of integrity judgment based on the confidence threshold will fail, judging v=1 keypoints as having a confidence greater than the confidence threshold, and misjudging incomplete pages as complete, resulting in misjudgment.

[0052] To overcome the above problems, one solution is to calculate the position loss and confidence loss only based on the key points with v=2. This solution completely abandons the position supervision signal of the key points with v=1. The key point detection model cannot learn the actual position of the key points with v=1, which affects the positioning accuracy of the key points with v=1.

[0053] Another solution is to reduce the weight of confidence loss. This approach does not distinguish between v=1 and v=2, which leads to the keypoint detection model predicting a low confidence level for keypoints with v=2, and may also result in misjudgments in integrity assessment.

[0054] The inventors of this application discovered that none of the above three schemes can simultaneously satisfy the two objectives of "preserving keypoint position learning at v=1" and "suppressing keypoint confidence at v=1," as these two objectives are coupled in existing loss designs. Based on this discovery, this application proposes a method to enable the keypoint confidence output by the keypoint detection model to accurately reflect the visibility state of the keypoints (v=1 → low confidence, v=2 → high confidence) without sacrificing keypoint position supervision at v=1, thereby supporting high-precision automatic target integrity detection.

[0055] Based on the above findings, a flowchart of an implementation method for training a key point detection model provided in this application is shown below. Figure 1 As shown, it may include:

[0056] Step S101: Obtain training samples.

[0057] The training samples include images and annotation information of at least one key point of the target object in the images. The annotation information of each key point includes the coordinates and visibility state of the key point. The visibility state is one of at least two states including a first state and a second state. The first state indicates that the key point exists but is not visible, and the second state indicates that the key point is fully visible.

[0058] The image can refer to an RGB image or a grayscale image of the target to be detected. The image can be captured by a general image acquisition unit (such as a camera) or a medical image captured by medical equipment.

[0059] The target to be detected may include, but is not limited to, any of the following: pages, human body, industrial parts, etc.

[0060] Key points can refer to feature points on the target to be detected that have specific semantic meaning, such as the four corner points of a book page (in the case of a single book page), the six corner points of a book page (in the case of two book pages), the joint points of the human body (usually 16 to 25), the key points of the hand (usually 21), or the center of the positioning hole of a part, etc.

[0061] The annotation information for each keypoint includes not only its position coordinates (x, y) in the image coordinate system, but also its visibility state, which characterizes its visibility. The visibility state is defined as a discrete variable that includes at least a first state and a second state. The first state indicates that the keypoint exists physically, but is not visible in the current image, possibly because the keypoint is outside the image's shooting boundary or is occluded by a foreground object (such as a finger or other book). The second state indicates that the keypoint exists and is fully visible in the image, is not occluded, and is located within the frame.

[0062] In some embodiments, the visibility state can also include a third state, indicating that the keypoint does not exist in the physical entity (e.g., the truncated target object does not have that corner point). Such keypoints are usually not involved in any loss calculation. For example, a single page only needs to detect 4 corner points. However, some keypoint detection schemes, in order to unify the use of the same model output head for single and double pages, fix the number of keypoint slots to 6. For a single page, only 4 slots are used to label the visibility state (first or second state) of the four corner points, and the other two slots are labeled as the third state. For the case of two pages (i.e., the book is open, and both pages are in the frame), 6 slots are used to label the visibility state (first or second state) of the six corner points. This avoids designing detection heads with different output dimensions for single and double pages, and the model structure is simpler.

[0063] For example, in a single page integrity detection scenario, if the captured image only contains the top left, top right, and bottom left corners of the page, while the bottom right corner is outside the camera's field of view, then the visibility state of the bottom right corner is marked as the first state (v=1), and its coordinates can be the theoretical projection position of the corner in the real world (even if the position is outside the image pixel range), while the visibility states of the other three corners within the image are marked as the second state (v=2).

[0064] In actual training, training is usually carried out in batches, that is, each batch of training sample images is input into the key point detection model to be trained, and the key point detection model processes each image in the same way.

[0065] Step S102: Input the image into the key point detection model to be trained to obtain the predicted coordinates and prediction confidence of the key points.

[0066] The keypoint detection model to be trained can refer to a neural network model that has not yet converged or is being iteratively optimized. Its architecture can be based on convolutional neural networks (CNNs) or Transformers, such as the YOLO-Pose series of models. The model is configured to receive an input image and output a prediction result (including predicted coordinates and prediction confidence) corresponding to each preset keypoint. The predicted coordinates are the location information of the keypoint in the image space estimated by the model, usually represented by normalized coordinates (coordinate values ​​range from [0,1]) or pixel coordinates; the prediction confidence is the probability estimate of whether the keypoint exists and whether it is visible, and its value range is usually [0,1].

[0067] Step S103: Calculate the position loss based on the key points whose visibility states are the first and second states.

[0068] Position loss is a metric used to measure the difference between the model's predicted coordinates and the labeled true coordinates. Its role can be to guide the model to optimize coordinate regression parameters and improve the accuracy of key point localization.

[0069] In this embodiment, the set of keypoints used to calculate the position loss covers all keypoints with visibility states of first state (existing but not visible) and second state (fully visible). This means that regardless of whether a keypoint is visible within the frame, as long as it physically exists (i.e., it is not non-existent), its coordinate information is included in the scope of position supervision.

[0070] The location loss can be calculated using methods such as Object Keypoint Similarity (OKS) loss, smoothed L1 loss, or mean squared error (MSE). Optionally, a location supervision mask can be pre-constructed, where different elements correspond to different keypoints. Elements for keypoints in both the first and second states are set to valid (e.g., a value of 1), while elements for non-existent keypoints are set to invalid (e.g., a value of 0). When calculating the location loss, the location supervision mask can be used to quickly determine which keypoints (keypoints with a mask of 1) participate in the location loss calculation and which keypoints (keypoints with a mask of 0) do not.

[0071] For key points involved in the location loss calculation, the location loss is calculated based on the predicted coordinates and the labeled true coordinates of the key point.

[0072] Step S104: Calculate the confidence loss based at least on the key points of the second state.

[0073] Confidence loss is a metric used to measure the difference between the model's predicted confidence and the expected supervision target. Its core objective is to decouple "preserving the location learning of keypoints with the first visibility state" from "suppressing the confidence of keypoints with the first visibility state," thereby correcting the problem of excessive supervision of the first-state keypoint confidence in existing technologies. This application provides two optional implementation strategies to achieve this goal:

[0074] Optionally, the confidence loss can be calculated based on keypoints whose visibility state is in the second state. That is, the confidence loss is calculated only for keypoints whose visibility state is in the second state (fully visible). Under this strategy, the system can construct a confidence mask, in which different elements correspond to different keypoints. Only when the visibility state of a keypoint is in the second state is the corresponding element set to valid (e.g., a value of 1), while the corresponding element for a keypoint whose visibility state is not in the second state is set to invalid (e.g., a value of 0). When calculating the confidence loss, the confidence mask can be used to quickly determine which keypoints (keypoints with a mask of 1) participate in the calculation of the confidence loss and which keypoints (keypoints with a mask of 0) do not participate in the calculation of the confidence loss.

[0075] For keypoints involved in the confidence loss calculation, the confidence loss is calculated based on the predicted coordinates of the keypoint and the pre-defined confidence supervision target corresponding to the visibility state of the keypoint. The confidence loss can be the binary cross-entropy (BCE) loss.

[0076] As an example, the confidence supervision target for keypoints in the second visibility state can be high (e.g., 1.0), indicating that the model should be confident that the point is visible. In this way, the model learns that only fully visible points have high confidence, while points that exist but are not visible (first state) will not be trained as high-confidence outputs.

[0077] Optionally, the confidence loss can be calculated based on keypoints with visibility in the first and second states, in such a way that the confidence supervision intensity applied to keypoints with visibility in the first state is less than the confidence supervision intensity applied to keypoints with visibility in the second state.

[0078] Confidence supervision strength refers to the degree of influence of key points in different visibility states on the loss function and model parameter updates when calculating confidence loss. The higher the confidence supervision strength of key points in any state, the stronger the driving effect of key points in that state on the model learning to output a specific confidence value; the lower the confidence supervision strength of key points in any state, the weaker the driving effect of key points in that state on the model learning to output a specific confidence value.

[0079] Supervision intensity can be achieved through one or more of the following methods:

[0080] The supervision strength is achieved by adjusting the size of the confidence level supervision target: the closer the confidence level supervision target is to 1, the greater the pressure on the model to output high confidence, and the higher the supervision strength; the closer the confidence level supervision target is to 0, the greater the pressure on the model to output low confidence, or the weaker the constraint on the model to output high confidence, and the lower the supervision strength.

[0081] This is achieved by adjusting the loss weights: after calculating the confidence loss for keypoints in different visibility states, different weight coefficients are assigned. The larger the weight coefficient, the greater the contribution of that part of the loss to the total loss, and the higher the supervision strength; the smaller the weight coefficient, the lower the supervision strength.

[0082] This strategy calculates confidence loss based on keypoints with visibility states of both state 1 and state 2, but applies different levels of supervision to each. Specifically, stronger confidence supervision (e.g., larger weights or a confidence supervision target of 1.0) is applied to keypoints in state 2, while weaker confidence supervision (e.g., smaller weights or a confidence supervision target between 0 and 1, but significantly lower than the confidence supervision target for state 2) is applied to keypoints in state 1. This weaker supervision signal corresponding to state 1 tells the model that although the point exists (and therefore participates in the position loss), its visibility is uncertain or low, and therefore a higher confidence score should not be output as it is for state 2. This strategy differentiates the confidence loss for keypoints with a visibility state of first state and keypoints with a visibility state of second state, breaking the traditional approach of treating keypoints with a visibility state of first state and keypoints with a visibility state of second state equally in confidence training. This allows the confidence score output by the model to truly reflect the visibility state of the keypoint: second state corresponds to high confidence score, and first state corresponds to low confidence score, thus providing a reliable basis for subsequent integrity judgment based on confidence score thresholds.

[0083] It should be noted that this application does not limit the execution order of steps S103 and S104. Step S103 can be executed first and then step S104, or step S104 can be executed first and then step S103, or both steps can be executed simultaneously.

[0084] The intensity of supervision can also be determined by whether or not a keypoint participates in the loss calculation: keypoints that participate in the confidence loss calculation have a certain level of supervision intensity; keypoints that do not participate in the confidence loss calculation have zero supervision intensity. The aforementioned calculation of confidence loss based only on keypoints with a visibility state of second state (fully visible) is equivalent to applying a confidence supervision intensity to keypoints with a visibility state of first state, which is less than the confidence supervision intensity applied to keypoints with a visibility state of second state.

[0085] Step S105: Update the parameters of the keypoint detection model based on the location loss and confidence loss.

[0086] The location loss and confidence loss can be weighted and summed to obtain the total loss. The parameters of the keypoint detection model are then updated with the goal of minimizing the total loss. The weights of the location loss and the confidence loss can be the same or different.

[0087] The backpropagation algorithm can be used to adjust the parameters of the keypoint detection model (such as weight parameters, bias parameters, etc.) based on the total loss. For specific update methods, please refer to existing solutions, which will not be elaborated here.

[0088] The keypoint detection model training method provided in this application, after acquiring training samples with labeled keypoint coordinates and visibility states (a first state indicating the keypoint exists but is not visible, or a second state indicating the keypoint is fully visible), incorporates both first-state and second-state keypoints into the position loss calculation during model training. This preserves the model's geometric learning ability for occluded keypoints, ensuring the accuracy of position prediction. Simultaneously, by calculating the confidence loss only for keypoints in the second visibility state or applying weak confidence supervision to keypoints in the first visibility state, the erroneous association of "existence equals high confidence" is successfully severed. This decoupling mechanism enables the trained model to output highly discriminative confidence scores (fully visible keypoints receive high confidence, while existing but not visible keypoints receive low confidence). Based on this, the model parameters are updated by combining the position loss and confidence loss, allowing the model to maintain high-precision localization while possessing accurate visibility perception capabilities. This improves the accuracy of target (e.g., page) integrity detection.

[0089] Furthermore, this application does not require modification of the underlying training framework source code or re-standardization of the dataset, making implementation simple.

[0090] In an optional embodiment, the flowchart for calculating the confidence loss based on key points with visibility states of first and second states is as follows: Figure 2 As shown, it may include:

[0091] Step S201: Calculate the first loss based on the prediction confidence of key points with visibility status of the first state and the first confidence supervision target.

[0092] The first confidence supervision objective is a pre-defined expected confidence value corresponding to the first state, used to guide the model's learning.

[0093] Step S202: Calculate the second loss based on the prediction confidence of key points whose visibility state is the second state and the second confidence supervision target; the first confidence supervision target is less than the second confidence supervision target.

[0094] The second confidence supervision objective is a pre-defined expected confidence value for the corresponding second state, used to guide model learning.

[0095] Optionally, if the absolute value of the difference between the first confidence level supervision target and the second confidence level supervision target is greater than a preset difference threshold, it indicates that the first confidence level supervision target is significantly lower than the second confidence level supervision target.

[0096] As an example, the first confidence level supervision target is 0.0 or a very small value close to 0.0 (e.g., 0.1 or 0.05), which aims to convey to the model the supervision signal that although the key point physically exists, it should not be regarded as a valid feature with high confidence in the current field of vision.

[0097] The second confidence level supervision target can be 1.0 or a high value close to 1.0 (e.g., 0.9 or 0.8).

[0098] By providing low-target-value supervision for keypoints in the first visibility state, the model can effectively suppress inflated confidence scores for invisible keypoints. The differentiated setting of the first and second confidence supervision targets allows the model to clearly distinguish the different semantics of the existence-but-not-visible and fully-visible states in the confidence dimension, breaking the coupled logic of treating the two as equals in existing technologies.

[0099] It should be noted that this application does not limit the execution order of steps S201 and S202. Step S201 can be executed first and then step S202, or step S202 can be executed first and then step S201, or both steps can be executed simultaneously.

[0100] Step S203: Weight the first loss and the second loss to obtain the confidence loss.

[0101] The weight of the first loss is less than or equal to the weight of the second loss. In other words, when updating model parameters through backpropagation, the contribution of the confidence error from keypoints with a visibility state of first state (existing but not visible) to the model parameter update is limited to a low level, while the confidence error from keypoints with a visibility state of second state (fully visible) dominates or at least has an equal contribution.

[0102] pass Figure 2 In the illustrated embodiment, this application implements hierarchical confidence supervision of key points in different visibility states. Specifically, based on the fact that the first confidence supervision target is smaller than the second confidence supervision target, the learning target of invisibility (i.e., low confidence) is clearly defined from the numerical level of the supervision signal. Simultaneously, the weight of the first loss is smaller than the weight of the second loss, limiting the interference of invisible key points on model parameter updates from the perspective of the optimization process's intensity. The synergistic effect of these two mechanisms allows the key point detection model to retain the key point position prediction capability in the first state (guaranteed by the position loss) while accurately suppressing its prediction confidence to a low level, and accurately assigning high confidence to key points in the second state. This decoupling mechanism effectively solves the problem of missed detection of page integrity caused by artificially high confidence of key points with v=1 in existing technologies, significantly improving the model's accuracy in judging target integrity without modifying the original data annotation.

[0103] In an optional embodiment, another implementation flowchart for calculating the confidence loss based on key points with visibility states of first and second states is shown below. Figure 3 As shown, it may include:

[0104] Step S301: Calculate the first loss based on the prediction confidence of the key point with the first confidence state and the first confidence supervision target.

[0105] The first confidence supervision objective is a pre-defined expected confidence value corresponding to the first state, used to guide the model's learning.

[0106] Step S302: Calculate the second loss based on the prediction confidence of key points with visibility status in the second state and the second confidence supervision target; the first confidence supervision target is equal to the second confidence supervision target.

[0107] The second confidence supervision objective is a pre-defined expected confidence value for the corresponding second state, used to guide model learning.

[0108] The first and second confidence level supervision targets can both be 1.0 or high values ​​close to 1.0 (e.g., 0.9 or 0.8).

[0109] It should be noted that this application does not limit the execution order of steps S301 and S302. Step S301 can be executed first and then step S302, or step S302 can be executed first and then step S301, or both steps can be executed simultaneously.

[0110] Step S303: The first loss and the second loss are weighted and summed to obtain the confidence loss; the weight of the first loss is less than the weight of the second loss.

[0111] In this embodiment, the weights assigned to the first loss are strictly less than the weights assigned to the second loss. Through this weight difference configuration, although the keypoints for visibility states one and two use the same confidence supervision target (e.g., both 1.0), the gradient contribution from the first-state keypoints is significantly suppressed during backpropagation to update the model parameters, while the gradient contribution from the second-state keypoints dominates. For example, the weight of the first loss is 0.2, while the weight of the second loss is 1.0.

[0112] pass Figure 3 The embodiment shown in this application provides an alternative to decoupled confidence training without changing the labeling. By setting the first confidence supervision target equal to the second confidence supervision target, the consistency of data labeling is maintained, and the data preprocessing process is simplified. Simultaneously, by employing a weighting strategy where the first loss weight is less than the second loss weight, a difference in confidence supervision intensity is introduced during the loss calculation stage. The low weight of the first loss limits the dominant role of invisible keypoints in the model gradient, preventing the model from overfitting to existing but invisible sample features and avoiding artificially high confidence scores during inference. The high weight of the second loss ensures that the confidence scores of fully visible keypoints converge sufficiently to high values. The synergistic effect of these two approaches preserves the opportunity for first-state keypoints to participate in confidence learning (beneficial for the potential association learning of location information) while effectively reducing their misleading risk in the confidence dimension. Ultimately, this enables the trained keypoint detection model to output a highly discriminative confidence score, providing a reliable basis for subsequent target integrity detection based on confidence thresholds.

[0113] In an optional embodiment, the keypoint detection model training method provided in this application may further include:

[0114] Before acquiring training samples, at a preset callback time point, a pre-registered callback function is invoked to configure the function used to calculate the confidence loss as the target loss function;

[0115] The target loss function is used to calculate the confidence loss based on keypoints with a visibility state of the second state; or, the confidence loss is calculated based on keypoints with visibility states of the first and second states in such a way that the confidence supervision intensity applied to keypoints with visibility states of the first state is less than the confidence supervision intensity applied to keypoints with visibility states of the second state.

[0116] The preset callback time point can refer to a specific initialization phase in the lifecycle of a deep learning training framework, specifically, a moment before the start of the training loop, after the optimizer is built, or before the first batch of data is loaded. Optionally, the preset callback time point can be the end of the pre-training routine, i.e., a moment before the first batch of data is loaded.

[0117] Pre-registered callback functions can refer to custom function objects bound to the aforementioned time points via the framework's registration interface (e.g., add_callback) when the training script starts. The purpose of this callback function is to dynamically intercept and modify the internal logic of the loss calculation module before the actual training runs.

[0118] The target loss function is a configured confidence loss calculation unit that is configured to execute one of the following two strategies:

[0119] Strategy 1: Calculate the confidence loss only for keypoints with a visibility state of second state (i.e., the keypoint is fully visible, v=2). Keypoints with a visibility state of first state (i.e., the keypoint exists but is not visible, v=1) are not included in the calculation of confidence loss.

[0120] Strategy 2: Calculate the confidence loss based on keypoints with visibility states of first and second states. However, the confidence supervision intensity applied to keypoints with visibility state of first state is significantly less than the supervision intensity applied to keypoints with visibility state of second state. For example, the confidence supervision target for keypoints with v=2 can be set to 1.0, while the confidence supervision target for keypoints with v=1 can be set to 0.0 or 0.2, thereby decoupling position supervision preservation and confidence suppression at the loss function level.

[0121] By calling the callback function at the callback time point, the system can replace the default confidence loss calculation logic of the training framework with the target loss function proposed in this application without modifying the source code of the training framework.

[0122] This application achieves non-intrusive deployment of training strategies by dynamically configuring the target loss function. The execution entity (training framework) automatically loads the preset loss calculation rules at pre-defined callback time points, without requiring manual intervention in the underlying codebase, ensuring the integrity of the native training process while flexibly injecting differentiated supervision strategies. The resulting configuration directly affects subsequent training iterations, ensuring that the model parameter update direction conforms to the decoupled training requirements of this application. This fundamentally solves the contradiction between positional accuracy and confidence discrimination ability in existing technical solutions, significantly improving the accuracy of target integrity detection based on keypoint confidence, and without requiring re-labeling of the dataset or modification of the source code of mainstream deep learning frameworks, making it practical for engineering applications and easy to migrate.

[0123] In an optional embodiment, the keypoint detection model training method provided in this application may further include:

[0124] Before inputting the image into the keypoint detection model to be trained, the system randomly determines whether to perform online occlusion enhancement processing on the image with a preset probability.

[0125] The preset probability is a pre-defined hyperparameter used to control the frequency of occlusion enhancement during training. Its value typically ranges from 0 to 1; for example, it can be set to 0.5, meaning that approximately 50% of the training samples are subjected to occlusion enhancement in each iteration. This step aims to increase the diversity of training data, thereby improving the robustness of the model.

[0126] As an example, the system can generate a random number that follows a uniform distribution. If the random number is less than or equal to a preset probability, the system determines whether to perform online occlusion enhancement on the image; otherwise, it determines not to perform online occlusion enhancement. This stochastic decision-making mechanism ensures the diversity of training data distribution and avoids the model overfitting to perfectly visible ideal samples.

[0127] If it is determined that no online occlusion enhancement processing will be performed on the image, the image will be input into the keypoint detection model to be trained.

[0128] When the random decision result is not to perform online occlusion enhancement on the image, the current training samples remain in their original state, meaning the image content and the annotations of each keypoint (including coordinate annotations and visibility state annotations) remain unchanged. This step, as a branch logic case, directly feeds the original image into the keypoint detection model to be trained for forward propagation calculation, calculating the position loss and confidence loss based on the original annotation information of the original image. At this point, the model learns features based on the original visibility state distribution. Keypoints in the first visibility state (existing but not visible) continue to participate in the position loss calculation but do not participate in or weakly participate in the positive confidence sample supervision, while keypoints in the second visibility state (fully visible) participate normally in the calculation of all loss terms. This process ensures that the model can continuously learn basic features from the real-world data distribution.

[0129] If it is determined that online occlusion enhancement processing is to be performed on the image, at least one key point is selected from the key points whose visibility state is in the second state as the target key point for occlusion, and the visibility state of the target key point is updated to the first state, while the visibility states of other key points remain unchanged, thus obtaining the enhanced training sample; the image of the enhanced training sample is input into the key point detection model to be trained.

[0130] When the random decision result is to perform online occlusion enhancement processing on the image, the process of online occlusion enhancement processing on the image is triggered. This process may include: randomly selecting at least one keypoint as the target keypoint from all keypoints in the image whose visibility state is in the second state; and then applying visual occlusion at the location of the target keypoint in the image to simulate the occlusion effect in the physical world. The occlusion form can be any of the following: covering a color block of a specific shape, pasting a transparent PNG occluder, or applying local Gaussian blur, etc. A transparent PNG occluder refers to a PNG format image with an alpha transparency channel. This PNG format image is the same size as the image that needs online occlusion enhancement, and only the area used to cover the target keypoint is opaque, while other areas are transparent. In this way, when this PNG format image is overlaid on the image that needs online occlusion enhancement, the texture of the image that needs online occlusion enhancement can be preserved, making it closer to real occlusion.

[0131] While applying visual occlusion, it is also necessary to update the visibility state of the target key point to the first state, while keeping the visibility state of other key points (non-target key points) unchanged, and keeping the coordinates of all key points (including target key points and non-target key points) unchanged.

[0132] For example, before online occlusion enhancement, the annotation information (x, y, v) of the four key points in the image used for page integrity detection are: top left corner (x1, y1, 2), top right corner (x2, y2, 2), bottom left corner (x3, y3, 1), and bottom right corner (x4, y4, 2). Suppose that the top left corner is occluded through online occlusion enhancement, then the visibility state of the top left corner is modified to v=1, while the coordinates remain unchanged. The coordinates and visibility states of the other corner points remain unchanged. In other words, the annotation information of the enhanced training sample image is: top left corner (x1, y1, 1), top right corner (x2, y2, 2), bottom left corner (x3, y3, 1), and bottom right corner (x4, y4, 2).

[0133] By modifying images and annotations simultaneously, the generated enhanced training samples visually exhibit occlusion features and conform to the definition of the first state in terms of annotation, thus forming a closed loop with the aforementioned position / confidence decoupled training mechanism of this invention: these newly generated visibility states are key points of the first state and will automatically apply the loss calculation rules designed for the first state (i.e., participate in position regression but receive low confidence supervision), without the need for additional manual re-annotation costs.

[0134] After inputting the enhanced training sample images into the keypoint detection model to be trained, and obtaining the predicted coordinates and prediction confidence of the keypoint detection model, the location loss and confidence loss are calculated based on the annotation information of the images in the enhanced training samples.

[0135] By performing online occlusion enhancement on the training samples, on the one hand (data-level optimization), a large number of samples conforming to the first state definition are actively constructed, enriching the diversity of the training data; on the other hand (algorithm-level optimization), combined with the low-confidence supervision strategy applied to the key points of the first state in this invention, these online enhanced training samples can effectively guide the model to reduce the confidence score of occluded key points, avoiding the false alarm problem of integrity detection caused by the inflated confidence of v=1 samples in traditional solutions. The combined use of these two methods not only preserves the accuracy of key point location prediction but also strengthens the correspondence between confidence scores and visibility states. This allows the finally trained model to accurately locate and evaluate visibility in complex occlusion scenarios, thereby improving the reliability and generalization ability of downstream tasks such as page integrity detection.

[0136] After the keypoint detection model is trained, target integrity detection can be performed based on the trained keypoint detection model. For example... Figure 4 The diagram shown is a flowchart of one implementation of the target integrity detection method provided in this application, which may include:

[0137] Step S401: Use the key point detection model to detect key points of the target object in the image to be detected, and obtain the predicted coordinates and prediction confidence of each key point of the target object; the key point detection model is trained by the key point detection model training method described above.

[0138] The image to be detected can be any of the following: a page image, a hand image, a human body image, a medical image, an industrial part image, etc.

[0139] If the image to be detected is a page image, then the images in the training samples used to train the keypoint detection model are page images.

[0140] Similarly, if the image to be detected is a hand image, then the images in the training samples used to train the keypoint detection model are hand images.

[0141] If the image to be detected is a human image, then the images in the training samples used to train the keypoint detection model are human images.

[0142] If the image to be detected is a medical image, then the images in the training samples used to train the keypoint detection model are medical images.

[0143] If the image to be detected is an image of an industrial part, then the images in the training samples used to train the keypoint detection model are images of industrial parts.

[0144] Step S402: Determine whether the target object in the image to be detected is complete based on the predicted coordinates and prediction confidence of each key point.

[0145] Optionally, if the prediction confidence of each key point is greater than or equal to the confidence threshold, and the predicted coordinates of each key point are all within the image boundary, then the target object in the image to be detected is determined to be complete; otherwise, the target object in the image to be detected is determined to be incomplete.

[0146] The keypoint detection model trained using the aforementioned training method can clearly distinguish the different semantics of the two states of existence but not visible and complete visibility in terms of confidence dimension. Therefore, based on the coordinate threshold (image boundary) and confidence threshold, it can accurately determine whether the target object in the image to be detected is complete, thereby improving the accuracy of target integrity judgment and reducing the misjudgment rate of target integrity judgment based on keypoint detection.

[0147] Optionally, the confidence threshold is 0.25, but it can also be set to other values ​​as needed. This application does not impose any specific limitations on it.

[0148] Optionally, if the prediction confidence of each keypoint is greater than or equal to the confidence threshold, and the predicted coordinates of each keypoint are all within the safety boundary, then the target object in the image to be detected is determined to be complete; otherwise, the target object in the image to be detected is determined to be incomplete. The safety boundary is the boundary formed by indenting the physical boundary of the image to be detected inward by a preset proportion.

[0149] To further reduce the misjudgment rate of target integrity assessment based on key point detection, the predicted coordinates of each key point need to be within the safety boundary.

[0150] As an example, suppose the width of the image to be detected is w, the height is h, the coordinates of the top left vertex are [0,0], the coordinates of the bottom right vertex are [w-1,h-1], and the preset scale is r; then, the coordinates of the top left vertex of the safety boundary are [w×r,h×r], and the coordinates of the bottom right vertex are [w-1-w×r,h-1-w×r]. Obviously, the width of the safety boundary is w-2×w×r+1, and the height is h-2×h×r+1.

[0151] For example, if the width of the image to be detected is 1000 pixels and the height is 500 pixels, and the preset ratio is 5%, then the coordinates of the top left vertex of the safety boundary are [50, 25] and the coordinates of the bottom right vertex are [949, 474]; the width of the safety boundary is 900 and the height is 450.

[0152] Optionally, the preset ratio can be 5%, or it can be set to other values ​​as needed. This application does not impose specific limitations.

[0153] By setting safety boundaries, measurement errors caused by key point coordinates falling exactly on the edge pixels of the image can be avoided, ensuring that the key points are indeed located within the effective area of ​​the image.

[0154] Optionally, if the prediction confidence of each key point is greater than or equal to the confidence threshold, the predicted coordinates of each key point are all within the safety boundary, and the ratio of the area of ​​the polygon enclosed by each key point to the total area of ​​the image to be detected is greater than the ratio threshold, then the target object in the image to be detected is determined to be complete; otherwise, the target object in the image to be detected is determined to be incomplete.

[0155] To further reduce the false positive rate of target integrity judgment based on keypoint detection, the ratio of the area of ​​the polygon enclosed by each keypoint to the total area of ​​the image to be detected must be greater than a proportional threshold. The proportional threshold measures the minimum proportion of the keypoint distribution range. Optionally, the proportional threshold can be 0.5%, or it can be set to other values ​​as needed; this application does not impose specific limitations.

[0156] This application ensures that a target object is considered complete only when it meets the requirements in three dimensions: semantic credibility, spatial location rationality, and structural coverage integrity. This significantly improves the robustness of target integrity detection and effectively avoids misjudging incomplete targets with key points outside the frame or severely occluded as complete.

[0157] The target integrity detection method of this application, based on the keypoint detection model obtained by the aforementioned model training method, outputs a keypoint prediction confidence level that can accurately reflect the visibility status of keypoints, thus giving the aforementioned confidence threshold judgment reliable physical meaning. Furthermore, by combining the spatial constraints of the safety boundary and the proportional constraints of the polygon area, the system not only focuses on the existence of keypoints (confidence level), but also on the reasonableness of their location (safety boundary) and the sufficiency of their distribution (area ratio). This multi-dimensional fusion judgment mechanism effectively solves the false alarm problem caused by the artificially high confidence level of keypoints with v=1 in existing technologies, providing high-quality data input assurance for subsequent tasks.

[0158] In an optional embodiment, the target integrity detection method provided in this application may further include:

[0159] If the predicted coordinates and prediction confidence of each key point of at least two target objects are obtained, calculate the area of ​​the detection box for each target object.

[0160] When a keypoint detection model infers on a single image and detects two or more target objects, the system obtains the predicted coordinates of all keypoints corresponding to each target object. Based on these predicted coordinates, a bounding box, or detection box, is constructed for each target object. The area of ​​the detection box is calculated based on the product of the width and height of the rectangle, and its value reflects the size of the pixel range occupied by the target object in the image. Specifically, for the i-th target object, if its keypoint predicted coordinate set is {(x... i ,y i If the coordinates of the top left corner of the detection box are (x, y), then the coordinates of the top left corner of the detection box are (x, y). min ,y min ) and the coordinates of the lower right corner (x max ,y max The area S of the detection box for the i-th target object is determined by the minimum and maximum values ​​of the coordinates of all key points, respectively. i =(x max -x min )×(y max -y min By traversing all detected target objects, the system sequentially calculates the area of ​​each detection box, forming an area set {S1, S2, ... S}.n}, where n is greater than or equal to 2, representing the number of target objects identified. This process provides a quantitative basis for subsequent object selection.

[0161] The target object corresponding to the largest detection box is identified as the object that needs to be inspected for integrity.

[0162] This application compares the detection box area values ​​of all target objects, identifies the maximum value, and locks the target object corresponding to this maximum value as the only object in the current image that needs to be checked for integrity. Other target objects with smaller areas are ignored and not included in the subsequent integrity judgment process (i.e., integrity checks based on confidence thresholds, safety boundaries, and polygon areas are no longer performed). The principle behind this approach is that in practical applications (such as shooting book pages in children's English education), users typically focus on the main object located in the center of the frame, closest to the lens, and most completely displayed. This main object usually has the largest detection box area in the image. For example, when a camera captures two books, one large and one small, laid out on a table, the larger book is often closer to the lens or more completely in the center of the field of view. Using it as the detection subject aligns with the user's intention, effectively filtering out distracting secondary targets in the background, ensuring the system's output integrity conclusion is unique and accurate, and avoiding contradictory judgments caused by detecting multiple targets separately.

[0163] The following section uses the page integrity detection task as an example to compare the effectiveness of the target integrity detection method of this application for page integrity detection with the effectiveness of existing technologies for page integrity detection.

[0164] The main process of this application for page integrity detection is the same as that of existing technologies, which predicts the corner coordinates and confidence scores of the page based on a keypoint detection model, and then determines whether the page is intact based on a threshold. The difference lies in the training method of the keypoint detection model in this application. This application considers area and safety boundaries when determining the page integrity based on the threshold. Table 1 shows a comparison of various indicators between the page integrity detection method based on this application and the page integrity detection methods based on existing technologies.

[0165] Table 1

[0166]

[0167] Wherein, TP represents a true positive instance, which is both a positive and a predicted positive instance; FP represents a false positive instance, which is both a negative and a predicted positive instance; TN represents a true negative instance, which is both a negative and a predicted negative instance; and FN represents a false negative instance, which is both a positive and a predicted negative instance.

[0168] This application reduced FP by 69%, FN by 99%, improved precision by 5.21 percentage points, improved recall by 6.35 percentage points, improved F1 score by 5.77 percentage points, and improved accuracy by 4.27 percentage points.

[0169] Corresponding to the method embodiments, this application also provides an electronic device. (See reference...) Figure 5 As shown, it illustrates a structural schematic diagram of an electronic device suitable for implementing embodiments of this application. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0170] like Figure 5 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. When the electronic device is powered on, the RAM 503 also stores various programs and data required for the operation of the electronic device. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0171] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, memory cards, hard drives, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0172] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the key point detection model training methods and / or target integrity detection methods provided in this application.

[0173] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the key point detection model training methods and / or target integrity detection methods provided in this application.

[0174] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] In the above embodiments, the functionality can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented entirely or partially as a computer program product. Those skilled in the art can use different methods to implement the described functions for each specific solution, but such implementation should not be considered beyond the scope of this application.

[0177] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0178] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0179] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A keypoint detection model training method, characterized in that, include: Acquire training samples, which include an image and annotation information of at least one key point of a target object in the image. The annotation information of each key point includes the coordinates and visibility state of the key point. The visibility state is one of at least two states including a first state and a second state. The first state indicates that the key point exists but is not visible, and the second state indicates that the key point is fully visible. The image is input into the key point detection model to be trained to obtain the predicted coordinates and prediction confidence of the key points; Based on the key points whose visibility states are the first state and the second state, calculate the position loss; Calculate the confidence loss based on keypoints whose visibility state is the second state; or, calculate the confidence loss based on keypoints whose visibility state is the first state and the second state, in such a way that the confidence supervision intensity applied to keypoints whose visibility state is the first state is less than the confidence supervision intensity applied to keypoints whose visibility state is the second state. The parameters of the keypoint detection model are updated based on the location loss and the confidence loss.

2. The method according to claim 1, characterized in that, Calculate the confidence loss based on key points with visibility states of first and second states, including: The first loss is calculated based on the prediction confidence of key points with visibility in the first state and the first confidence supervision target; The second loss is calculated based on the prediction confidence of key points whose visibility state is the second state and the second confidence supervision target; the first confidence supervision target is less than the second confidence supervision target; The confidence loss is obtained by weighted summation of the first loss and the second loss; the weight of the first loss is less than or equal to the weight of the second loss.

3. The method according to claim 1, characterized in that, Calculate the confidence loss based on key points with visibility states of first and second states, including: The first loss is calculated based on the prediction confidence of key points with visibility in the first state and the first confidence supervision target; The second loss is calculated based on the prediction confidence of key points whose visibility state is the second state and the second confidence supervision target; the first confidence supervision target is equal to the second confidence supervision target; The confidence loss is obtained by weighted summation of the first loss and the second loss; the weight of the first loss is less than the weight of the second loss.

4. The method according to any one of claims 1-3, characterized in that, Before obtaining training samples, the following is also included: At the preset callback time point, the pre-registered callback function is invoked to configure the function used to calculate the confidence loss as the target loss function; The target loss function is used to calculate the confidence loss based on keypoints with a visibility state of the second state; or, the confidence loss is calculated based on keypoints with visibility states of the first and second states in such a way that the confidence supervision intensity applied to keypoints with visibility states of the first state is less than the confidence supervision intensity applied to keypoints with visibility states of the second state.

5. The method according to any one of claims 1-3, characterized in that, Also includes: Before inputting the image into the key point detection model to be trained, it is randomly determined with a preset probability whether to perform online occlusion enhancement processing on the image; If it is determined that the image will not be subjected to online occlusion enhancement processing, the image will be input into the key point detection model to be trained; If it is determined that the image will be subjected to online occlusion enhancement processing, at least one key point is selected from the key points whose visibility state is the second state as the target key point for occlusion, and the visibility state of the target key point is updated to the first state, while the visibility states of other key points remain unchanged, thus obtaining the enhanced training sample. The images of the enhanced training samples are input into the keypoint detection model to be trained.

6. A target integrity detection method, characterized in that, include: The key point detection model is used to detect key points of the target object in the image to be detected, and the predicted coordinates and prediction confidence of each key point of the target object are obtained. The keypoint detection model is trained using the keypoint detection model training method as described in any one of claims 1-5; Based on the predicted coordinates and prediction confidence of each key point, it is determined whether the target object in the image to be detected is complete.

7. The method according to claim 6, characterized in that, Determining whether the target object in the image to be detected is complete based on the predicted coordinates and prediction confidence of each key point includes: If the prediction confidence of each key point is greater than or equal to the confidence threshold, the predicted coordinates of each key point are all within the safe boundary, and the ratio of the area of ​​the polygon enclosed by each key point to the total area of ​​the image to be detected is greater than the ratio threshold, then the target object in the image to be detected is determined to be complete; otherwise, the target object in the image to be detected is determined to be incomplete. The security boundary is the boundary formed by indenting the physical boundary of the image to be detected by a preset proportion.

8. The method according to claim 6, characterized in that, Also includes: If the predicted coordinates and prediction confidence of each key point of at least two target objects are obtained, calculate the area of ​​the detection box for each target object; The target object corresponding to the largest detection box is identified as the object that needs to be inspected for integrity.

9. An electronic device, characterized in that, The electronic device includes at least one processor and a memory connected to the processor; wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to enable the electronic device to implement the key point detection model training method as described in any one of claims 1 to 5, and / or to implement the target integrity detection method as described in any one of claims 6 to 8.

10. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the keypoint detection model training method as described in any one of claims 1 to 5, and / or to implement the target integrity detection method as described in any one of claims 6 to 8.

11. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the key point detection model training method as described in any one of claims 1 to 5, and / or, the target integrity detection method as described in any one of claims 6 to 8.