A method, system, electronic device, and storage medium for annotating human body features.
By establishing a communication connection between the laser annotation device and the terminal device, and using a human image recognition model, the adaptability and user-friendliness issues of human feature annotation in existing technologies have been resolved. This has enabled simple and easy-to-use automatic human feature annotation, improving the recognition capabilities of non-professional users and the real-time interactive capabilities of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUBEI GANGDING TECHNOLOGY CO LTD
- Filing Date
- 2025-05-24
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for automatic human feature annotation suffer from insufficient adaptability, difficult-to-understand technical terms, high data acquisition costs, and difficulty in real-time system interaction, leading to confusion for non-professional users and making it impossible to effectively establish spatial cognition of human features.
By establishing a communication connection between the laser annotation device and the terminal device, the laser projection area is calibrated using a multi-view geometric method. Combined with a human image recognition model and a similarity matching algorithm, human features are automatically annotated. User information is obtained through voice or text, the camera is controlled to capture images, and the annotation results are displayed through laser projection.
It achieves simple and easy-to-use automatic human feature annotation, enabling non-professional users to identify human body parts and skeletons, reducing data acquisition costs, improving the system's real-time interactive capabilities, and enhancing the user experience.
Smart Images

Figure CN120600248B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human feature recognition technology, and more specifically, to a method, system, electronic device, and storage medium for labeling human features. Background Technology
[0002] With the aging population and increased public awareness of health management, the demand for home health monitoring has exploded. Traditional medical testing relies heavily on specialized equipment and technicians. While existing home medical devices can measure basic indicators such as blood pressure and blood oxygen, there are still significant technological gaps in terms of vital sign recognition, operational complexity, and result readability.
[0003] Current technologies for automatic human feature annotation still suffer from the following prominent problems: First, most existing solutions are computer vision-based annotation systems that are insufficiently adaptable to complex human postures and individual differences. Especially when dealing with non-standardized body shapes, special movements, or occluded scenes, the accuracy of skeletal and muscle annotation drops significantly, easily leading to keypoint drift or misjudgment of anatomical structures. Second, existing solutions generally use professional medical terminology and lack semantic conversion mechanisms for the general public. Ordinary users struggle to understand professional medical terms such as "coracoid process" and "anterior superior iliac spine," creating a cognitive gap. Third, traditional methods rely on large amounts of precisely annotated medical image data for model training, but acquiring high-quality medical data is costly and involves privacy and ethical issues, limiting the algorithm's generalization ability. Furthermore, existing systems mostly employ centralized computing architectures, making real-time interactive annotation on mobile devices difficult, and lacking multimodal guidance mechanisms (such as AR enhancement and 3D dynamic demonstrations). Non-professional users are prone to operational confusion and cannot effectively establish spatial cognition of human features.
[0004] Therefore, there is an urgent need for a new method for annotating human body features, which can automatically annotate human body features and make it easier for people without medical knowledge to perform simple human body feature recognition. Summary of the Invention
[0005] This invention addresses the technical problems existing in the prior art by providing a method, system, electronic device, and storage medium for annotating human body features.
[0006] According to a first aspect of the present invention, a method for annotating human body features is provided, comprising:
[0007] Establish a communication connection with the laser marking device, and calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0008] Obtain the human feature information to be labeled sent by the user, and collect corresponding human images based on the human feature information to be labeled;
[0009] Based on the trained human image recognition model, the coordinate positions of the human features to be labeled corresponding to the human image are identified.
[0010] The human features to be annotated are annotated to obtain the annotation results, and the annotation results and the coordinate positions are sent to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0011] Furthermore, calibrating the laser projection area of the laser marking device within a preset spatial coordinate system includes:
[0012] A checkerboard calibration board is printed and placed on a flat scene. Multiple images of the checkerboard from different perspectives are captured by the laser marking device, and each image contains laser highlights.
[0013] The pose of the laser marking device relative to the checkerboard coordinate system is calculated based on the multi-view geometric method.
[0014] Based on the geometric intersection of the laser spot and the calibration pose in the chessboard, the reprojection error is established as the optimization objective function.
[0015] Given the optimization objective function, the relative positional relationship between the laser marking device and the target device is calculated.
[0016] Furthermore, the calculation of the pose of the laser marking device relative to the checkerboard coordinate system based on the multi-view geometry method includes:
[0017] The pose is represented as
[0018] The laser is oriented fixed forward, let d be the direction. B =[0,0,1] T The starting position of the laser beam when image i is captured. and towards d (i) Represented as
[0019]
[0020] The checkerboard plane is located at Z=0, and passes through the normal vector n=[0,0,1]. T Let x0 be the center point of the checkerboard grid, then the intersection point of the laser beam and the plane is... for
[0021]
[0022] The intersection point is projected back onto the image, i.e., its position in the camera coordinate system is...
[0023]
[0024] Where π(·) is the normalized projection function divided by the third dimension, and K is the camera intrinsic parameter matrix. To predict the coordinates of the laser point in the image.
[0025] Furthermore, the step of acquiring the human feature information to be labeled sent by the user, and collecting corresponding human images based on the human feature information to be labeled, includes:
[0026] Acquire voice or text information sent by the user, and identify the human feature information to be labeled from the voice or text information;
[0027] Based on the mapping relationship between the human feature information to be labeled and the preset human body region, the camera is controlled to rotate to the corresponding human body region to capture human body images.
[0028] Furthermore, the step of identifying the coordinate positions of the human features to be labeled corresponding to the human image captured based on the trained human image recognition model includes:
[0029] Based on the trained human image recognition model, the human image is recognized to obtain feature recognition results;
[0030] Based on a similarity matching algorithm, the feature recognition result is matched with the human feature information to be labeled to determine whether they match.
[0031] If a match fails, the human body image will be re-acquired.
[0032] Furthermore, the step of matching the feature recognition result with the human feature information to be labeled based on the similarity matching algorithm includes:
[0033]
[0034] Where I is the human body image captured by the camera, F(I) is the feature recognition result, Q is the human body feature information to be labeled input by the user, τ is the similarity threshold, and S(·,·) is the similarity function.
[0035] Furthermore, the step of annotating the human features to be annotated to obtain the annotation result includes:
[0036] The human features to be labeled are explained in text and defined by a bounding box, wherein the bounding box defines the human feature recognition area.
[0037] According to a second aspect of the present invention, a system for annotating human body features is provided, comprising:
[0038] An initialization module is used to establish a communication connection with the laser marking device and to calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0039] The image acquisition module is used to acquire human feature information to be labeled sent by the user, and to acquire corresponding human images based on the human feature information to be labeled.
[0040] The feature recognition module is used to identify the coordinate position of the human features to be labeled in the human image based on the trained human image recognition model.
[0041] An automatic annotation module is used to annotate the human features to be annotated to obtain annotation results, and send the annotation results and the coordinate positions to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0042] According to a third aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the processor is configured to implement the steps of the above-described method for annotating human features when executing a computer management program stored in the memory.
[0043] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer management program is stored, wherein the computer management program, when executed by a processor, implements the steps of the above-described method for annotating human features.
[0044] This invention provides a method, system, electronic device, and storage medium for annotating human body features. It can automatically annotate human body features, including body parts, bones, muscles, etc., and the annotation process is very simple. Users only need to send information, which makes it easy for people without medical knowledge to have a simple identification of human body features. Attached Figure Description
[0045] Figure 1 This is a flowchart of a method for annotating human body features according to an embodiment of the present invention;
[0046] Figure 2 A system structure diagram for annotating human body features provided in an embodiment of the present invention;
[0047] Figure 3 A schematic diagram of an embodiment of the electronic device provided in this invention. Detailed Implementation
[0048] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0049] Human body feature recognition involves complex anatomical knowledge and medical terminology, which requires a significant amount of time and effort for non-professionals to learn and understand. The lack of intuitive and easy-to-understand tools to assist in recognition makes it difficult to correlate recognition results with actual human body parts.
[0050] Therefore, embodiments of the present invention provide a method for automatically labeling human features to address the shortcomings of existing solutions.
[0051] Figure 1 A flowchart of a method for annotating human features provided in an embodiment of the present invention is shown below. Figure 1 The methods include:
[0052] 101. Establish a communication connection with the laser marking device, and calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0053] 102. Obtain the human feature information to be labeled sent by the user, and collect corresponding human images based on the human feature information to be labeled;
[0054] 103. Based on the trained human image recognition model, identify the coordinate positions of the human features to be labeled corresponding to the human image captured;
[0055] 104. The human features to be annotated are annotated to obtain the annotation results, and the annotation results and the coordinate positions are sent to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0056] It should be noted that the execution subject of the embodiments of the present invention is a terminal device, including but not limited to any terminal processing device with camera function such as mobile phones, tablets, and computers used by users.
[0057] Understandably, this terminal device has a corresponding annotation application installed. When users need to make annotations, they need to open the application and follow the instructions on the program to enable the annotation function.
[0058] First, regarding the hardware, this embodiment of the invention involves placing the laser marking device on the terminal device and establishing a communication connection between the laser marking device and the terminal device. The communication connection methods include, but are not limited to, Bluetooth, wireless networks, and mobile communication networks. The terminal device is a common device, such as a mobile phone or tablet. Software is pre-installed on the terminal device; after launching the software, the user can select to connect to the laser marking device and enable its function. The laser marking device is a small-sized laser projection device that can project the image commands transmitted from the terminal device onto a specific area using laser magnification. The installation position of the laser marking device is adjusted to maximize the overlap between the camera's shooting area and the laser projection area of the laser marking device. Then, the two devices are rigidly connected, maintaining their relative positional relationship.
[0059] Specifically, in step 101, a stable communication link is established with the laser marking device via Wi-Fi, Bluetooth, or wired network to ensure the real-time performance and reliability of data transmission, and to determine the correspondence between the laser projection area of the laser marking device and the preset spatial coordinate system.
[0060] In step 102, the system receives a description of the human features to be labeled from the user via voice, text, or graphics. Based on the acquired human feature information, a suitable human body shooting angle and distance are selected, and a human body image containing the specified features is captured using a camera. The system can prompt the user to adjust their posture or position to ensure that a clear, complete, and compliant human body image is captured.
[0061] Next, in step 103, the acquired human body images are input into the trained human image recognition model. This model is based on deep learning algorithms and can automatically identify human body parts, bones, muscles, and other features in the images.
[0062] Based on the model's recognition results, the coordinate position of the human feature to be labeled in the image is determined, and then, combined with the transformation parameters calibrated in step 101, the coordinate position is transformed to a preset spatial coordinate system.
[0063] Finally, in step 104, based on the coordinate positions determined in step 103, the human features to be annotated in the image are labeled, generating intuitive annotation results, such as drawing marker points, outlines, or text descriptions on the image. The annotation results and coordinate position information are sent to the laser annotation device via a communication link. Based on the received coordinate position information, the laser annotation device adjusts the direction and position of the laser projection within a preset spatial coordinate system, ensuring the laser is accurately projected onto the actual position of the human feature to be annotated, and displays the annotation results, such as using laser spots of different colors or shapes to mark the parts or features to be annotated.
[0064] This invention provides a method for annotating human body features, which can automatically annotate human body features, including body parts, bones, muscles, etc. The annotation process is very simple; users only need to send information, which makes it easy for people without medical knowledge to have a simple recognition of human body features.
[0065] Based on the above embodiments, the step of calibrating the laser projection area of the laser marking device in a preset spatial coordinate system includes:
[0066] A printed checkerboard calibration board is placed on a flat scene, and multiple images of the checkerboard from different perspectives are captured by the laser marking device, with each image containing laser highlights;
[0067] The pose of the laser marking device relative to the checkerboard coordinate system is calculated based on the multi-view geometric method.
[0068] Based on the geometric intersection of the laser spot and the calibration pose in the chessboard, the reprojection error is established as the optimization objective function.
[0069] Given the optimization objective function, the relative positional relationship between the laser and the laser marking device is calculated, that is, the position t of the laser starting point in the camera coordinate system.
[0070] It is understandable that the area projected by the installed laser marking device may be different from that of this terminal device, and they need to be normalized to the same coordinate system in order to accurately project and mark the positioning position. Therefore, the embodiments of the present invention first need to complete the spatial position calibration between the two devices, so as to clarify the relative spatial position between the laser marking device and this terminal device, which will facilitate subsequent laser marking.
[0071] Specifically, this embodiment of the invention establishes a carrier coordinate system (a right-handed coordinate system, for example, the horizontal rightward axis is the X-axis, the vertical downward axis is the Y-axis, and the outward-facing camera axis is the Z-axis) based on the terminal device. Then, the position coordinates (i.e., relative positional relationship) of the laser marking device are calibrated in this coordinate system, aligning the two devices in a unified spatial coordinate system. This ensures that the laser mark on the laser marking device aligns with the image on the terminal device during device movement. The calibration method employs an "initial value + optimization" strategy. First, values are given along the horizontal, vertical, and outward-facing camera directions on the X, Y, and Z axes, respectively. These values are not precise but serve as the initial value t0 for the relative positional relationship. Then, a real-scale checkerboard calibration board is prepared, and a function for constructing the geometric intersection relationship between the laser beam and the calibration board is built. The translation residual Δt is optimized, ultimately obtaining the precise relative positional relationship t = t0 + Δt. Implementation details are as follows:
[0072] (1) Checkerboard and Image Processing: A standard checkerboard is printed and flattened. The laser of the laser marking device and the camera of this terminal device are turned on to take multiple images of the checkerboard from different perspectives. The laser spot can be seen in each image. The coordinates of the corner points of the checkerboard and the position coordinates of the laser spot are obtained by using image processing tools. At the same time, each image i provides a laser point observation. The corresponding laser beam is the ray at the intersection of the laser starting point and the checkerboard plane;
[0073] (2) Laser beam geometric modeling: The pose of the camera of this terminal device relative to the chessboard coordinate system is estimated using a multi-view geometric method, denoted as... The laser is oriented fixed forward, let d be the direction. B =[0,0,1] T The starting position of the laser beam when image i is captured. and towards d (i) It can be represented as
[0074]
[0075] Assume the chessboard plane lies at Z = 0 (via the normal vector n = [0, 0, 1)). T Let x0 be the center point of the checkerboard grid, then the intersection point of the laser beam and the plane is... for
[0076]
[0077] Then, this intersection point is projected back onto the image, i.e., its position in the camera coordinate system is...
[0078]
[0079] Where π(·) is the normalized projection function divided by the third dimension, and K is the camera intrinsic parameter matrix. To predict the coordinates of the laser point in the image.
[0080] (3) Optimization Solution: Based on the geometric intersection of the laser beam and the calibration plate, the reprojection error is established as the optimization objective function, defined as follows:
[0081]
[0082] The objective function is To minimize this, based on the least squares principle, the optimal translation deviation can be obtained when N ≥ 4. The final precise relative positional relationship is as follows:
[0083] t = t0 + Δt
[0084] Based on the above embodiments, the step of acquiring the human feature information to be labeled sent by the user, and collecting corresponding human images based on the human feature information to be labeled, includes:
[0085] Acquire voice or text information sent by the user, and identify the human feature information to be labeled from the voice or text information;
[0086] Based on the mapping relationship between the human feature information to be labeled and the preset human body region, the camera is controlled to rotate to the corresponding human body region to capture human body images.
[0087] It is understood that the terminal device used in the embodiments of the present invention needs to receive the feature information that the user wants to annotate before it can enable the function for annotation. There are generally two ways to receive user information: the user tells the terminal through voice input or the user tells the terminal through text input.
[0088] In practical use, the user opens the annotation app installed on the terminal. If voice input is required, the speaker is turned on to receive the user's voice. For example, if the user says "annotate my knees," the system converts it into the text "annotate my knees," and then performs natural language processing on the acquired text information to identify key information. For example, in the text "annotate my knees," the system identifies "knee" as the human feature information to be annotated.
[0089] Furthermore, the terminal device has a built-in mapping relationship between human body feature information and human body regions. Taking "knee" as an example, the terminal device looks up the mapping relationship to determine the area where the "knee" corresponds to the connection between the lower leg and the thigh.
[0090] Finally, based on the above mapping relationship, the terminal device controls the camera to rotate to the corresponding area of the human body. For example, the terminal device controls the camera to rotate to the user's knee area to capture images.
[0091] Based on the above embodiments, the step of identifying the coordinate position of the human feature to be labeled corresponding to the human image captured by the trained human image recognition model includes:
[0092] Based on the trained human image recognition model, the human image is recognized to obtain feature recognition results;
[0093] Based on a similarity matching algorithm, the feature recognition result is matched with the human feature information to be labeled to determine whether they match.
[0094] If a match fails, the human body image will be re-acquired.
[0095] Specifically, after acquiring human images captured by a camera, the terminal device in this embodiment of the invention uses a pre-trained deep learning model (such as ResNet, YOLO, or Mask R-CNN) to optimize feature recognition of human bones, muscles, joints, acupoints, etc., through transfer learning. It can be understood that this process is a pre-training process, employing several commonly used neural network models. A pre-labeled image training set is input to obtain the trained model, which can automatically identify the feature names and locations in the image, ultimately outputting a value containing the location coordinates, confidence level, and category label of the target feature.
[0096] Furthermore, in this embodiment of the invention, the recognition results output by the model are subjected to feature verification using a similarity matching algorithm, and the image I captured by the terminal camera is processed in real time to extract human feature representations of the target region in the image. Then, a similarity matching calculation is performed between the target information Q provided by the user, resulting in a matching score S(F(I),Q). The matching determination function is defined as follows:
[0097]
[0098] Where I is the human image captured by the camera, F(I) is the feature recognition result, Q is the human feature information to be labeled by the user, τ is the similarity threshold, and S(·,·) is the similarity function (such as cosine similarity, Euclidean distance, etc.). The result of the judgment function Match(I,Q) is used for the logic executed by the system: when it equals 0, it prompts the user that the image is not registered and the camera position needs to be readjusted; when it equals 1, it indicates that the match is successful, and the image I is labeled according to the Q information.
[0099] Based on the above embodiments, the step of annotating the human features to be annotated to obtain the annotation result includes:
[0100] The human features to be labeled are explained in text and defined by a bounding box, wherein the bounding box defines the human feature recognition area.
[0101] Understandably, the annotation process includes: defining the target area selected by the user in the form of an annotation box, explaining the user's selected target in the form of text, and visually displaying the user's selected target in the form of a three-dimensional or two-dimensional diagram.
[0102] In this embodiment of the invention, the annotation of human features currently adopts two forms of annotation methods. The first annotation form is a text explanation, which can be either simplified or detailed. For example, if the feature is annotated as the kneecap, its scientific name will be displayed on the knee: "patella". If the user selects the detailed type, it will be displayed as: "Patella, this is a round bone in front of the knee that protects the knee joint and provides leverage support when the leg is straight."
[0103] The second type of annotation is the annotation box annotation, which uses laser-projected contour marks to clearly mark the actual range of the target human body features in the laser projection, helping users to quickly locate the feature area.
[0104] It is understood that users can simultaneously select two annotation formats or one of them. The shape and size of the annotation box are dynamically adjusted according to the annotation features and the acquired image. This embodiment of the invention does not impose specific limitations on this.
[0105] It should be noted that the dynamic annotation method provided in this embodiment of the invention is a dynamic adjustment method, which will accurately follow the feature positions located in the image for projection movement. For example, when the human body posture changes (such as raising an arm), the annotation box and text will also automatically adjust their position and shape accordingly.
[0106] Figure 2 A system structure diagram for annotating human features provided in an embodiment of the present invention, such as... Figure 2 As shown, a system for labeling human body features includes an initialization module 201, an image acquisition module 202, a feature recognition module 203, and an automatic labeling module 204, wherein:
[0107] The initialization module 201 is used to establish a communication connection with the laser marking device and to calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0108] The image acquisition module 202 is used to acquire human feature information to be labeled sent by the user, and to acquire corresponding human images based on the human feature information to be labeled;
[0109] The feature recognition module 203 is used to identify the coordinate position of the human features to be labeled corresponding to the human image captured by the trained human image recognition model.
[0110] The automatic annotation module 204 is used to annotate the human features to be annotated to obtain annotation results, and send the annotation results and the coordinate positions to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0111] It is understood that the system for annotating human body features provided by the present invention corresponds to the methods for annotating human body features provided in the foregoing embodiments, and will not be described again here.
[0112] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating an embodiment of the electronic device provided in this invention. For example... Figure 3 As shown, this embodiment of the invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it performs the following steps:
[0113] Establish a communication connection with the laser marking device, and calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0114] Obtain the human feature information to be labeled sent by the user, and collect corresponding human images based on the human feature information to be labeled;
[0115] Based on the trained human image recognition model, the coordinate positions of the human features to be labeled corresponding to the human image are identified.
[0116] The human features to be annotated are annotated to obtain the annotation results, and the annotation results and the coordinate positions are sent to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0117] An embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0118] Establish a communication connection with the laser marking device, and calibrate the laser projection area of the laser marking device in a preset spatial coordinate system;
[0119] Obtain the human feature information to be labeled sent by the user, and collect corresponding human images based on the human feature information to be labeled;
[0120] Based on the trained human image recognition model, the coordinate positions of the human features to be labeled corresponding to the human image are identified.
[0121] The human features to be annotated are annotated to obtain the annotation results, and the annotation results and the coordinate positions are sent to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions.
[0122] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0123] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0125] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0126] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0127] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0128] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for annotating human body features, characterized in that, include: Establish a communication connection with the laser marking device, and calibrate the laser projection area of the laser marking device in a preset spatial coordinate system so that the two devices are aligned in a unified spatial coordinate system, so that the laser mark of the laser marking device can be aligned with the image of this terminal device when the devices are moving. Obtain the human feature information to be labeled sent by the user, and collect corresponding human images based on the human feature information to be labeled; Based on the trained human image recognition model, the coordinate positions of the human features to be labeled corresponding to the human image are identified. The human features to be annotated are annotated to obtain annotation results, and the annotation results and the coordinate positions are sent to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions. The step of calibrating the laser projection area of the laser marking device within a preset spatial coordinate system includes: A printed checkerboard calibration board is placed on a flat scene, and multiple images of the checkerboard from different perspectives are captured by the laser marking device, with each image containing laser highlights; Based on the multi-view geometric method, the pose of the laser marking device relative to the checkerboard coordinate system is calculated; based on the geometric intersection of the laser spot and the marking pose in the checkerboard, the reprojection error is established as the optimization objective function. Given the optimization objective function, the relative positional relationship with the laser marking device is calculated; The method of identifying the coordinates of the human features to be labeled in the human image based on the trained human image recognition model includes: Based on a trained human image recognition model, the human image is recognized to obtain feature recognition results; based on a similarity matching algorithm, the feature recognition results are matched with the human feature information to be labeled. If a match fails, the human body image will be re-acquired.
2. The method for annotating human body features according to claim 1, characterized in that, The multi-view geometry method is used to calculate the pose of the laser marking device relative to the checkerboard coordinate system, including: The pose is represented as ; The laser is oriented fixed forward. The starting position of the laser beam when image i is captured and orientation Represented as The chessboard plane is located in Through the normal vector And the center point of the chessboard This indicates the intersection point of the laser beam and the plane. for The intersection point is projected back onto the image, i.e., its position in the camera coordinate system is... in The normalization is divided by the projection function in the third dimension, where K is the camera intrinsic parameter matrix. To predict the coordinates of the laser point in the image.
3. The method for annotating human body features according to claim 1, characterized in that, The step of acquiring the human feature information to be labeled sent by the user, and collecting corresponding human images based on the human feature information to be labeled, includes: Acquire voice or text information sent by the user, and identify the human feature information to be labeled from the voice or text information; Based on the mapping relationship between the human feature information to be labeled and the preset human body region, the camera is controlled to rotate to the corresponding human body region to capture human body images.
4. The method for annotating human body features according to claim 1, characterized in that, The similarity-based matching algorithm matches whether the feature recognition result matches the human feature information to be labeled, including: Where I represents the human body image captured by the camera. This is the feature recognition result, where Q is the human feature information to be labeled input by the user, and τ is the similarity threshold. This is a similarity function.
5. The method for annotating human body features according to claim 1, characterized in that, The annotation of the human features to be annotated to obtain the annotation result includes: The human features to be labeled are explained in text and defined by a bounding box, wherein the bounding box defines the human feature recognition area.
6. A system for labeling human body features, characterized in that, include: An initialization module is used to establish a communication connection with the laser marking device and to calibrate the laser projection area of the laser marking device in a preset spatial coordinate system; The image acquisition module is used to acquire human feature information to be labeled sent by the user, and to acquire corresponding human images based on the human feature information to be labeled. The feature recognition module is used to identify the coordinate position of the human features to be labeled in the human image based on the trained human image recognition model. An automatic annotation module is used to annotate the human features to be annotated to obtain annotation results, and send the annotation results and the coordinate positions to the laser annotation device so that the laser annotation device can project and display the annotation results according to the coordinate positions. The step of calibrating the laser projection area of the laser marking device in a preset spatial coordinate system includes: printing a checkerboard calibration board and placing it on a planar scene, and acquiring multiple images of the checkerboard from different perspectives captured by the laser marking device, with each image containing a laser highlight. Based on the multi-view geometric method, the pose of the laser marking device relative to the checkerboard coordinate system is calculated; based on the geometric intersection of the laser spot and the marking pose in the checkerboard, the reprojection error is established as the optimization objective function. Given the optimization objective function, the relative positional relationship between the laser marking device and the target device is calculated.
7. An electronic device, characterized in that, The device includes a memory and a processor, wherein the processor is used to execute computer management programs stored in the memory to implement the steps of the method for annotating human features as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer management program, which, when executed by a processor, implements the steps of the method for annotating human features as described in any one of claims 1-5.