Key point detection method, device, computer device and storage medium

By performing the fusion process of feature extraction and target object detection on the detected image, the key points of the target object are detected directly from the image, solving the problem of misrelevance in traditional methods, improving the detection accuracy and saving training costs.

CN114332484BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111329254.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-07-18
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Traditional key point detection methods lack the overall information of the target object, which leads to false associations when the key points are associated with the target object, reducing the detection accuracy.

Method used

By performing feature extraction and target object detection on the image to be detected, a first feature map and a second feature map are generated and fused to determine the key point feature parameters of the target object, thereby directly detecting the key points of the target object from the first feature map.

Benefits of technology

Improve the accuracy of key point detection of target objects, avoid the step of associating key points with their target objects, and save training time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332484B_ABST
    Figure CN114332484B_ABST
Patent Text Reader

Abstract

The present application relates to a key point detection method, apparatus, computer device, and storage medium, belonging to the field of artificial intelligence technology. The method includes: performing feature extraction processing on an image to be detected to obtain a first feature map of the image to be detected; performing target object detection processing on the first feature map to obtain a second feature map of the target object; fusing the first feature map and the second feature map to obtain a fused feature map; determining key point feature parameters of the target object based on the fused feature map; and detecting key points of the target object from the first feature map based on the key point feature parameters. Using this method can improve the accuracy of key point detection of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence technology, and more particularly to the field of image processing technology, especially a key point detection method, device, computer device, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, key point detection technology has emerged. Key point detection refers to locating the positions of key points from the image to be detected. In traditional technology, usually in a bottom-up manner, all key points in the image are first identified, and then through auxiliary information and post-processing means, the identified key points are associated with the target objects to which they belong, so as to obtain the final key point detection result.

[0003] However, traditional key point detection methods lack the overall information of the target object. Therefore, when associating key points with the target objects to which they belong, mis-association is likely to occur, resulting in a low accuracy rate of key point detection in the target object. Summary of the Invention

[0004] Based on this, it is necessary to provide a key point detection method, device, computer device, and storage medium that can improve the accuracy rate of key point detection for the above technical problems.

[0005] A key point detection method, the method comprising:

[0006] Performing feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected;

[0007] Performing target object detection processing on the first feature map to obtain a second feature map of the target object;

[0008] Fusing the first feature map and the second feature map to obtain a fused feature map;

[0009] Based on the fused feature map, determining key point feature parameters of the target object;

[0010] Based on the key point feature parameters, detecting key points of the target object from the first feature map.

[0011] A key point detection device, the device comprising:

[0012] An extraction module, configured to perform feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected;

[0013] A detection module, configured to perform target object detection processing on the first feature map to obtain a second feature map of the target object;

[0014] A fusion module, configured to fuse the first feature map and the second feature map to obtain a fused feature map;

[0015] A determination module, configured to determine key point feature parameters of the target object based on the fused feature map;

[0016] The detection module is configured to detect key points of the target object from the first feature map based on the key point feature parameters.

[0017] In one embodiment, the extraction module is further configured to obtain an original feature map of the image to be detected; perform convolution on the original feature map to obtain a convolved feature map; perform upsampling on the original feature map to obtain an upsampled feature map; fuse the convolved feature map and the upsampled feature map to obtain a fused feature map; perform convolution on the fused feature map to obtain the first feature map of the image to be detected.

[0018] In one embodiment, there are multiple target objects, and different types of target objects are included among the multiple target objects; the detection module is further configured to perform convolution on the first feature map to obtain multiple intermediate feature maps; perform convolution on the multiple intermediate feature maps to fuse the features of target objects of the same type into the same feature map, so as to obtain a second feature map corresponding to each type respectively.

[0019] In one embodiment, the detection module is further configured to perform target object detection processing on the first feature map to obtain a first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value respectively; the first probability value is used to represent the probability that a target object exists at the position of the corresponding pixel point; divide the first probability feature map into a preset number of first image blocks with the same size; for each of the first image blocks, select the first probability value with the largest probability value from the first image block as the first target probability value; determine the pixel points corresponding to the probability values greater than the first preset probability value of the first target probability value as first target pixel points; generate the second feature map of the target object according to the first target pixel points.

[0020] In one embodiment, the second feature map is generated by an object detection network in a trained key point detection model; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network; the determining module is further configured to input the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the target object; the detecting module is further configured to use the key point feature parameters as the convolutional parameters of the second convolutional network, and perform convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

[0021] In one embodiment, the detecting module is further configured to use the key point feature parameters as the convolutional parameters of the second convolutional network, so that the second convolutional network determines a target region in the first feature map based on the key point feature parameters; the target region is the region of the key points of the target object in the first feature map; the key points of the target object are detected from the target region based on the second convolutional network.

[0022] In one embodiment, the apparatus further includes: a training module, configured to obtain a sample image containing a target object; input the sample image into a key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained; predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained; determine a first loss value between the predicted attribute information and the target attribute information of the target object; determine a second loss value between the predicted key point information and the target key point information of the target object; determine a target loss value according to the first loss value and the second loss value; perform iterative training on the key point detection model to be trained in a direction of reducing the target loss value until the iterative stop condition is met, and obtain a trained key point detection model.

[0023] In one embodiment, the to-be-trained key point detection network includes a to-be-trained first convolutional network; the predicted attribute information includes a predicted object heat map; the predicted key point information includes a predicted key point heat map; the training module is further configured to predict, through the to-be-trained target detection network, the predicted object heat map of the target object in the sample image; fuse the predicted object heat map and the feature map of the sample image to obtain a sample fusion feature map, and input the sample fusion feature map into the to-be-trained first convolutional network to output the predicted key point feature parameters; based on the predicted key point feature parameters, predict the key points of the target object from the feature map of the sample image, and generate the predicted key point heat map of the target object based on the predicted key points.

[0024] In one embodiment, the predicted object heat map is obtained by performing heat map coordinate transformation on the coordinates of the center point of the target object predicted in the sample image by the to-be-trained target detection network; the predicted attribute information further includes the predicted size information of the bounding box corresponding to the target object, and the conversion error corresponding to the center point of the target object; the conversion error is the error generated when performing heat map coordinate transformation on the coordinates of the center point.

[0025] In one embodiment, the detection module is further configured to perform convolution on the first feature map according to the key point feature parameters to obtain a second probability feature map; each pixel point in the second probability feature map corresponds to a second probability value; the second probability value is used to represent the probability that a key point exists at the position of the corresponding pixel point; divide the second probability feature map into a preset number of second image blocks with the same size; for each of the second image blocks, select the second probability value with the largest probability value from the second image block as the second target probability value; determine the pixel points corresponding to the probability values greater than the second preset probability value in the second target probability value as the second target pixel points; use the second target pixel points as the key points of the target object.

[0026] In one embodiment, the to-be-detected image is an image collected in a point reading scenario; the target object is an input entity used to trigger point reading in the point reading scenario; the device further includes: a point reading module, configured to determine a target point reading text based on the key points of the input entity; and perform point reading processing based on the target point reading text.

[0027] In one embodiment, there are multiple input entities, and the multiple input entities include different types of input entities; the point reading module is further configured to, according to the priorities respectively corresponding to each type in the different types of input entities, use the input entity corresponding to the highest priority type as the target input entity; determine the key points of the target input entity as the target key points, and determine the target point reading text pointed to by the target key points.

[0028] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] Perform feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected;

[0030] Perform target object detection processing on the first feature map to obtain a second feature map of the target object;

[0031] Fuse the first feature map and the second feature map to obtain a fused feature map;

[0032] Based on the fused feature map, determine the key point feature parameters of the target object;

[0033] Based on the key point feature parameters, detect the key points of the target object from the first feature map.

[0034] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0035] Perform feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected;

[0036] Perform target object detection processing on the first feature map to obtain a second feature map of the target object;

[0037] Fuse the first feature map and the second feature map to obtain a fused feature map;

[0038] Based on the fused feature map, determine the key point feature parameters of the target object;

[0039] Based on the key point feature parameters, detect the key points of the target object from the first feature map.

[0040] A computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0041] Perform feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected;

[0042] Perform object detection processing on the first feature map to obtain a second feature map of the target object;

[0043] Fuse the first feature map and the second feature map to obtain a fused feature map;

[0044] Based on the fused feature map, determine the key point feature parameters of the target object;

[0045] Based on the key point feature parameters, detect the key points of the target object from the first feature map.

[0046] The above key point detection method, device, computer device, and storage medium can obtain the first feature map of the image to be detected by performing feature extraction processing on the image to be detected. By performing object detection processing on the first feature map, a second feature map of the target object including the overall information of the target object can be obtained. By fusing the first feature map and the second feature map, a fused feature map can be obtained. Based on the fused feature map, the key point feature parameters of the target object are determined. Since the image to be detected is changing, the obtained key point feature parameters will also change dynamically with the image to be detected. Furthermore, based on the key point feature parameters, the key points of the target object can be directly detected from the first feature map, avoiding the step of associating the key points with their corresponding target objects, and improving the accuracy of key point detection of the target object. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is an application environment diagram of the key point detection method in an embodiment;

[0048] Figure 2 It is a flowchart of the key point detection method in an embodiment;

[0049] Figure 3 It is a structural diagram of the key point detection model in an embodiment;

[0050] Figure 4 It is a predicted object heat map of the target object in an embodiment;

[0051] Figure 5 It is a predicted object heat map of all target objects in the sample image in an embodiment;

[0052] Figure 6 It is a predicted key point heat map of the key points in an embodiment;

[0053] Figure 7 It is a predicted key point heat map of all key points in the sample image in an embodiment;

[0054] Figure 8 Schematic diagram of all key points detected from an image to be detected in one embodiment;

[0055] Figure 9 Schematic diagram of target key points detected from an image to be detected in one embodiment;

[0056] Figure 10 Schematic diagram of dot-reading processing based on target dot-reading text in one embodiment;

[0057] Figure 11 Schematic diagram of dot-reading processing based on target dot-reading text in another embodiment;

[0058] Figure 12 Schematic flow chart of a key point detection method in another embodiment;

[0059] Figure 13 Structural block diagram of a key point detection device in one embodiment;

[0060] Figure 14 Structural block diagram of a key point detection device in another embodiment;

[0061] Figure 15 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0062] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] The key point detection method provided by the present application can be applied to an application scenario such as Figure 1 In this application scenario, the user touches the text in the book 102 by hand to obtain the corresponding image 104 to be detected. The computer device 106 can acquire the image 104 to be detected, and can perform feature extraction processing on the image 104 to be detected to obtain the first feature map of the image 104 to be detected, and perform target object detection processing on the first feature map to obtain the second feature map of the target object. The computer device 106 can fuse the first feature map and the second feature map to obtain a fused feature map, and determine the key point feature parameters of the target object based on the fused feature map. The computer device 106 can detect the key points of the target object from the first feature map based on the key point feature parameters.

[0064] Among them, the computer device 106 may include a terminal and a server. The terminal may be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, portable wearable devices, vehicle-mounted terminals, and reading devices. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0065] It should be noted that the key point detection method in some embodiments of this application uses artificial intelligence technology. For example, the first feature map of the image to be detected and the second feature map of the target object belong to the feature maps obtained by using artificial intelligence technology for feature extraction. In addition, the key points of the target object also belong to the key points detected by using artificial intelligence technology.

[0066] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0067] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0068] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. The key point detection method in some embodiments of this application uses computer vision technology. For example, when a computer device performs feature extraction processing on a to-be-detected image to obtain a first feature map of the to-be-detected image, it belongs to the feature map obtained by using computer vision technology for feature extraction.

[0069] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0070] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0071] In one embodiment, as Figure 2 shown, a key point detection method is provided. In this embodiment, taking the application of this method to the Figure 1 computer device 106 as an example, the method includes the following steps:

[0072] Step 202, perform feature extraction processing on the to-be-detected image to obtain a first feature map of the to-be-detected image.

[0073] Among them, the image to be detected is the image for which key point detection is to be performed. The first feature map is the feature map of the image to be detected itself.

[0074] Specifically, the computer device can obtain the image to be detected and perform feature extraction processing on the obtained image to be detected to obtain the first feature map of the image to be detected.

[0075] In one embodiment, the image to be detected can be a single-channel or multi-channel image. For example, the image to be detected can be a single-channel grayscale image or a 3-channel RGB (Red, Green, Blue) image.

[0076] In one embodiment, the computer device can obtain the image to be detected and perform scaling processing on the image to be detected according to a preset image size. For example, perform scaling processing on the image to be detected according to an image size of 320*320. Furthermore, the computer device can perform feature extraction processing on the scaled image to be detected to obtain the first feature map of the image to be detected.

[0077] In one embodiment, a trained key point detection model can be run in the computer device, and the trained key point detection model includes a feature extraction network. The computer device can obtain the image to be detected and input the image to be detected into the feature extraction network to perform feature extraction processing on the image to be detected through the feature extraction network to obtain the first feature map of the image to be detected.

[0078] In one embodiment, the feature extraction network includes a backbone network. The computer device can obtain the image to be detected and input the image to be detected into the backbone network to perform preliminary feature extraction processing on the image to be detected through the backbone network to obtain the original feature map of the image to be detected. Furthermore, the computer device can perform further feature extraction processing on the original feature map to obtain the first feature map of the image to be detected. Among them, the original feature map is the feature map obtained by performing preliminary feature extraction processing on the image to be detected.

[0079] Step 204, perform target object detection processing on the first feature map to obtain the second feature map of the target object.

[0080] Among them, the target object is the detection object as the target. The second feature map is the feature map of the target object itself.

[0081] Specifically, the computer device can perform convolution on the first feature map to obtain the features of the target object in the first feature map. Furthermore, the computer device can generate the second feature map of the target object based on the features of the target object in the first feature map.

[0082] In one embodiment, a trained key point detection model can be run on a computer device. The computer device can input a first feature map into the trained key point detection model, and perform convolution on the first feature map through the trained key point detection model to obtain the features of the target object in the first feature map. The computer device can generate a second feature map of the target object based on the features of the target object in the first feature map.

[0083] In one embodiment, a trained key point detection model can be run on a computer device, and a target detection network is included in the trained key point detection model. The computer device can input a first feature map into the target detection network to perform convolution on the first feature map through the target detection network to obtain the features of the target object in the first feature map. The computer device can generate a second feature map of the target object based on the features of the target object in the first feature map.

[0084] In one embodiment, the computer device can perform convolution on the first feature map for feature learning to obtain an intermediate feature map. The computer device can then perform convolution on the intermediate feature map to obtain a second feature map of the target object. Among them, the intermediate feature map is a feature map in an intermediate state during the process of performing target object detection processing on the first feature map to generate the second feature map of the target object.

[0085] In one embodiment, the computer device can perform convolution on the intermediate feature map to obtain a first probability feature map. Furthermore, the computer device can perform max pooling on the first probability feature map and generate a second feature map of the target object based on the result after max pooling. Among them, the first probability feature map is used to represent the probability that there is a target object in the first feature map.

[0086] Step 206: Fuse the first feature map and the second feature map to obtain a fused feature map.

[0087] In one embodiment, when the computer device fuses the first feature map and the second feature map, specifically, it can perform feature splicing on the first feature map and the second feature map to obtain a spliced feature map, and use the spliced feature map as the fused feature map. It should be noted that when the first feature map and the second feature map are fused, specifically, it can also be other feature fusion methods other than feature splicing. This embodiment does not limit the specific method of feature fusion.

[0088] Step 208: Determine the key point feature parameters of the target object based on the fused feature map.

[0089] Among them, the key point feature parameters are parameters used to represent key point features. The key point features are features used to represent key points.

[0090] Specifically, the computer device can extract the features of the key points of the target object from the fused feature map. Furthermore, the computer device can determine the key point feature parameters of the target object based on the features of the key points of the target object.

[0091] In one embodiment, a trained key point detection model can be run in the computer device. The computer device can extract the features of the key points of the target object from the fused feature map through the trained key point detection model. Furthermore, the computer device can determine the key point feature parameters of the target object based on the features of the key points of the target object through the trained key point detection model again.

[0092] In one embodiment, a trained key point detection model can be run in the computer device. The trained key point detection model includes a key point detection network, and the key point detection network includes a first convolutional network. The computer device can input the fused feature map into the first convolutional network to perform convolution on the fused feature map through the first convolutional network to obtain the key point feature parameters of the target object.

[0093] Step 210, detect the key points of the target object from the first feature map based on the key point feature parameters.

[0094] Specifically, the computer device can determine the position area where the key points of the target object are located based on the key point feature parameters. Furthermore, the computer device can detect the key points of the target object from the first feature map based on the determined position area.

[0095] In one embodiment, a trained key point detection model can be run in the computer device. The trained key point detection model includes a key point detection network, and the key point detection network includes a second convolutional network. The computer device can use the key point feature parameters as the convolutional parameters of the second convolutional network. Furthermore, input the first feature map into the second convolutional network and perform convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

[0096] In one embodiment, the computer device can perform convolution on the first feature map according to the key point feature parameters to obtain a second probability feature map. Furthermore, the computer device can perform maximum pooling processing on the second probability feature map and detect the key points of the target object from the first feature map based on the result of the maximum pooling processing. The second probability feature map is used to represent the probability that the key points of the target object exist in the first feature map.

[0097] In the above key point detection method, by performing feature extraction processing on the image to be detected, a first feature map of the image to be detected can be obtained. By performing target object detection processing on the first feature map, a second feature map of the target object including the overall information of the target object can be obtained. By fusing the first feature map and the second feature map, a fused feature map can be obtained. Based on the fused feature map, the key point feature parameters of the target object are determined. Since the image to be detected is variable, the obtained key point feature parameters will also change dynamically with the image to be detected. Furthermore, based on the key point feature parameters, the key points of the target object can be directly detected from the first feature map, avoiding the step of associating the key points with their corresponding target objects, and improving the accuracy of key point detection of the target object.

[0098] Meanwhile, compared with the traditional top-down key point detection method, that is, first detecting the target object through a target detection model, and then detecting the key points of the target object through a key point detection model independent of the target detection model, the present application proposes a brand-new key point detection method. The key point detection method of the present application only needs to train one model to achieve key point detection of the target object, instead of separately training two independent models as in the traditional top-down key point detection method, saving time costs.

[0099] In one embodiment, performing feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected includes: obtaining the original feature map of the image to be detected; performing convolution on the original feature map to obtain a convolved feature map; performing upsampling on the original feature map to obtain an upsampled feature map; fusing the convolved feature map and the upsampled feature map to obtain a fused feature map; and performing convolution on the fused feature map to obtain the first feature map of the image to be detected.

[0100] Among them, performing upsampling on the original feature map means performing size enlargement processing on the original feature map. The fused feature map is the feature map generated by fusing the convolved feature map and the upsampled feature map.

[0101] Specifically, the computer device can obtain the original feature map of the image to be detected and perform convolution on the obtained original feature map to obtain a convolved feature map. The computer device can perform upsampling on the original feature map, that is, perform size enlargement processing on the original feature map, to obtain an upsampled feature map. Furthermore, the computer device can perform feature fusion on the convolved feature map and the upsampled feature map to obtain a fused feature map, and perform convolution on the fused feature map to obtain the first feature map of the image to be detected.

[0102] In one embodiment, a trained key point detection model can be run on a computer device. The trained key point detection model includes a feature extraction network, and the feature extraction network includes a backbone network feature convolutional network. The computer device can input the image to be detected into the backbone network to perform preliminary feature extraction processing on the image to be detected through the backbone network, and obtain the original feature map of the image to be detected. Furthermore, the computer device can input the original feature map into the feature convolutional network to perform convolution on the original feature map through the feature convolutional network, obtain the convolved feature map, perform upsampling on the original feature map to obtain the upsampled feature map, and fuse the convolved feature map and the upsampled feature map to obtain the fused feature map. Finally, perform convolution on the re-fused feature map to obtain the first feature map of the image to be detected.

[0103] In one embodiment, the computer device can input the image to be detected into the backbone network. For example, input the image to be detected with a size of 320*320 into the backbone network to perform preliminary feature extraction processing on the image to be detected through the backbone network, and obtain the original feature map of the image to be detected. Furthermore, the computer device can input the original feature map into the feature convolutional network to perform 1*1 convolution on the original feature map through the feature convolutional network to obtain the convolved feature map. At the same time, perform upsampling on the original feature map in the way of FPN (Feature Pyramid Networks) to obtain the upsampled feature map, and fuse the convolved feature map and the upsampled feature map to obtain the fused feature map. For example, a fused feature map with a size of 80*80 can be obtained. Finally, perform 3*3 convolution on the re-fused feature map to obtain the first feature map of the image to be detected. The number of the first feature maps is N, and N is a natural number.

[0104] In one embodiment, the backbone network can be any neural network. For example, the backbone network can be any one of MoileNetV1 (Mobile Network Version 1), MoileNetV2 (Mobile Network Version 2), VGG (Visual Geometry Group Network), ResNet (Residual Network), etc.

[0105] In the above embodiment, by performing convolution on the original feature map of the image to be detected, a more abstract convolved feature map can be obtained. By performing upsampling on the original feature map, a more specific upsampled feature map can be obtained. Furthermore, by fusing the convolved feature map and the upsampled feature map, a fused feature map can be obtained. By performing convolution on the fused feature map, a better first feature map of the image to be detected can be obtained, and thus the key point detection accuracy of the target object can be further improved.

[0106] In one embodiment, there are multiple target objects, and different types of target objects are included among the multiple target objects. Performing target object detection processing on the first feature map to obtain the second feature map of the target object includes: performing convolution on the first feature map to obtain multiple intermediate feature maps; performing convolution on the multiple intermediate feature maps to fuse the features of target objects of the same type into the same feature map, thereby obtaining a second feature map corresponding to each type.

[0107] Specifically, the computer device can perform convolution on the first feature map to obtain multiple intermediate feature maps, and perform convolution on the multiple intermediate feature maps to fuse the features of target objects of the same type into the same feature map, thereby obtaining a second feature map corresponding to each type. It can be understood that there are multiple types of target objects, and one type of target object corresponds to one second feature map.

[0108] In one embodiment, the computer device can perform 3*3 convolution on the first feature map for feature learning to obtain multiple intermediate feature maps. Further, the computer device can perform 1*1 convolution on the multiple intermediate feature maps to fuse the features of target objects of the same type into the same feature map, thereby obtaining a second feature map corresponding to each type. For example, if there are two types of target objects, namely the first type and the second type, the target objects of the first type correspond to one second feature map for representing the target objects, and the target objects of the second type correspond to another second feature map for representing the target objects.

[0109] In the above embodiment, by performing convolution on the first feature map for feature learning, multiple intermediate feature maps are obtained. By performing convolution on the multiple intermediate feature maps again, the features of target objects of the same type can be fused into the same feature map, thereby obtaining a second feature map corresponding to each type, so that different types of target objects can be detected, and the detection accuracy of the target objects can be improved.

[0110] In one embodiment, performing target object detection processing on the first feature map to obtain the second feature map of the target object includes: performing target object detection processing on the first feature map to obtain the first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value; the first probability value is used to represent the probability that a target object exists at the position corresponding to the pixel point; dividing the first probability feature map into a preset number of first image blocks with the same size; for each first image block, selecting the first probability value with the largest probability value from the first image block as the first target probability value; determining the pixel points corresponding to the probability values where the first target probability value is greater than the first preset probability value as the first target pixel points; generating the second feature map of the target object according to the first target pixel points.

[0111] Among them, the first image block is an image block obtained by dividing the first probability feature map. The first target probability value is the first probability value serving as the target. The first target pixel point is a pixel point where the target object actually exists at the corresponding position.

[0112] Specifically, the computer device can perform target object detection processing on the first feature map to obtain the first probability feature map, and divide the first probability feature map into a preset number of first image blocks of the same size. For each first image block, the computer device can select the first probability value with the largest probability value from the first image block as the first target probability value. The computer device can compare the first target probability value with the first preset probability value, and determine the pixel points corresponding to the probability values where the first target probability value is greater than the first preset probability value as the first target pixel points. Furthermore, the computer device can generate a second feature map of the target object according to the first target pixel points.

[0113] In the above embodiment, by performing target object detection processing on the first feature map, the first probability feature map corresponding to the first feature map can be obtained. The first probability feature map is divided into multiple first image blocks, and the first probability value with the largest probability value is selected from each first image block as the first target probability value. Then, the pixel points corresponding to the probability values where the first target probability value is greater than the first preset probability value are determined as the first target pixel points. It can be understood that the position where the first target pixel points are located is the position where the target object is located. Furthermore, according to the first target pixel points, a second feature map of the target object can be generated, thereby improving the detection accuracy of the target object.

[0114] In one embodiment, the computer device can perform convolution on the first feature map to obtain an intermediate feature map. The computer device can perform target object detection processing on the intermediate feature map to obtain the first probability feature map, and divide the first probability feature map into a preset number of first image blocks of the same size. For each first image block, the computer device can select the first probability value with the largest probability value from the first image block as the first target probability value. The computer device can compare the first target probability value with the first preset probability value, and determine the pixel points corresponding to the probability values where the first target probability value is greater than the first preset probability value as the first target pixel points. Furthermore, the computer device can generate a second feature map of the target object according to the first target pixel points.

[0115] In one embodiment, the second feature map is generated by the object detection network in the trained key point detection model; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network. Based on the fused feature map, determining the key point feature parameters of the target object includes: inputting the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the target object; based on the key point feature parameters, detecting the key points of the target object from the first feature map includes: using the key point feature parameters as the convolutional parameters of the second convolutional network, and performing convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

[0116] Among them, the convolutional parameters are the parameters required for the second convolutional network to perform convolution operations.

[0117] Specifically, the trained key point detection model includes an object detection network and a key point detection network, and the key point detection network includes a first convolutional network and a second convolutional network. The computer device can input the first feature map into the object detection network to perform object detection processing on the first feature map through the object detection network to obtain the second feature map of the target object. The computer device can fuse the first feature map and the second feature map to obtain the fused feature map. The computer device can input the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the target object. Furthermore, the computer device can use the key point feature parameters as the convolutional parameters of the second convolutional network and perform convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

[0118] In one embodiment, the computer device can use the key point feature parameters as the convolutional parameters of the second convolutional network so that the second convolutional network determines the specific position of the target object in the first feature map based on the key point feature parameters. Furthermore, the computer device can, based on the second convolutional network, detect the key points of the target object from the first feature map according to the specific position of the target object.

[0119] In one embodiment, the computer device can obtain a sample image containing a target object and input the sample image into a key point detection model to be trained, so as to predict a prediction result corresponding to the sample image through the key point detection model to be trained. The computer device can determine the loss value between the prediction result and the sample result corresponding to the sample image, and iteratively train the key point detection model to be trained in the direction of reducing the loss value until the trained key point detection model is obtained when the iteration stop condition is met. Among them, the sample image is a training image for training the key point detection model to be trained. The prediction result is the result predicted by the key point detection model to be trained based on the input sample image during the process of training the key point detection model to be trained. The sample result is the result pre-annotated for the sample image.

[0120] In one embodiment, the first convolutional network can be a dynamic convolutional kernel, and the parameters of the dynamic convolutional kernel, that is, the key point feature parameters, can change dynamically with different inputs.

[0121] In the above embodiment, by inputting the fused feature map into the first convolutional network for convolution, the key point feature parameters of the target object associated with the input can be output. It can be understood that the key point feature parameters can change dynamically with different inputs. Furthermore, after using the key point feature parameters as the convolutional parameters of the second convolutional network, by performing convolution on the first feature map through the second convolutional network, the key points of the target object can be detected from the first feature map, further improving the accuracy of key point detection of the target object.

[0122] In one embodiment, using the key point feature parameters as the convolutional parameters of the second convolutional network and performing convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map includes: using the key point feature parameters as the convolutional parameters of the second convolutional network so that the second convolutional network determines a target region in the first feature map based on the key point feature parameters; the target region is the region of the key points of the target object in the first feature map; and detecting the key points of the target object from the target region based on the second convolutional network.

[0123] Specifically, the computer device can use the key point feature parameters as the convolutional parameters of the second convolutional network so that the second convolutional network can determine a target region in the first feature map based on the key point feature parameters. Furthermore, the computer device can detect the key points of the target object from the target region based on the second convolutional network.

[0124] In the above embodiments, by using the key point feature parameters as the convolution parameters of the second convolutional network, the second convolutional network can determine the target region in the first feature map based on the key point feature parameters, so that the key points of the target object can be detected from the target region by the second convolutional network, further improving the efficiency and accuracy of key point detection.

[0125] In one embodiment, the steps of obtaining the trained key point detection model include: obtaining the trained key point detection model, including: acquiring a sample image containing a target object; inputting the sample image into the key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained; predicting the predicted attribute information of the target object in the sample image by the target detection network to be trained, and predicting the predicted key point information of the target object by the key point detection network to be trained; determining a first loss value between the predicted attribute information and the target attribute information of the target object; determining a second loss value between the predicted key point information and the target key point information of the target object; determining a target loss value according to the first loss value and the second loss value; iteratively training the key point detection model to be trained in the direction of reducing the target loss value until the iterative stop condition is met, and obtaining the trained key point detection model.

[0126] Among them, the predicted attribute information is the attribute information predicted by the key point detection model to be trained based on the target object in the input sample image during the training process of the key point detection model to be trained. The predicted key point information is the key point information predicted by the key point detection model to be trained based on the target object in the input sample image during the training process of the key point detection model to be trained. The target attribute information is the pre-annotated attribute information for the target object in the sample image. The target key point information is the pre-annotated key point information for the target object in the sample image. The first loss value is the error between the predicted attribute information and the target attribute information of the target object. The second loss value is the error between the predicted key point information and the target key point information of the target object. The target loss value is the loss value as the target.

[0127] Specifically, the computer device can obtain a sample image containing the target object and input the sample image into the key point detection model to be trained. The computer device can predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained. The computer device can use the difference between the predicted attribute information and the target attribute information of the target object as the first loss value, and the difference between the predicted key point information and the target key point information of the target object as the second loss value. The computer device can perform weighted fusion on the first loss value and the second loss value to obtain the target loss value, and iteratively train the key point detection model to be trained in the direction of reducing the target loss value until the iterative stop condition is met, and then obtain the trained key point detection model.

[0128] In one embodiment, the above prediction results include predicted attribute information and predicted key point information, and the above sample results include target attribute information and target key point information.

[0129] In one embodiment, the key point detection network to be trained includes a first convolutional network to be trained, the predicted attribute information includes a predicted object feature map, and the predicted key point information includes a predicted key point feature map. The computer device can predict the predicted object feature map of the target object in the sample image through the target detection network to be trained. The computer device can fuse the predicted object feature map and the feature map of the sample image, and input the fused feature map of the predicted object feature map and the feature map of the sample image into the first convolutional network to be trained to output the predicted key point feature parameters. The computer device can predict the key points of the target object from the feature map of the sample image based on the predicted key point feature parameters, and generate the predicted key point feature map of the target object based on the predicted key points. The predicted object feature map is the feature map predicted by the key point detection model to be trained based on the target object in the input sample image during the process of training the key point detection model to be trained. The predicted key point feature map is the key point feature map predicted by the key point detection model to be trained based on the target object in the input sample image during the process of training the key point detection model to be trained.

[0130] In the above embodiments, by inputting a sample image containing a target object into a key point detection model to be trained including a target detection network to be trained and a key point detection network to be trained, the predicted attribute information of the target object in the sample image can be quickly predicted through the target detection network to be trained, and the predicted key point information of the target object can be quickly predicted through the key point detection network to be trained. Determine a first loss value between the predicted attribute information and the target attribute information of the target object, and determine a second loss value between the predicted key point information and the target key point information of the target object, so that the target loss value can be accurately determined according to the first loss value and the second loss value. Iteratively train the key point detection model to be trained in the direction of reducing the target loss value. When the iteration stop condition is met, a trained key point detection model can be obtained, so that the finally obtained key point detection model has the ability to detect both the target object and the key points of the target object.

[0131] In one embodiment, the key point detection network to be trained includes a first convolutional network to be trained; the predicted attribute information includes a predicted object heat map; the predicted key point information includes a predicted key point heat map; predicting the predicted attribute information of the target object in the sample image through the target detection network to be trained and predicting the predicted key point information of the target object through the key point detection network to be trained includes: predicting the predicted object heat map of the target object in the sample image through the target detection network to be trained; fusing the predicted object heat map and the feature map of the sample image to obtain a sample fusion feature map, and inputting the sample fusion feature map into the first convolutional network to be trained to output predicted key point feature parameters; based on the predicted key point feature parameters, predicting the key points of the target object from the feature map of the sample image, and generating a predicted key point heat map of the target object based on the predicted key points.

[0132] Among them, the predicted object heat map is a heat map predicted by the key point detection model to be trained based on the target object in the input sample image during the training of the key point detection model to be trained. The predicted key point heat map is a key point heat map predicted by the key point detection model to be trained based on the target object in the input sample image during the training of the key point detection model to be trained.

[0133] Specifically, the computer can predict the predicted object heat map of the target object in the sample image through the target detection network to be trained, and perform feature fusion on the predicted object heat map and the feature map of the sample image to obtain a sample fusion feature map. The computer device can input the obtained sample fusion feature map into the first convolutional network to be trained to perform convolution on the sample fusion feature map and output predicted key point feature parameters. The computer device can predict the key points of the target object from the feature map of the sample image based on the predicted key point feature parameters, and generate a predicted key point heat map of the target object based on the predicted key points.

[0134] In one embodiment, the predicted attribute information may include, in addition to the predicted object heatmap, information corresponding to any other attributes included in the target object. The predicted key point information may include, in addition to the predicted key point heatmap, other information that can characterize the key points of the target object.

[0135] In the above embodiment, through the target detection network to be trained, the predicted object heatmap of the target object in the sample image can be quickly predicted. By fusing the predicted object heatmap and the feature map of the sample image, a sample fusion feature map can be obtained. By inputting the sample fusion feature map into the first convolutional network to be trained, the predicted key point feature parameters associated with the input can be output. Furthermore, based on the predicted key point feature parameters associated with the input, the key points of the target object can be accurately predicted from the feature map of the sample image, and based on the predicted key points, the predicted key point heatmap of the target object can be quickly generated. In this way, through the first convolutional network, the training of target object detection and the training of key point detection of the target object can be combined to achieve multi-task joint training.

[0136] In one embodiment, the predicted object heatmap is obtained by performing heatmap coordinate transformation on the coordinates of the center point of the target object in the sample image predicted by the target detection network to be trained; the predicted attribute information further includes the predicted size information of the bounding box corresponding to the target object, and the conversion error corresponding to the center point of the target object; the conversion error is the error generated when performing heatmap coordinate transformation on the coordinates of the center point.

[0137] Among them, the bounding box is a graphic box that encloses the target object, for example, a rectangular box. The predicted size information is the size information of the bounding box predicted by the key point detection model to be trained.

[0138] Specifically, the computer device can predict the coordinates of the center point of the target object in the sample image based on the target detection network to be trained, and after obtaining the center point of the target object, perform heatmap coordinate transformation on the coordinates of the center point to obtain the predicted object heatmap. The computer device can determine the bounding box corresponding to the target object based on the key point detection model to be trained, and predict the size of the bounding box to obtain the predicted size information. The computer device can obtain the conversion error corresponding to the center point of the target object when performing heatmap coordinate transformation on the coordinates of the center point.

[0139] In one embodiment, the predicted size information of the bounding box corresponding to the target object may specifically include the predicted height information and predicted width information of the bounding box.

[0140] In one embodiment, such as Figure 3As shown in the figure, the key point detection model includes a feature extraction network 301, an object detection network 302, and a key point detection network 303. Among them, the feature extraction network 301 includes a backbone network, and the key point detection network 303 includes a first convolutional network 3031 and a second convolutional network 3032. The computer device can input the 320*320 image to be detected into the backbone network to obtain an 80*80 original feature map, perform 1*1 convolution on the original feature map to obtain a convolved feature map, perform upsampling on the original feature map to obtain an upsampled feature map, fuse the convolved feature map and the upsampled feature map to obtain a fused feature map, and perform 3*3 convolution on the fused feature map to obtain a first feature map of the 80*80 target object. The computer device can perform 3*3 convolution on the first feature map to obtain multiple intermediate feature maps. The computer device can input the intermediate feature maps into the object detection network 302, and perform 1*1 convolution on the multiple intermediate feature maps through the object detection network 302 to fuse the features of the same type of target objects into the same feature map, obtaining a second feature map corresponding to each type. The computer device can merge and input the first feature map and the second feature map into the first convolutional network 3031 to obtain key point feature parameters, and use the key point feature parameters as the convolutional parameters of the second convolutional network 3032, and perform convolution on the first feature map through the second convolutional network 3032 to detect the key points of the target object from the first feature map.

[0141] It should be noted that the computer device can detect the key points of the target object through the trained key point detection model. Among them, the trained key point detection model can be obtained by iteratively training the key point detection model to be trained. During the process of iteratively training the key point detection model, refer to Figure 3 , the computer device can predict the coordinates of the center point of the target object in the sample image based on the object detection network 302, and after obtaining the center point of the target object, perform heat map coordinate transformation on the coordinates of the center point to obtain the predicted object heat map. At the same time, the computer device can determine the bounding box corresponding to the target object based on the object detection network 302, and predict the width and height of the bounding box to obtain predicted width information and predicted height information. In addition, when the computer device performs heat map coordinate transformation on the coordinates of the above center point, it can obtain the transformation error corresponding to the center point of the target object. It can be understood that the width information, height information, and transformation error corresponding to the target object are used to assist in training the key point detection model, and in the actual application process of the key point detection model, it is not necessary to output the width information, height information, and transformation error corresponding to the target object. It is only necessary to output the second feature map corresponding to the target object, and merge the first feature map and the second feature map and input them into the key point detection network to detect the key points of the target object.

[0142] In one embodiment, as Figure 4 shown, the sample image 401 includes a target object (i.e., a hand), and the computer device can predict the coordinates of the center point of the target object (i.e., the center point of the hand) in the sample image based on the key point detection model to be trained. Furthermore, after obtaining the center point of the target object, the computer device can perform heat map coordinate transformation on the coordinates of this center point to obtain the corresponding predicted object heat map 402. It can be understood that the white dot in the predicted object heat map 402 represents the center point of the target object.

[0143] In one embodiment, if there are three target objects in the sample image, then, as Figure 5 shown, the computer device can predict the coordinates of the center points of these three target objects in the sample image through the key point detection model to be trained. Furthermore, after obtaining the coordinates of the center points of these three target objects, the computer device can perform heat map coordinate transformation on the coordinates of these three center points to obtain the corresponding predicted object heat map. It can be understood that the three white dots in this predicted object heat map represent the center points of these three target objects.

[0144] In one embodiment, as Figure 6 shown, the sample image 601 includes a target object (i.e., a hand), and the computer device can predict the coordinates of the key points of the target object in the sample image based on the key point detection model to be trained. After obtaining the coordinates of the key points of the target object, the computer device can perform heat map coordinate transformation on the coordinates of this key point to obtain the corresponding predicted key point heat map 602. It can be understood that the white dot in the predicted key point heat map 602 represents the key point of the target object.

[0145] In one embodiment, if there are two key points in the sample image, then, as Figure 7 shown, the computer device can predict the coordinates of these two key points in the sample image through the key point detection model to be trained. After obtaining the coordinates of these two key points, the computer device can perform heat map coordinate transformation on the coordinates of these two key points to obtain the corresponding predicted key point heat map. It can be understood that the two white dots in this predicted key point heat map represent the two key points in the sample image.

[0146] In the above embodiments, the predicted size information and conversion error in the predicted attribute information can play a very good auxiliary role in the training process of target object detection, further improving the accuracy of target object detection.

[0147] In one embodiment, based on the key-point feature parameters, detecting the key points of the target object from the first feature map includes: performing convolution on the first feature map according to the key-point feature parameters to obtain a second probability feature map; each pixel point in the second probability feature map corresponds to a second probability value respectively; the second probability value is used to represent the probability that there is a key point at the position corresponding to the pixel point; dividing the second probability feature map into a preset number of second image blocks with the same size; for each second image block, selecting the second probability value with the largest probability value from the second image block as the second target probability value; determining the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value as the second target pixel points; and taking the second target pixel points as the key points of the target object.

[0148] Wherein, the second image block is an image block obtained by dividing the second probability feature map. The second target probability value is the second probability value used as the target. The second target pixel point is the pixel point where there is actually a key point of the target object at the corresponding position.

[0149] Specifically, the computer device can perform convolution on the first feature map according to the key-point feature parameters to obtain a second probability feature map, and divide the second probability feature map into a preset number of second image blocks with the same size. For each second image block, the computer device can select the second probability value with the largest probability value from the second image block as the second target probability value. The computer device can compare the second target probability value with the second preset probability value, and determine the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value as the second target pixel points. Furthermore, the computer device can take the second target pixel points as the key points of the target object.

[0150] In the above embodiment, by performing convolution on the first feature map according to the key-point feature parameters, a second probability feature map can be obtained. Divide the second probability feature map into multiple second image blocks, select the second probability value with the largest probability value from each second image block as the second target probability value, and then determine the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value as the second target pixel points. It can be understood that the position where the second target pixel points are located is the position where the key points of the target object are located. Furthermore, the second target pixel points can be directly used as the key points of the target object, further improving the accuracy of detecting the key points of the target object.

[0151] In one embodiment, the image to be detected is an image collected in a point-reading scenario; the target object is an input entity used to trigger point-reading in the point-reading scenario. The above method further includes: determining the target point-reading text based on the key points of the input entity; and performing point-reading processing based on the target point-reading text.

[0152] Among them, the input entity is an entity object used to trigger point reading. For example, the hand of the point reader, an ordinary pen for doing homework, or a dedicated point reading pen for point reading, etc. The target point reading text is the point reading text as the target.

[0153] Specifically, the computer device can determine the target point reading text that needs to be point read based on the key points of the input entity. Furthermore, the computer device can perform point reading processing based on the target point reading text.

[0154] In one embodiment, the point reading processing can specifically be to perform text recognition on the target point reading text and return the description information of the target point reading text. Among them, the description information is the information used to describe the target point reading text.

[0155] For example, if the target point reading text is an English word, the description information can include at least one of the pronunciation, Chinese translation, part of speech, singular and plural forms, and application examples of the English word.

[0156] In the above embodiment, based on the key points of the input entity, the target point reading text can be quickly determined, and then point reading processing can be performed based on the target point reading text, improving the point reading accuracy in the point reading scenario.

[0157] In one embodiment, there are multiple input entities, and different types of input entities are included among the multiple input entities. Determining the target point reading text based on the key points of the input entity includes: taking the input entity corresponding to the highest priority among the types corresponding to each type in the different types of input entities as the target input entity; determining the key points of the target input entity as the target key points, and determining the target point reading text pointed to by the target key points.

[0158] Among them, the target input entity is the input entity as the target. The target key points are the key points as the target.

[0159] Specifically, for each input entity, the computer device can pre-determine the priority corresponding to the type of the input entity based on the type of the input entity. Furthermore, the computer device can take the input entity corresponding to the highest priority among the types corresponding to each type in the different types of input entities as the target input entity. The computer device can determine the key points of the target input entity as the target key points and determine the target point reading text pointed to by the target key points.

[0160] In one embodiment, in the point reading scenario, the input entities in the image to be detected include a hand and a pen, such as Figure 8As shown, the computer device can detect the key points of all input entities (i.e., including the fingertips 801 and 802 of the hand and the tip 803 of the pen) through the trained key point detection model. If both types of input entities, the hand and the pen, are included in the image to be detected, and if the type with the highest priority is preset as the pen, then as Figure 9 shown, the computer device can finally determine the key points of the pen as the target key points 901.

[0161] In one embodiment, in the point reading scenario, as Figure 10 shown, if the target point reading text pointed to by the key points of the pen 1002 is "man", then the computer device 1001 can perform point reading processing based on the target point reading text "man" and return and display the description information 1003 of the target point reading text "man". As Figure 11 shown, if the target point reading text pointed to by the key points of the hand 1102 is "the", then the computer device 1101 can perform point reading processing based on the target point reading text "the" and return and display the description information 1103 of the target point reading text "the".

[0162] In the above embodiment, according to the priorities corresponding to each type among different types of input entities, the input entity corresponding to the type with the highest priority can be used as the target input entity, and then the key points of the target input entity can be determined as the target key points. In this way, the target point reading text pointed to by the target key points can be determined quickly and accurately, and further, the point reading accuracy in the point reading scenario is improved. At the same time, since point reading operations can be performed on each type of input entity, an uninterrupted interaction method is provided in the point reading scenario.

[0163] In one embodiment, as Figure 12 shown, a key point detection method is provided, which specifically includes the following steps:

[0164] Step 1202, obtain a sample image containing input entities and input the sample image into the key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained.

[0165] Step 1204, predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained.

[0166] In one embodiment, the key point detection network to be trained includes a first convolutional network to be trained; the predicted attribute information includes a predicted object heat map; and the predicted key point information includes a predicted key point heat map. The computer device can predict the predicted object heat map of the input entity in the sample image through the target detection network to be trained; fuse the predicted object heat map and the feature map of the sample image to obtain a sample fused feature map, and input the sample fused feature map into the first convolutional network to be trained to output predicted key point feature parameters; based on the predicted key point feature parameters, predict the key points of the input entity from the feature map of the sample image, and generate a predicted key point heat map of the input entity based on the predicted key points.

[0167] In one embodiment, the predicted object heat map is obtained by performing heat map coordinate transformation on the coordinates of the center point of the input entity predicted in the sample image by the target detection network to be trained; the predicted attribute information further includes the predicted size information of the bounding box corresponding to the input entity and the conversion error corresponding to the center point of the input entity; the conversion error is the error generated when performing heat map coordinate transformation on the coordinates of the center point.

[0168] Step 1206: Determine a first loss value between the predicted attribute information and the target attribute information of the input entity, determine a second loss value between the predicted key point information and the target key point information of the input entity, and determine a target loss value according to the first loss value and the second loss value.

[0169] Step 1208: Iteratively train the key point detection model to be trained in the direction of reducing the target loss value until the iteration stop condition is met, and obtain the trained key point detection model.

[0170] Step 1210: Obtain the original feature map of the image to be detected; the image to be detected is an image collected in a point reading scenario, perform convolution on the original feature map to obtain a convolved feature map, and perform upsampling on the original feature map to obtain an upsampled feature map.

[0171] Step 1212: Fuse the convolved feature map and the upsampled feature map to obtain a fused feature map, perform convolution on the fused feature map to obtain the first feature map of the image to be detected, and perform convolution on the first feature map to obtain an intermediate feature map.

[0172] Step 1214: Input the intermediate feature map into the target detection network in the trained key point detection model, perform input entity detection processing on the intermediate feature map to obtain a first probability feature map; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network.

[0173] Step 1216: Divide the first probability feature map into a preset number of first image patches of the same size; for each first image patch, select the first probability value with the largest probability value from the first image patch as the first target probability value.

[0174] Step 1218: Determine the pixel points corresponding to the probability values whose first target probability values are greater than the first preset probability value as the first target pixel points.

[0175] Step 1220: Generate a second feature map of the input entity according to the first target pixel points, and fuse the first feature map and the second feature map to obtain a fused feature map.

[0176] Step 1222: Input the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the input entity.

[0177] Step 1224: Use the key point feature parameters as the convolution parameters of the second convolutional network, and perform convolution on the first feature map through the second convolutional network to obtain a second probability feature map.

[0178] Step 1226: Divide the second probability feature map into a preset number of second image patches of the same size; for each second image patch, select the second probability value with the largest probability value from the second image patch as the second target probability value.

[0179] Step 1228: Determine the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value as the second target pixel points.

[0180] Step 1230: Use the second target pixel points as the key points of the input entity, and according to the priorities corresponding to each type in different types of input entities, use the input entity corresponding to the highest priority type as the target input entity.

[0181] Step 1232: Determine the key points of the target input entity as the target key points, determine the target point reading text pointed to by the target key points, and perform point reading processing based on the target point reading text.

[0182] The present application also provides an application scenario, which applies the above key point detection method. Specifically, the key point detection method can be applied to the key point detection scenario in the point reading service. The computer device can obtain a sample image containing the input entity, and input the sample image into the key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained. The predicted attribute information of the target object in the sample image is predicted through the target detection network to be trained, and the predicted key point information of the target object is predicted through the key point detection network to be trained. Determine the first loss value between the predicted attribute information and the target attribute information of the input entity. Determine the second loss value between the predicted key point information and the target key point information of the input entity; determine the target loss value according to the first loss value and the second loss value. Iteratively train the key point detection model to be trained in the direction of reducing the target loss value until the iterative stop condition is met, and obtain the trained key point detection model.

[0183] The computer device can obtain the original feature map of the image to be detected; the image to be detected is an image collected in the point reading scenario, and the original feature map is convolved to obtain the convolved feature map. The original feature map is upsampled to obtain the upsampled feature map. The convolved feature map and the upsampled feature map are fused to obtain the fused feature map. The fused feature map is convolved to obtain the first feature map of the image to be detected, and the first feature map is convolved to obtain the intermediate feature map.

[0184] The computer device can input the intermediate feature map into the target detection network in the trained key point detection model, and perform input entity detection processing on the intermediate feature map to obtain the first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value respectively; the first probability value is used to represent the probability that the input entity exists at the position of the corresponding pixel point; the trained key point detection model also includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network; the input entity is an input entity used to trigger point reading in the point reading scenario. The first probability feature map is divided into a preset number of first image blocks with the same size; for each first image block, the first probability value with the largest probability value is selected from the first image block as the first target probability value. The pixel points corresponding to the probability values whose first target probability values are greater than the first preset probability value are determined as the first target pixel points, and the second feature map of the input entity is generated according to the first target pixel points.

[0185] The computer device can fuse the first feature map and the second feature map to obtain a fused feature map, and input the fused feature map into a first convolutional network for convolution to output the key point feature parameters of the input entity. The key point feature parameters are used as the convolution parameters of the second convolutional network, and the first feature map is convolved through the second convolutional network to obtain a second probability feature map; each pixel point in the second probability feature map corresponds to a second probability value; the second probability value is used to represent the probability that there is a key point at the position of the corresponding pixel point. The second probability feature map is divided into a preset number of second image blocks with the same size; for each second image block, the second probability value with the largest probability value is selected from the second image block as the second target probability value. The pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value are determined as the second target pixel points. The second target pixel points are used as the key points of the input entity.

[0186] The computer device can, according to the priorities corresponding to each type in different types of input entities, use the input entity corresponding to the highest priority type as the target input entity. The key points of the target input entity are determined as the target key points, the target point reading text pointed to by the target key points is determined, and point reading processing is performed based on the target point reading text.

[0187] This application also provides another application scenario, which applies the above key point detection method. Specifically, the key point detection method can be applied to the key point detection scenario of a human face in the human face recognition process. The computer device can perform feature extraction processing on the image to be detected to obtain the first feature map of the image to be detected, and perform target human face detection processing on the first feature map to obtain the second feature map of the target human face. The first feature map and the second feature map are fused to obtain a fused feature map. Based on the fused feature map, the key point feature parameters of the target human face are determined, and based on the key point feature parameters, the key points of the target human face are detected from the first feature map.

[0188] It should be understood that although the steps in the flowcharts of the above embodiments are shown in sequence, these steps are not necessarily executed in sequence. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0189] In one embodiment, as Figure 13As shown, a key point detection device 1300 is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes:

[0190] An extraction module 1301, configured to perform feature extraction processing on the image to be detected, and obtain a first feature map of the image to be detected.

[0191] A detection module 1302, configured to perform target object detection processing on the first feature map, and obtain a second feature map of the target object.

[0192] A fusion module 1303, configured to fuse the first feature map and the second feature map, and obtain a fused feature map.

[0193] A determination module 1304, configured to determine key point feature parameters of the target object based on the fused feature map.

[0194] The detection module 1302 is further configured to detect key points of the target object from the first feature map based on the key point feature parameters.

[0195] In one embodiment, the extraction module 1301 is further configured to obtain an original feature map of the image to be detected; perform convolution on the original feature map to obtain a convolved feature map; perform upsampling on the original feature map to obtain an upsampled feature map; fuse the convolved feature map and the upsampled feature map to obtain a fused feature map; and perform convolution on the fused feature map to obtain the first feature map of the image to be detected.

[0196] In one embodiment, there are multiple target objects, and different types of target objects are included among the multiple target objects; the detection module 1302 is further configured to perform convolution on the first feature map to obtain multiple intermediate feature maps; perform convolution on the multiple intermediate feature maps to fuse the features of the target objects of the same type into the same feature map, and obtain second feature maps corresponding to each type respectively.

[0197] In one embodiment, the detection module 1302 is further configured to perform target object detection processing on the first feature map to obtain a first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value respectively; the first probability value is used to represent the probability that a target object exists at the position corresponding to the pixel point; divide the first probability feature map into a preset number of first image blocks with the same size; for each first image block, select the largest first probability value from the first image block as the first target probability value; determine the pixel points corresponding to the probability values greater than the first preset probability value as the first target pixel points; and generate a second feature map of the target object according to the first target pixel points.

[0198] In one embodiment, the second feature map is generated by a target detection network in a trained key point detection model; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network; the determination module 1304 is further configured to input the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the target object; the detection module 1302 is further configured to use the key point feature parameters as the convolutional parameters of the second convolutional network, and perform convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

[0199] In one embodiment, the detection module 1302 is further configured to use the key point feature parameters as the convolutional parameters of the second convolutional network, so that the second convolutional network determines a target region in the first feature map based on the key point feature parameters; the target region is the region of the key points of the target object in the first feature map; the key points of the target object are detected from the target region based on the second convolutional network.

[0200] In one embodiment, the apparatus further includes: a training module, configured to obtain a sample image containing a target object; input the sample image into the key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained; predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained; determine a first loss value between the predicted attribute information and the target attribute information of the target object; determine a second loss value between the predicted key point information and the target key point information of the target object; determine a target loss value according to the first loss value and the second loss value; perform iterative training on the key point detection model to be trained in a direction that reduces the target loss value until the iterative stop condition is met, and obtain the trained key point detection model.

[0201] In one embodiment, the key point detection network to be trained includes a first convolutional network to be trained; the predicted attribute information includes a predicted object heat map; the predicted key point information includes a predicted key point heat map; the training module is further configured to predict the predicted object heat map of the target object in the sample image through the target detection network to be trained; fuse the predicted object heat map and the feature map of the sample image to obtain a sample fused feature map, and input the sample fused feature map into the first convolutional network to be trained to output predicted key point feature parameters; predict the key points of the target object from the feature map of the sample image based on the predicted key point feature parameters, and generate a predicted key point heat map of the target object based on the predicted key points.

[0202] In one embodiment, the predicted object heatmap is obtained by performing heatmap coordinate transformation on the coordinates of the center point of the target object in the sample image predicted by the target detection network to be trained; the predicted attribute information further includes the predicted size information of the bounding box corresponding to the target object and the conversion error corresponding to the center point of the target object; the conversion error is the error generated when performing heatmap coordinate transformation on the coordinates of the center point.

[0203] In one embodiment, the detection module 1302 is further configured to perform convolution on the first feature map according to the key point feature parameters to obtain a second probability feature map; each pixel point in the second probability feature map corresponds to a second probability value respectively; the second probability value is used to represent the probability that there is a key point at the position of the corresponding pixel point; the second probability feature map is divided into a preset number of second image blocks of the same size; for each second image block, the second probability value with the largest probability value is selected from the second image block as the second target probability value; the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value are determined as the second target pixel points; the second target pixel points are used as the key points of the target object.

[0204] In one embodiment, the image to be detected is an image collected in a point reading scenario; the target object is an input entity used to trigger point reading in the point reading scenario; the apparatus further includes: a point reading module, configured to determine the target point reading text based on the key points of the input entity; and perform point reading processing based on the target point reading text.

[0205] In one embodiment, there are multiple input entities, and different types of input entities are included in the multiple input entities; the point reading module is further configured to use the input entity corresponding to the highest priority among the different types of input entities as the target input entity according to the priorities respectively corresponding to each type in the different types of input entities; determine the key points of the target input entity as the target key points, and determine the target point reading text pointed to by the target key points.

[0206] Reference Figure 14 , in one embodiment, the key point detection apparatus 1300 further includes a training module 1305 and a point reading module 1306.

[0207] The above key point detection device can obtain the first feature map of the image to be detected by performing feature extraction processing on the image to be detected. By performing target object detection processing on the first feature map, a second feature map of the target object including the overall information of the target object can be obtained. By fusing the first feature map and the second feature map, a fused feature map can be obtained. Based on the fused feature map, the key point feature parameters of the target object are determined. Since the image to be detected is variable, the obtained key point feature parameters will also change dynamically with the image to be detected. Furthermore, based on the key point feature parameters, the key points of the target object can be directly detected from the first feature map, avoiding the step of associating the key points with their corresponding target objects, and improving the accuracy of key point detection of the target object.

[0208] For the specific limitations of the key point detection device, reference can be made to the limitations of the key point detection method in the above text, which will not be elaborated here. Each module in the above key point detection device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0209] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, a key point detection method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0210] Those skilled in the art can understand, Figure 15The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0211] In one embodiment, a computer device is also provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0212] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0213] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0214] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0216] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0217] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A key point detection method, characterized in that, The method includes: Performing feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected; Performing target object detection processing on the first feature map through a target detection network in a trained key point detection model; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network; Fusing the first feature map and the second feature map to obtain a fused feature map; Inputting the fused feature map into the first convolutional network for convolution to output key point feature parameters of the target object; Using the key point feature parameters as convolutional parameters of the second convolutional network, and performing convolution on the first feature map through the second convolutional network to detect key points of the target object from the first feature map.

2. The method according to claim 1, wherein The performing feature extraction processing on the image to be detected to obtain a first feature map of the image to be detected includes: Obtaining an original feature map of the image to be detected; Performing convolution on the original feature map to obtain a convolved feature map; Performing upsampling on the original feature map to obtain an upsampled feature map; Fusing the convolved feature map and the upsampled feature map to obtain a fused feature map; Performing convolution on the fused feature map to obtain the first feature map of the image to be detected.

3. The method according to claim 1, characterized in that, There are multiple target objects, and different types of target objects are included among the multiple target objects; the performing target object detection processing on the first feature map through a target detection network in a trained key point detection model to obtain a second feature map of the target object includes: Performing convolution on the first feature map to obtain multiple intermediate feature maps; Performing convolution on the multiple intermediate feature maps to fuse the features of target objects of the same type into the same feature map to obtain a second feature map corresponding to each type respectively.

4. The method according to claim 1, wherein The performing target object detection processing on the first feature map through a target detection network in a trained key point detection model to obtain a second feature map of the target object includes: Performing target object detection processing on the first feature map to obtain a first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value respectively; the first probability value is used to represent the probability that a target object exists at the position of the corresponding pixel point; Dividing the first probability feature map into a preset number of first image blocks with the same size; for each of the first image blocks, selecting the first probability value with the largest probability value from the first image block as a first target probability value; Determining pixel points corresponding to probability values greater than a first preset probability value among the first target probability values as first target pixel points; Generating a second feature map of the target object according to the first target pixel points.

5. The method according to claim 1, wherein The using the key point feature parameters as convolutional parameters of the second convolutional network, and performing convolution on the first feature map through the second convolutional network to detect key points of the target object from the first feature map includes: Use the key point feature parameters as the convolution parameters of the second convolutional network, so that the second convolutional network determines a target region in the first feature map based on the key point feature parameters; the target region is the region of the key points of the target object in the first feature map; Detect the key points of the target object from the target region based on the second convolutional network.

6. The method according to claim 1, wherein The steps of obtaining the trained key point detection model include: Obtain a sample image containing a target object; Input the sample image into the key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained; Predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained; Determine a first loss value between the predicted attribute information and the target attribute information of the target object; Determine a second loss value between the predicted key point information and the target key point information of the target object; Determine a target loss value according to the first loss value and the second loss value; Iteratively train the key point detection model to be trained in a direction that reduces the target loss value until the iteration stop condition is met, and obtain the trained key point detection model.

7. The method according to claim 6, wherein The key point detection network to be trained includes a first convolutional network to be trained; the predicted attribute information includes a predicted object heat map; the predicted key point information includes a predicted key point heat map; The step of predicting the predicted attribute information of the target object in the sample image through the target detection network to be trained and predicting the predicted key point information of the target object through the key point detection network to be trained includes: Predict the predicted object heat map of the target object in the sample image through the target detection network to be trained; Fuse the predicted object heat map and the feature map of the sample image to obtain a sample fusion feature map, and input the sample fusion feature map into the first convolutional network to be trained to output key point feature parameters; Predict the key points of the target object from the feature map of the sample image based on the key point feature parameters, and generate a predicted key point heat map of the target object based on the predicted key points.

8. The method according to claim 7, wherein The predicted object heat map is obtained by performing heat map coordinate transformation on the coordinates of the center point of the target object predicted in the sample image by the target detection network to be trained; the predicted attribute information further includes the predicted size information of the bounding box corresponding to the target object, and the conversion error corresponding to the center point of the target object; the conversion error is the error generated when performing heat map coordinate transformation on the coordinates of the center point.

9. The method according to claim 1, wherein The step of using the key point feature parameters as the convolution parameters of the second convolutional network and performing convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map includes: Convolve the first feature map according to the key point feature parameters to obtain a second probability feature map; each pixel point in the second probability feature map corresponds to a second probability value respectively; the second probability value is used to represent the probability that a key point exists at the position of the corresponding pixel point. Divide the second probability feature map into a preset number of second image blocks with the same size; for each of the second image blocks, select the second probability value with the largest probability value from the second image block as the second target probability value. Determine the pixel points corresponding to the probability values whose second target probability values are greater than the second preset probability value as the second target pixel points. Use the second target pixel points as the key points of the target object.

10. The method according to any one of claims 1 to 9, characterized in that The image to be detected is an image collected in a point reading scenario. The target object is an input entity used to trigger point reading in the point reading scenario; the method further includes: Determine the target point reading text based on the key points of the input entity. Perform point reading processing based on the target point reading text.

11. The method according to claim 10, wherein There are multiple input entities, and different types of input entities are included in the multiple input entities; determining the target point reading text based on the key points of the input entity includes: According to the priorities corresponding to each type in the different types of input entities, use the input entity corresponding to the highest priority type as the target input entity. Determine the key points of the target input entity as the target key points. Determine the target point reading text pointed to by the target key points.

12. A key point detection device, characterized in that, The device includes: An extraction module, configured to perform feature extraction processing on the image to be detected to obtain the first feature map of the image to be detected. A detection module, configured to perform target object detection processing on the first feature map through a target detection network in a trained key point detection model to obtain a second feature map of the target object; the trained key point detection model further includes a key point detection network; the key point detection network includes a first convolutional network and a second convolutional network. A fusion module, configured to fuse the first feature map and the second feature map to obtain a fused feature map. A determination module, configured to input the fused feature map into the first convolutional network for convolution to output the key point feature parameters of the target object. The detection module is further configured to use the key point feature parameters as the convolution parameters of the second convolutional network, and perform convolution on the first feature map through the second convolutional network to detect the key points of the target object from the first feature map.

13. The key point detection device according to claim 12, characterized in that, The extraction module is further configured to obtain the original feature map of the image to be detected; perform convolution on the original feature map to obtain a convolved feature map; perform upsampling on the original feature map to obtain an upsampled feature map; fuse the convolved feature map and the upsampled feature map to obtain a fused feature map; perform convolution on the fused feature map to obtain the first feature map of the image to be detected.

14. The key point detection device according to claim 12, wherein The target objects are multiple, and different types of target objects are included among the multiple target objects; the detection module is further configured to perform convolution on the first feature map to obtain multiple intermediate feature maps; perform convolution on the multiple intermediate feature maps to fuse the features of the target objects of the same type into the same feature map, so as to obtain a second feature map corresponding to each type.

15. The key point detection device according to claim 12, characterized in that, The detection module is further configured to perform target object detection processing on the first feature map to obtain a first probability feature map; each pixel point in the first probability feature map corresponds to a first probability value respectively; the first probability value is used to represent the probability that a target object exists at the position corresponding to the pixel point; divide the first probability feature map into a preset number of first image blocks with the same size; for each of the first image blocks, select the first probability value with the largest probability value from the first image block as the first target probability value; determine the pixel points corresponding to the probability values greater than the first preset probability value of the first target probability value as the first target pixel points; generate the second feature map of the target object according to the first target pixel points.

16. The key point detection device according to claim 12, wherein The detection module is further configured to use the key point feature parameters as the convolution parameters of the second convolutional network, so that the second convolutional network determines a target region in the first feature map based on the key point feature parameters; the target region is the region of the key points of the target object in the first feature map; Based on the second convolutional network, the key points of the target object are detected from the target region.

17. The key point detection device according to claim 12, wherein, The device further includes a training module, and the training module is configured to obtain a sample image containing a target object; input the sample image into a key point detection model to be trained; the key point detection model to be trained includes a target detection network to be trained and a key point detection network to be trained; Predict the predicted attribute information of the target object in the sample image through the target detection network to be trained, and predict the predicted key point information of the target object through the key point detection network to be trained; Determine a first loss value between the predicted attribute information and the target attribute information of the target object; Determine a second loss value between the predicted key point information and the target key point information of the target object; determine a target loss value according to the first loss value and the second loss value; perform iterative training on the key point detection model to be trained in the direction of reducing the target loss value until the iterative stop condition is met, and obtain the trained key point detection model.

18. The key point detection device according to claim 17, characterized in that, The to-be-trained key point detection network includes a to-be-trained first convolutional network; the predicted attribute information includes a predicted object heat map; the predicted key point information includes a predicted key point heat map; the training module is further configured to, through the to-be-trained target detection network, predict a predicted object heat map of a target object in the sample image; fuse the predicted object heat map and the feature map of the sample image to obtain a sample fused feature map, and input the sample fused feature map into the to-be-trained first convolutional network to output predicted key point feature parameters; based on the predicted key point feature parameters, predict key points of the target object from the feature map of the sample image, and generate a predicted key point heat map of the target object based on the predicted key points.

19. The key point detection device according to claim 18, characterized in that, The predicted object heat map is obtained by performing heat map coordinate transformation on the coordinates of the center point of the target object predicted in the sample image by the to-be-trained target detection network; the predicted attribute information further includes predicted size information of the bounding box corresponding to the target object, and conversion error corresponding to the center point of the target object; the conversion error is the error generated when performing heat map coordinate transformation on the coordinates of the center point.

20. The key point detection device according to claim 12, wherein The detection module is further configured to perform convolution on the first feature map according to the key point feature parameters to obtain a second probability feature map; each pixel point in the second probability feature map respectively corresponds to a second probability value; the second probability value is used to represent the probability that there is a key point at the position of the corresponding pixel point; divide the second probability feature map into a preset number of second image blocks with the same size; for each of the second image blocks, select the second probability value with the largest probability value from the second image block as the second target probability value; determine the pixel points corresponding to the probability values where the second target probability value is greater than the second preset probability value as second target pixel points; Use the second target pixel points as the key points of the target object.

21. The key point detection device according to any one of claims 12 to 20, characterized in that The to-be-detected image is an image collected in a point reading scenario; The target object is an input entity used to trigger point reading in the point reading scenario; the device further includes: a point reading module, configured to determine a target point reading text based on the key points of the input entity. Perform point reading processing based on the target point reading text.

22. The key point detection device according to claim 21, wherein There are multiple input entities, and the multiple input entities include different types of input entities; the point reading module is further configured to, according to the priorities respectively corresponding to each type in the different types of input entities, use the input entity corresponding to the highest priority type as the target input entity; determine the key points of the target input entity as the target key points, and determine the target point reading text pointed to by the target key points.

23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 11.

24. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 11.

25. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by the processor, it implements the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Face attribute recognition method and device

    CN111144369A

  • Image detection method, device, equipment, storage medium and computer program product

    CN112597837A