Target Object Recognition Method, Device, Computer Equipment, and Storage Medium

Through object detection and human detection combined with image feature extraction, non-maximum suppression algorithm and interleaving ratio screening, the accuracy of fixed position recognition of the shooting equipment function function of the shooting equipment is solved, and the accurate recognition of the fixed position of the shooting equipment is achieved.

CN114170417BActive Publication Date: 2025-07-08ARASHI VISION INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111342860.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-07-08
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

In the prior art, a shooting device with the function of an invisible auxiliary device hides the auxiliary device in video or image data, resulting in the inability to accurately identify the fixed position of the shooting device.

Method used

By obtaining the image to be identified, object detection and human body detection are performed, image feature extraction, non-maximum suppression algorithm and interleaving ratio screening, the object detection box that meets the requirements is selected, and the object detection results are updated to identify the fixed position of the shooting device.

Benefits of technology

Accurate identification of fixed positions of the shooting equipment is achieved, reducing the occurrence of error detection and improving the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170417B_ABST
    Figure CN114170417B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, storage medium, and computer program product for identifying a target object. The method includes: obtaining an image to be recognized; performing object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized, where the object detection result includes a detection category and a detection box; when there is an object detection category in the object detection result, determining an object detection box corresponding to the object detection category, and performing human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized; screening the object detection box according to the human detection box, and updating the object detection result according to the screening result; and obtaining a target object recognition result according to the updated object detection result. By using this method, accurate recognition of the target object can be achieved on the basis of screening out the object detection boxes that may cause false detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a method, apparatus, computer device, and storage medium for target object recognition. Background Art

[0002] With the development of image processing technology, for video or image data captured by a shooting device, target detection can be used to identify the fixed position of the shooting device.

[0003] In traditional technologies, the commonly used recognition method is to determine the fixed position of the shooting device by recognizing auxiliary devices used to assist the shooting device in the video or image data. For example, the auxiliary device can specifically refer to a selfie stick used to assist the shooting device.

[0004] However, for a shooting device with an invisible auxiliary device function, since it hides the auxiliary device in the captured video or image data through physical means, it will cause the fixed position of the shooting device to be unable to be accurately recognized in the video or image data. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for target object recognition that can accurately recognize the fixed position of a shooting device.

[0006] In a first aspect, this application provides a method for target object recognition. The method includes:

[0007] Obtain an image to be recognized;

[0008] Perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized, where the target detection result includes a detection category and a detection box;

[0009] When there is a target detection category in the target detection result, determine the target detection box corresponding to the target detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0010] Screen the target detection box according to the human detection box, and update the target detection result according to the screening result;

[0011] Obtain a target object recognition result according to the updated target detection result.

[0012] In one embodiment, performing target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized includes:

[0013] Extract image feature information corresponding to the image to be recognized by performing image feature extraction on the image to be recognized;

[0014] Perform object detection based on the image feature information to obtain a first detection box corresponding to the image to be recognized;

[0015] Filter the first detection box to obtain an object detection result corresponding to the image to be recognized.

[0016] In one embodiment, filtering the first detection box to obtain an object detection result corresponding to the image to be recognized includes:

[0017] Obtain the confidence of the first detection box corresponding to the first detection box;

[0018] Filter the first detection box according to the confidence of the first detection box and a preset confidence threshold to obtain a second detection box;

[0019] Use the non-maximum suppression algorithm to filter the second detection box to obtain candidate detection boxes;

[0020] Obtain an object detection result corresponding to the image to be recognized according to the candidate detection boxes.

[0021] In one embodiment, using the non-maximum suppression algorithm to filter the second detection box to obtain candidate detection boxes includes:

[0022] Classify the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category;

[0023] Use the non-maximum suppression algorithm to filter the detection boxes of the same category in the set of detection boxes of the same category to obtain candidate detection boxes.

[0024] In one embodiment, filtering the object detection box according to the human body detection box and updating the object detection result according to the filtering result includes:

[0025] Determine the intersection-over-union ratio between the human body detection box and the object detection box;

[0026] Filter the object detection box according to the intersection-over-union ratio and a preset intersection-over-union ratio threshold to obtain a third detection box, and the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union ratio threshold;

[0027] Update the object detection result according to the third detection box.

[0028] In one embodiment, obtaining an object recognition result according to the updated object detection result includes:

[0029] Determine the confidence of the second detection box corresponding to the detection box in the updated object detection result;

[0030] Sort the confidence levels of the second detection boxes to determine the detection box corresponding to the highest detection box confidence level;

[0031] Obtain the target object recognition result based on the detection box corresponding to the highest detection box confidence level.

[0032] In a second aspect, the present application also provides a target object recognition device. The device includes:

[0033] An acquisition module, configured to acquire an image to be recognized;

[0034] A first detection module, configured to perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized, where the target detection result includes a detection category and a detection box;

[0035] A second detection module, configured to determine a target detection box corresponding to the target detection category when there is a target detection category in the target detection result, and perform human body detection on the image to be recognized to obtain a human body detection box corresponding to the image to be recognized;

[0036] A screening module, configured to screen the target detection boxes according to the human body detection boxes, and update the target detection result according to the screening result;

[0037] A processing module, configured to obtain the target object recognition result according to the updated target detection result.

[0038] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0039] Acquire an image to be recognized;

[0040] Perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized, where the target detection result includes a detection category and a detection box;

[0041] When there is a target detection category in the target detection result, determine a target detection box corresponding to the target detection category, and perform human body detection on the image to be recognized to obtain a human body detection box corresponding to the image to be recognized;

[0042] Screen the target detection boxes according to the human body detection boxes, and update the target detection result according to the screening result;

[0043] Obtain the target object recognition result according to the updated target detection result.

[0044] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0045] Obtain the image to be recognized;

[0046] Perform object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized, where the object detection result includes a detection category and a detection box;

[0047] When there is an object detection category in the object detection result, determine the object detection box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0048] Screen the object detection box according to the human detection box, and update the object detection result according to the screening result;

[0049] Obtain an object recognition result according to the updated object detection result.

[0050] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0051] Obtain the image to be recognized;

[0052] Perform object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized, where the object detection result includes a detection category and a detection box;

[0053] When there is an object detection category in the object detection result, determine the object detection box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0054] Screen the object detection box according to the human detection box, and update the object detection result according to the screening result;

[0055] Obtain an object recognition result according to the updated object detection result.

[0056] The above object recognition method, device, computer device, storage medium, and computer program product obtain the image to be recognized, perform object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized. When there is an object detection category in the object detection result, determine the object detection box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized, which can use the human detection box to screen the object detection box, screen out the object detection boxes that meet the requirements, so as to reduce the occurrence of false detections. By updating the object detection result according to the screening result and obtaining the object recognition result according to the updated object detection result, accurate recognition of the object can be achieved on the basis of screening out the object detection boxes that will cause false detections. Brief Description of the Drawings

[0057] Figure 1 It is a schematic flowchart of a target object recognition method in an embodiment;

[0058] Figure 2 It is a schematic diagram of a trained target detection model in an embodiment;

[0059] Figure 3 It is a schematic diagram of a residual network in an embodiment;

[0060] Figure 4 It is a schematic diagram of an FPN (feature pyramid networks) network in an embodiment;

[0061] Figure 5 It is a schematic flowchart of a target object recognition method in another embodiment;

[0062] Figure 6 It is a structural block diagram of a target object recognition device in an embodiment;

[0063] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Embodiment

[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0065] In one embodiment, as Figure 1 shown, a target object recognition method is provided. In this embodiment, it is exemplified that the method is applied to a server. It can be understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, portable wearable devices, action cameras, panoramic cameras or pan-tilt cameras, etc. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0066] Step 102, obtain the image to be recognized.

[0067] Among them, the image to be recognized refers to the image captured by the imaging device and requires determining the target object. The target object refers to the fixed position of the imaging device and / or the photographer and / or the auxiliary imaging device in the image to be recognized. Among them, the imaging device includes but is not limited to single-lens reflex cameras, mirrorless cameras, mobile phones with photo-taking and video-recording functions, action cameras, panoramic cameras, gimbal cameras, drones, etc.

[0068] Specifically, for the video or image data captured by the imaging device, it is usually necessary to identify the fixed position of the imaging device. When identifying the fixed position of the imaging device, the server will obtain the image to be recognized corresponding to the video or image data captured by the imaging device. Among them, when the imaging device captures image data, the image data can be directly used as the image to be recognized. When the imaging device captures video data, it is necessary to extract frames from the video data, and the obtained video frames are used as the images to be recognized.

[0069] Furthermore, after obtaining the image to be recognized, the server will preprocess the image to be recognized to obtain an image to be recognized with a unified size. For example, the processing method can be that the server first adjusts the size of the image to be recognized and then performs normalization processing. Among them, adjusting the size of the image to be recognized can resize it to (1024 * 512 (width * height)). The following is an example of normalization. When the average pixel of the image to be recognized is (0.485, 0.456, 0.406) and the standard pixel is (0.229, 0.224, 0.225) (where the channel order of the pixel is RGB, that is, the pixel is (r, g, b)), taking the R channel as an example, the normalized pixel r2 = (r1 / 255 - 0.485) / 0.229.

[0070] Step 104, perform object detection on the image to be recognized to obtain the object detection result corresponding to the image to be recognized. The object detection result includes the detection category and the detection box.

[0071] Among them, object detection refers to detecting the detection categories and detection bounding boxes of the regions in the image to be recognized that may be target objects. The object detection results include detection categories and detection bounding boxes, where the detection categories can be set according to needs, and the detection bounding boxes are used to mark the positions of the detected regions in the image to be recognized. For example, when the shooting device is a panoramic camera, due to the particularity of the panoramic camera compared with other cameras, its viewing angle can cover a range of 180 degrees up and down, 360 degrees or 180 degrees left and right. Therefore, the detection categories of the corresponding target objects (i.e., the fixed positions of the panoramic camera) can be divided into three types. One is that the target object does not appear in the image to be recognized. The second is that only the target object appears in the image to be recognized, such as the positions of fixed devices, such as the vehicle head, tripod, etc. The third is that the photographer and the panoramic camera appear in the image to be recognized at the same time. Further, for the third case, it can be further divided into the photographer holding the panoramic camera and the photographer holding the panoramic camera through a selfie stick.

[0072] Specifically, after the server obtains the image to be recognized, it will extract features from the image to be recognized to obtain the image feature information corresponding to the image to be recognized, and then perform object detection based on the image feature information to obtain the first detection bounding box corresponding to the image to be recognized. By screening the first detection bounding box, the object detection result is obtained. Among them, each first detection bounding box has a corresponding detection category, which will be displayed simultaneously after object detection.

[0073] Specifically, the server can implement object detection on the image to be recognized by obtaining a trained object detection model. Among them, the trained object detection model can be obtained by training the initial object detection model with sample images carrying category annotations and detection bounding box annotations. Here, the initial object detection model refers to an untrained object detection model. For example, the trained object detection model can specifically refer to an object detection model based on the FPN network. This embodiment does not specifically limit the object detection model here.

[0074] Step 106, when there is an object detection category in the object detection result, determine the object detection bounding box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain the human detection bounding box corresponding to the image to be recognized.

[0075] Among them, the target detection category refers to the category that is preset and needs to further confirm whether there is an incorrect detection. For example, when the shooting device is a panoramic camera, the target detection category refers to the third category, that is, the situation where the photographer and the panoramic camera appear in the image to be recognized at the same time. At this time, it is necessary to further verify whether there is a photographer in the image to be recognized. The target detection box refers to the detection box in the image to be recognized corresponding to the target detection category. In the target detection result, the detection category and the detection box are in one-to-one correspondence, so that the target detection box corresponding to the target detection category can be determined. Human detection refers to performing target detection on the image to be recognized to detect whether there is a human body in the image to be recognized. The human detection box is used to mark the position of the human body existing in the image to be recognized.

[0076] Specifically, after obtaining the target detection result, the server will compare the detection category in the target detection result with the preset target detection category. When the target detection category exists in the target detection result, the target detection box corresponding to the target detection category is determined, and human detection is performed on the image to be recognized to obtain the human detection box corresponding to the image to be recognized, so as to use the human detection box to screen the target detection box and reduce the occurrence of incorrect detection.

[0077] Step 108: Screen the target detection boxes according to the human detection boxes, and update the target detection result according to the screening result.

[0078] Specifically, the server will calculate the intersection-over-union ratio between each human detection box and the target detection box, and screen the target detection boxes according to the intersection-over-union ratio, and screen out the target detection boxes whose intersection-over-union ratio with the human detection box is greater than or equal to the preset intersection-over-union ratio threshold. Update the target detection result according to the screened target detection boxes, that is, delete the un-screened target detection boxes from the target detection result. Among them, the preset intersection-over-union ratio threshold can be set according to needs. For example, the preset intersection-over-union ratio threshold can be specifically 0.3.

[0079] Step 110: Obtain the target object recognition result according to the updated target detection result.

[0080] Specifically, after updating the target detection result, the server will determine the confidence level corresponding to each detection box in the updated target detection result, sort the confidence levels, and screen out the detection box corresponding to the highest confidence level as the detection box of the target object to obtain the target object recognition result. The target object recognition result includes the detection box of the target object and the detection category corresponding to the detection box.

[0081] Further, after obtaining the target object recognition result, the to-be-recognized image can be analyzed using the target object recognition result. For example, in the project of automatic video editing, it is relatively important for us to obtain the shooting subject. For instance, when a certain person is located as the shooting subject, when other objects occlude the shooting subject, we will consider this shot to be of poor quality. Another example, when there are at least two or more subjects in the to-be-recognized image, we will consider that the non-shooting subjects are more worthy of attention.

[0082] The above target object recognition method obtains the to-be-recognized image, performs target detection on the to-be-recognized image to obtain the target detection result corresponding to the to-be-recognized image. When there is a target detection category in the target detection result, the target detection box corresponding to the target detection category is determined, and human body detection is performed on the to-be-recognized image to obtain the human body detection box corresponding to the to-be-recognized image. The human body detection box can be used to screen the target detection box, and the target detection box that meets the requirements is screened out to reduce the occurrence of false detections. By updating the target detection result according to the screening result and obtaining the target object recognition result based on the updated target detection result, accurate recognition of the target object can be achieved on the basis of screening out the target detection boxes that will cause false detections.

[0083] In one embodiment, performing target detection on the to-be-recognized image to obtain the target detection result corresponding to the to-be-recognized image includes:

[0084] Performing image feature extraction on the to-be-recognized image to obtain the image feature information corresponding to the to-be-recognized image;

[0085] Performing target detection according to the image feature information to obtain the first detection box corresponding to the to-be-recognized image;

[0086] Screening the first detection box to obtain the target detection result corresponding to the to-be-recognized image.

[0087] Specifically, the server will perform image feature extraction on the to-be-recognized image to obtain the image feature information corresponding to the to-be-recognized image, perform target detection according to the image feature information to obtain the first detection box corresponding to the to-be-recognized image, and then screen the first detection box to filter out the detection boxes that do not meet the requirements to obtain the target detection result corresponding to the to-be-recognized image.

[0088] Specifically, the server can implement target detection on the to-be-recognized image by obtaining the trained target detection model to obtain the first detection box corresponding to the to-be-recognized image. For example, Figure 2As shown, when the trained object detection model is an object detection model based on the FPN network, the object detection model first uses the residual network and the FPN network to extract image features from the image to be recognized, so as to obtain multi-scale image feature information, and then inputs the multi-scale image feature information into the class subnet and the detection box subnet respectively to obtain the detection class and the detection box. Among them, the structure of the residual network can be as Figure 3 shown, and the structure of the FPN network can be as Figure 4 shown.

[0089] In this embodiment, by first extracting image features from the image to be recognized to obtain image feature information, and then performing object detection on the image feature information to obtain the first detection box, the object detection of the image to be recognized can be realized. By screening the first detection box, the detection boxes that meet the requirements can be screened out to obtain the object detection result corresponding to the image to be recognized.

[0090] In one embodiment, screening the first detection box to obtain the object detection result corresponding to the image to be recognized includes:

[0091] Obtaining the confidence of the first detection box corresponding to the first detection box;

[0092] According to the confidence of the first detection box and the preset confidence threshold, screening the first detection box to obtain the second detection box;

[0093] Using the non-maximum suppression algorithm to screen the second detection box to obtain the candidate detection box;

[0094] According to the candidate detection box, obtaining the object detection result corresponding to the image to be recognized.

[0095] Among them, the confidence of the first detection box refers to the confidence corresponding to the first detection box in the object detection result. The confidence of the first detection box is output simultaneously with the first detection box, and is used to represent the authenticity of the first detection box. That is, the higher the confidence of the first detection box, the more credible the corresponding first detection box is. Non-maximum suppression refers to suppressing elements that are not maximum values, which can be understood as local maximum search. This local represents a neighborhood, and there are two variable parameters in the neighborhood, one is the dimension of the neighborhood, and the other is the size of the neighborhood. In this embodiment, it mainly refers to suppressing detection boxes with non-maximum confidence.

[0096] Specifically, the server will obtain the confidence of the first detection box corresponding to the first detection box, filter the confidence of the first detection box according to the preset confidence threshold, retain the confidence of the first detection box that is greater than or equal to the preset confidence threshold, and then retain the first detection box corresponding thereto as the second detection box according to the retained confidence of the first detection box, so as to realize the screening of the first detection box. After screening out the second detection box, the server will further use the non-maximum suppression algorithm to screen the second detection box to suppress the detection boxes with non-maximum confidence, obtain the candidate detection boxes, and use the candidate detection boxes and the detection categories corresponding to the candidate detection boxes as the target detection results corresponding to the image to be recognized. Among them, the preset confidence threshold can be set by yourself according to needs.

[0097] In this embodiment, by obtaining the confidence of the first detection box corresponding to the first detection box, the first detection box can be screened by using the confidence of the first detection box and the preset confidence threshold to filter out the detection boxes with low confidence, obtain the second detection box, and by using the non-maximum suppression algorithm to screen the second detection box, the number of detection boxes can be reduced to obtain the candidate detection boxes, so that the target detection results can be obtained according to the candidate detection boxes.

[0098] In one embodiment, using the non-maximum suppression algorithm to screen the second detection box to obtain the candidate detection boxes includes:

[0099] Classify the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category;

[0100] Use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain the candidate detection boxes.

[0101] Specifically, the non-maximum suppression in this embodiment is mainly for the detection boxes of the same category. Therefore, after obtaining the second detection box, the server will classify the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category, and then use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain the candidate detection boxes.

[0102] Specifically, the method of screening the detection boxes of the same category in the set of detection boxes of the same category using the non-maximum suppression algorithm is as follows: First, sort the detection boxes of the same category according to the confidence level, determine the target detection box of the same category corresponding to the highest confidence level, then calculate the intersection-over-union ratio between other detection boxes of the same category in the set of detection boxes of the same category and the target detection box of the same category, remove the detection boxes of the same category whose intersection-over-union ratio with the target detection box of the same category is greater than or equal to the preset non-maximum intersection-over-union ratio threshold, use the target detection box of the same category as the candidate detection box and remove it from the set of detection boxes of the same category, then sort the remaining detection boxes of the same category in the set of detection boxes of the same category after removing both the detection boxes of the same category and the target detection box of the same category according to the confidence level to determine a new target detection box of the same category, and use the new target detection box of the same category to screen other remaining detection boxes of the same category again until the set of detection boxes of the same category is empty, and obtain the candidate detection boxes according to the target detection box of the same category obtained each time.

[0103] In this embodiment, by classifying the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category, and using the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category, the detection boxes can be refined to obtain candidate detection boxes.

[0104] In one embodiment, screening the target detection boxes according to the human body detection boxes and updating the target detection results according to the screening results includes:

[0105] Determine the intersection-over-union ratio between the human body detection box and the target detection box;

[0106] According to the intersection-over-union ratio and the preset intersection-over-union ratio threshold, screen the target detection boxes to obtain the third detection boxes, and the intersection-over-union ratio corresponding to the third detection boxes is greater than or equal to the preset intersection-over-union ratio threshold;

[0107] Update the target detection results according to the third detection boxes.

[0108] Among them, the intersection-over-union ratio refers to the overlapping rate of two detection boxes, that is, the ratio of their intersection to their union. The most ideal situation is complete overlap, that is, the ratio is 1. In this embodiment, it mainly refers to the overlapping rate between the human body detection box and the target detection box.

[0109] Specifically, after obtaining the human body detection box, the server will calculate the intersection-over-union ratio between the human body detection box and the target detection box, and screen the target detection boxes according to the intersection-over-union ratio and the preset intersection-over-union ratio threshold, so as to screen out the target detection boxes whose intersection-over-union ratio with the human body detection box is greater than or equal to the preset intersection-over-union ratio threshold as the third detection boxes, and update the target detection results according to the third detection boxes, that is, delete the target detection boxes in the target detection results that are not the third detection boxes. Among them, the preset intersection-over-union ratio threshold can be set by yourself according to needs.

[0110] In this embodiment, by determining the intersection over union ratio between the human detection frame and the target detection frame, screening the target detection frame according to the intersection over union ratio and a preset intersection over union threshold to obtain a third detection frame, and updating the target detection result according to the third detection frame, it is possible to use the human detection frame to detect whether there is a human in the target detection frame, so as to verify the existence of a photographer and reduce the occurrence of false detections.

[0111] In one embodiment, according to the updated target detection result, the target object recognition result includes:

[0112] Determine the confidence of the second detection frame corresponding to the detection frame in the updated target detection result;

[0113] Sort the confidence of the second detection frame to determine the detection frame corresponding to the highest detection frame confidence;

[0114] According to the detection frame corresponding to the highest detection frame confidence, obtain the target object recognition result.

[0115] Specifically, after updating the target detection result, the server will determine the confidence of the second detection frame corresponding to the detection frame in the updated target detection result, sort the confidence of the second detection frame to determine the highest detection frame confidence in the confidence of the second detection frame, and then determine the detection frame corresponding to the highest detection frame confidence, and use the detection frame corresponding to the highest detection frame confidence and the detection category corresponding to this detection frame as the target object recognition result.

[0116] In this embodiment, by determining the confidence of the second detection frame corresponding to the detection frame in the updated target detection result, sorting the confidence of the second detection frame to determine the detection frame corresponding to the highest detection frame confidence, it is possible to determine the target object recognition result according to the detection frame corresponding to the highest detection frame confidence.

[0117] In one embodiment, when the target detection category does not exist in the target detection result, the server will directly determine the confidence corresponding to each detection frame in the target detection result, sort the confidence to screen out the detection frame corresponding to the highest confidence as the detection frame of the target object, and obtain the target object recognition result. The target object recognition result includes the detection frame of the target object and the detection category corresponding to this detection frame.

[0118] In one embodiment, as Figure 5 shown, a flow diagram is used to illustrate the target object recognition method of the present application. The target object recognition method specifically includes the following steps:

[0119] Step 502, obtain the image to be recognized;

[0120] Step 504, perform image feature extraction on the image to be recognized to obtain image feature information corresponding to the image to be recognized;

[0121] Step 506, perform object detection based on the image feature information to obtain a first detection box corresponding to the image to be recognized;

[0122] Step 508, obtain the confidence of the first detection box corresponding to the first detection box;

[0123] Step 510, screen the first detection box according to the confidence of the first detection box and a preset confidence threshold to obtain a second detection box;

[0124] Step 512, classify the second detection box according to the detection category of the second detection box to obtain a set of detection boxes of the same category;

[0125] Step 514, use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain candidate detection boxes;

[0126] Step 516, obtain an object detection result corresponding to the image to be recognized according to the candidate detection box, where the object detection result includes a detection category and a detection box;

[0127] Step 518, when there is an object detection category in the object detection result, determine the object detection box corresponding to the object detection category, and perform human body detection on the image to be recognized to obtain a human body detection box corresponding to the image to be recognized;

[0128] Step 520, determine the intersection-over-union ratio between the human body detection box and the object detection box;

[0129] Step 522, screen the object detection box according to the intersection-over-union ratio and a preset intersection-over-union ratio threshold to obtain a third detection box, where the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union ratio threshold;

[0130] Step 524, update the object detection result according to the third detection box;

[0131] Step 526, determine the confidence of the second detection box corresponding to the detection box in the updated object detection result;

[0132] Step 528, sort the confidence of the second detection box to determine the detection box corresponding to the highest detection box confidence;

[0133] Step 530, obtain an object recognition result according to the detection box corresponding to the highest detection box confidence.

[0134] In one embodiment, taking the shooting device as a panoramic camera as an example, the technical concept of the object recognition method of the present application is described.

[0135] For a panoramic camera with an invisible selfie stick function, since it hides the selfie stick in the captured video or image data by physical means, it makes it difficult for the algorithm to identify the fixed position of the camera (i.e., the target object). How to identify the fixed position of the camera in the video or image data captured by a panoramic camera with an invisible selfie stick function is an urgent problem to be solved.

[0136] Due to the particularity of the panoramic camera compared to other cameras, its viewing angle can cover a range of 180 degrees up and down, 360 degrees or 180 degrees left and right. Therefore, the corresponding target object recognition can be divided into four cases: one is that no target object appears in the image to be recognized (i.e., they are all invisible together), the second is that the target object appears in the image to be recognized (such as a fixed device, such as the front of a vehicle, a tripod, etc.), the third is that the photographer holding the panoramic camera appears in the image to be recognized, and the fourth is that the photographer holds the panoramic camera through a selfie stick and appears in the image to be recognized. Since the target object cannot be recognized by image recognition technology in the first case, this application mainly focuses on the latter several cases.

[0137] Based on the analysis of the above situations and combined with specific image analysis, we can infer the target object from some image features. For case two, larger distortions will appear in some edge parts of the image to be recognized. For example, in the video or image data captured by a panoramic camera set on the front of a vehicle, larger distortions will appear in the lower half of the image to be recognized; for case three, the photographer will appear in the image to be recognized and the distortion of the hand is serious; for case four, the photographer will appear in the image to be recognized and there are obvious actions of holding a selfie stick. When the above image features are not detected, it can be attributed to case one, mainly including situations such as being photographed by others or the fixed objects such as cameras being too small to be visible in the image to be recognized.

[0138] Therefore, we can infer the fixed position of the panoramic camera by detecting the appearance of these image phenomena. Utilizing the powerful extraction ability of the convolutional neural network, we can transform it into a multi-object detection problem. That is, for the above-mentioned Situation 2 to Situation 4, we perform object detection on the image to be recognized by training an object detection model, obtaining the corresponding detection categories (including the categories corresponding to Situation 2, Situation 3, and Situation 4 respectively) and detection bounding boxes. At the same time, according to the particularity of the problem, we can improve the accuracy of object detection through some measures. When Situation 3 and Situation 4 occur, the photographer will definitely appear in the image. We can combine a human detector to reduce some false detections. On a single image to be recognized, there is at most one fixed object of the camera. We can utilize this feature to further reduce the occurrence of false detections. For example, we can finally sort the confidences of all detection bounding boxes to determine the detection bounding box corresponding to the highest confidence as the recognition result of the target object, that is, the fixed position where the panoramic camera is located.

[0139] It should be noted that in this application, mainly through the image feature information of the image to be recognized, we detect whether there is a fixed position of the shooting device in the image to be recognized. Using human detection is only to improve the detection accuracy and utilize some existing prior knowledge to avoid the occurrence of some obvious detection errors.

[0140] Specifically, for the detection bounding boxes of Situation 3 and Situation 4, it means that there may be a photographer in this area. At this time, we can use human detection to detect whether there is a human in this area to further verify the existence of a photographer in this area. Furthermore, in this application, we can also use a hand detector for human detection because when the above situations occur, the photographer's hand will definitely appear in the picture and there will be corresponding image features, that is, hand distortion or the action of holding a selfie stick.

[0141] It should be clear that there is at most one fixed position of the shooting device in a single image to be recognized. This is a very useful piece of information. The object detection result may include multiple detection bounding boxes. In addition to using the non-maximum suppression algorithm to reduce the number of detection bounding boxes, we can also select only the detection bounding box with the highest confidence as the recognition result of the target object. It should be noted that at this time, the confidence corresponding to this recognition result of the target object should exceed the preset confidence threshold; otherwise, it belongs to Situation 1, that is, no target object appears in the image to be recognized.

[0142] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0143] Based on the same inventive concept, the embodiments of the present application also provide a target object recognition device for implementing the above-mentioned target object recognition method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the target object recognition device provided below can refer to the limitations on the target object recognition method in the above text, and will not be repeated here.

[0144] In one embodiment, as Figure 6 shown, a target object recognition device is provided, including: an acquisition module 602, a first detection module 604, a second detection module 606, a screening module 608, and a processing module 610, where:

[0145] The acquisition module 602 is configured to acquire an image to be recognized;

[0146] The first detection module 604 is configured to perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized. The target detection result includes a detection category and a detection box;

[0147] The second detection module 606 is configured to, when there is a target detection category in the target detection result, determine a target detection box corresponding to the target detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0148] The screening module 608 is configured to screen the target detection boxes according to the human detection box, and update the target detection result according to the screening result;

[0149] The processing module 610 is configured to obtain a target object recognition result according to the updated target detection result.

[0150] The above-mentioned target object recognition device obtains an image to be recognized, performs target detection on the image to be recognized, obtains a target detection result corresponding to the image to be recognized, determines a target detection box corresponding to the target detection category when there is a target detection category in the target detection result, and performs human body detection on the image to be recognized to obtain a human body detection box corresponding to the image to be recognized. It can use the human body detection box to screen the target detection box, screen out the target detection boxes that meet the requirements, so as to reduce the occurrence of false detections. By updating the target detection result according to the screening result, and obtaining the target object recognition result according to the updated target detection result, it can accurately recognize the target object on the basis of screening out the target detection boxes that will cause false detections.

[0151] In one embodiment, the first detection module is further configured to extract image features from the image to be recognized, obtain image feature information corresponding to the image to be recognized, perform target detection according to the image feature information, obtain a first detection box corresponding to the image to be recognized, and screen the first detection box to obtain a target detection result corresponding to the image to be recognized.

[0152] In one embodiment, the first detection module is further configured to obtain a first detection box confidence corresponding to the first detection box, screen the first detection box according to the first detection box confidence and a preset confidence threshold to obtain a second detection box, use the non-maximum suppression algorithm to screen the second detection box to obtain a candidate detection box, and obtain a target detection result corresponding to the image to be recognized according to the candidate detection box.

[0153] In one embodiment, the first detection module is further configured to classify the second detection box according to the detection category of the second detection box to obtain a set of detection boxes of the same category, and use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain a candidate detection box.

[0154] In one embodiment, the screening module is further configured to determine the intersection-over-union ratio between the human body detection box and the target detection box, screen the target detection box according to the intersection-over-union ratio and a preset intersection-over-union threshold to obtain a third detection box, the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union threshold, and update the target detection result according to the third detection box.

[0155] In one embodiment, the processing module is further configured to determine the second detection box confidence corresponding to the detection box in the updated target detection result, sort the second detection box confidence, determine the detection box corresponding to the highest detection box confidence, and obtain the target object recognition result according to the detection box corresponding to the highest detection box confidence.

[0156] Each module in the above target object recognition device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0157] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as target detection categories. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a target object recognition method.

[0158] Those skilled in the art can understand that Figure 7 the structure shown in

[0159] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0160] Obtain an image to be recognized;

[0161] Perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized. The target detection result includes a detection category and a detection box;

[0162] When there is a target detection category in the target detection result, determine the target detection box corresponding to the target detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0163] Screen the target detection box according to the human detection box, and update the target detection result according to the screening result;

[0164] Obtain a target object recognition result according to the updated target detection result.

[0165] In one embodiment, when the processor executes the computer program, the following steps are further implemented: extracting image features from the image to be recognized to obtain image feature information corresponding to the image to be recognized, performing object detection based on the image feature information to obtain a first detection box corresponding to the image to be recognized, screening the first detection box to obtain a target detection result corresponding to the image to be recognized.

[0166] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining a first detection box confidence corresponding to the first detection box, screening the first detection box according to the first detection box confidence and a preset confidence threshold to obtain a second detection box, using a non-maximum suppression algorithm to screen the second detection box to obtain candidate detection boxes, and obtaining a target detection result corresponding to the image to be recognized according to the candidate detection boxes.

[0167] In one embodiment, when the processor executes the computer program, the following steps are further implemented: classifying the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category, and using a non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain candidate detection boxes.

[0168] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining an intersection-over-union ratio between the human detection box and the target detection box, screening the target detection box according to the intersection-over-union ratio and a preset intersection-over-union threshold to obtain a third detection box, where the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union threshold, and updating the target detection result according to the third detection box.

[0169] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining the second detection box confidence corresponding to the detection box in the updated target detection result, sorting the second detection box confidence, determining the detection box corresponding to the highest detection box confidence, and obtaining a target object recognition result according to the detection box corresponding to the highest detection box confidence.

[0170] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0171] Obtaining the image to be recognized;

[0172] Performing object detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized, where the target detection result includes a detection category and a detection box;

[0173] When there is a target detection category in the target detection result, determining the target detection box corresponding to the target detection category, and performing human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized;

[0174] Screen the target detection boxes according to the human detection boxes, and update the target detection results according to the screening results;

[0175] Obtain the target object recognition result according to the updated target detection result.

[0176] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extract image features from the image to be recognized to obtain image feature information corresponding to the image to be recognized, perform target detection according to the image feature information to obtain a first detection box corresponding to the image to be recognized, and screen the first detection box to obtain a target detection result corresponding to the image to be recognized.

[0177] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtain the confidence of the first detection box corresponding to the first detection box, screen the first detection box according to the confidence of the first detection box and the preset confidence threshold to obtain a second detection box, and use the non-maximum suppression algorithm to screen the second detection box to obtain a candidate detection box, and obtain a target detection result corresponding to the image to be recognized according to the candidate detection box.

[0178] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: classify the second detection box according to the detection category of the second detection box to obtain a set of detection boxes of the same category, and use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain a candidate detection box.

[0179] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determine the intersection-over-union ratio between the human detection box and the target detection box, screen the target detection box according to the intersection-over-union ratio and the preset intersection-over-union ratio threshold to obtain a third detection box, the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union ratio threshold, and update the target detection result according to the third detection box.

[0180] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determine the confidence of the second detection box corresponding to the detection box in the updated target detection result, sort the confidence of the second detection box, determine the detection box corresponding to the highest detection box confidence, and obtain the target object recognition result according to the detection box corresponding to the highest detection box confidence.

[0181] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0182] Obtain the image to be recognized;

[0183] Perform object detection on the image to be recognized, and obtain the object detection result corresponding to the image to be recognized. The object detection result includes the detection category and the detection box;

[0184] When there is an object detection category in the object detection result, determine the object detection box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain the human detection box corresponding to the image to be recognized;

[0185] Screen the object detection box according to the human detection box, and update the object detection result according to the screening result;

[0186] According to the updated object detection result, obtain the target object recognition result.

[0187] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: extract image features from the image to be recognized to obtain the image feature information corresponding to the image to be recognized, perform object detection according to the image feature information to obtain the first detection box corresponding to the image to be recognized, and screen the first detection box to obtain the object detection result corresponding to the image to be recognized.

[0188] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: obtain the confidence of the first detection box corresponding to the first detection box, screen the first detection box according to the confidence of the first detection box and the preset confidence threshold to obtain the second detection box, use the non-maximum suppression algorithm to screen the second detection box to obtain the candidate detection box, and obtain the object detection result corresponding to the image to be recognized according to the candidate detection box.

[0189] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: classify the second detection box according to the detection category of the second detection box to obtain the set of detection boxes of the same category, and use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain the candidate detection box.

[0190] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: determine the intersection-over-union ratio between the human detection box and the object detection box, screen the object detection box according to the intersection-over-union ratio and the preset intersection-over-union ratio threshold to obtain the third detection box, and the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union ratio threshold, and update the object detection result according to the third detection box.

[0191] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: determine the confidence of the second detection box corresponding to the detection box in the updated object detection result, sort the confidence of the second detection box, determine the detection box corresponding to the highest confidence of the detection box, and obtain the target object recognition result according to the detection box corresponding to the highest confidence of the detection box.

[0192] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0193] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0194] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for identifying a target object, characterized in that, The target object is a fixed position of a panoramic camera, and the method includes: Obtain an image to be recognized; Perform object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized. The object detection result includes a detection category and a detection box. The detection category includes that the target object does not appear in the image to be recognized, only the target object appears in the image to be recognized, and both the photographer of the image to be recognized and the panoramic camera appear in the image to be recognized; When there is an object detection category in the object detection result where both the photographer of the image to be recognized and the panoramic camera appear in the image to be recognized, determine the object detection box corresponding to the object detection category, and perform human detection on the image to be recognized to obtain a human detection box corresponding to the image to be recognized; Screen the object detection box according to the human detection box, and update the object detection result according to the screening result; Obtain an object recognition result according to the updated object detection result.

2. The method according to claim 1, characterized in that The performing object detection on the image to be recognized to obtain an object detection result corresponding to the image to be recognized includes: Extract image features from the image to be recognized to obtain image feature information corresponding to the image to be recognized; Perform object detection according to the image feature information to obtain a first detection box corresponding to the image to be recognized; Screen the first detection box to obtain an object detection result corresponding to the image to be recognized.

3. The method according to claim 2, wherein The screening the first detection box to obtain an object detection result corresponding to the image to be recognized includes: Obtain a confidence level of the first detection box corresponding to the first detection box; Screen the first detection box according to the confidence level of the first detection box and a preset confidence threshold to obtain a second detection box; Use the non-maximum suppression algorithm to screen the second detection box to obtain candidate detection boxes; Obtain an object detection result corresponding to the image to be recognized according to the candidate detection boxes.

4. The method according to claim 3, characterized in that, The using the non-maximum suppression algorithm to screen the second detection box to obtain candidate detection boxes includes: Classify the second detection boxes according to the detection categories of the second detection boxes to obtain a set of detection boxes of the same category; Use the non-maximum suppression algorithm to screen the detection boxes of the same category in the set of detection boxes of the same category to obtain candidate detection boxes.

5. The method according to claim 1, characterized in that, The screening the object detection box according to the human detection box and updating the object detection result according to the screening result includes: Determine the intersection-over-union ratio between the human detection box and the object detection box; Screen the object detection box according to the intersection-over-union ratio and a preset intersection-over-union threshold to obtain a third detection box, and the intersection-over-union ratio corresponding to the third detection box is greater than or equal to the preset intersection-over-union threshold; Update the object detection result according to the third detection box.

6. The method according to claim 1, characterized in that, The obtaining an object recognition result according to the updated object detection result includes: Determine the confidence level of the second detection box corresponding to the detection box in the updated object detection result; Sort the confidence levels of the second detection frames to determine the detection frame corresponding to the highest detection frame confidence level; Obtain the target object recognition result according to the detection frame corresponding to the highest detection frame confidence level.

7. An object recognition device, characterized in that, The target object is a fixed position of a panoramic camera, and the device includes: An acquisition module, configured to acquire an image to be recognized; A first detection module, configured to perform target detection on the image to be recognized to obtain a target detection result corresponding to the image to be recognized, where the target detection result includes a detection category and a detection frame, and the detection category includes that the target object does not appear in the image to be recognized, only the target object appears in the image to be recognized, and both the photographer of the image to be recognized and the panoramic camera appear in the image to be recognized; A second detection module, configured to, when there is a target detection category in the target detection result that both the photographer of the image to be recognized and the panoramic camera appear in the image to be recognized, determine a target detection frame corresponding to the target detection category, and perform human body detection on the image to be recognized to obtain a human body detection frame corresponding to the image to be recognized; A screening module, configured to screen the target detection frame according to the human body detection frame, and update the target detection result according to the screening result; A processing module, configured to obtain a target object recognition result according to the updated target detection result.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Image recognition method and device, terminal equipment and computer readable storage medium

    CN113158869A