Method, device, computer equipment and storage medium for object detection in panoramic images

By using preset object deformation adaptation convolution operators in convolution neural networks, convolution processing on panoramic images is solved, and the detection result deviation caused by distortion during object detection in panoramic images is achieved, and a higher target detection accuracy is achieved.

CN114005052BActive Publication Date: 2025-05-06ARASHI VISION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111233006.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-05-06
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

When detecting objects in panoramic images in the prior art, the rectangular frame of the detection result cannot reasonably frame the deformation and extension targets due to distortion, resulting in deviations in the detection result.

Method used

A convolutional neural network containing preset target deformation adaptation convolution operators is used to convolutionize the panoramic image to obtain target category and target boundary position point data, including coordinates beyond the image boundary, in order to more accurately detect the target.

Benefits of technology

By presetting the target deformation adaptation convolution operator, the target deformation in the panoramic image can be effectively adapted to the target deformation in the panoramic image, the accuracy of object detection is improved, and all target areas, including the detection target of the image boundary, can be accurately detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005052B_ABST
    Figure CN114005052B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, computer equipment and storage medium for target detection in panoramic images. The method obtains a panoramic image to be detected; performs convolution processing on the panoramic image to be detected through a convolutional neural network including a preset target deformation adaptive convolution operator, and obtains the target category and target boundary position point data of the detection target in the panoramic image to be detected, and the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; according to the target category and target boundary position point data of the detection target, obtains the target detection result corresponding to the panoramic image to be detected. The present application can effectively determine the area of ​​all detection targets including the detection targets arranged at both ends of the panoramic image, thereby improving the accuracy of panoramic image target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a method, device, computer equipment and storage medium for detecting objects in panoramic images. Background Art

[0002] With the development of artificial intelligence technology, computer vision technology has also been more and more widely used. Computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify, track and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain "information" from images or multidimensional data. Object detection in panoramic images is one of the research objects of computer vision. Object detection in panoramic images belongs to the sub-field of object detection. It is a computer vision technology based on the statistical characteristics and semantic information of panoramic targets. It can simultaneously obtain the category information and location information of targets in panoramic images.

[0003] A panoramic image is a special image with an aspect ratio of 2:1, which is composed of multiple images. It follows the longitude and latitude expansion method, where the width of the image is the latitude 0-2π, and the height of the image is the longitude 0-π. Therefore, it can record all information of 360 degrees horizontally and 180 degrees in pitch. At present, when performing target detection on a panoramic image, target detection is generally performed on the flat expanded image of the panoramic image, so some objects in the panoramic image will be distorted, resulting in the rectangular frame of the detection result not being able to reasonably frame the deformed and extended target, resulting in deviations in the detection result.

[0004] To address the problem that some objects in panoramic images are distorted, resulting in deviations in detection results, BFoV (Bounding Field-of-Views) can be used to represent targets in panoramic images. BFoV regards the panoramic image as a sphere, uses the latitude and longitude coordinates of the target to represent its center point, and uses its two field of view angles in the horizontal and vertical directions to represent the space it occupies. However, due to the unique symmetry of BFoV, the detection area it represents still does not always contain tilted or deformed targets in the panoramic image, which affects the detection effect of target detection. Summary of the invention

[0005] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for object detection in panoramic images, which can improve the accuracy of object detection in panoramic images, in order to address the above technical problems.

[0006] A method for detecting an object in a panoramic image, the method comprising:

[0007] Acquire the panoramic image to be detected;

[0008] Performing convolution processing on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator to obtain target categories and target boundary position point data of the detected target in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0009] According to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained.

[0010] In one of the embodiments, performing convolution processing on the panoramic image to be detected by a convolutional neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detection target in the panoramic image to be detected includes:

[0011] Inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator, and obtaining a heat map, target category, and target boundary position point data corresponding to an initial detection target in the panoramic image to be detected;

[0012] Filtering the initial detection target according to the thermal map to obtain the detection target;

[0013] Obtain the target category corresponding to the detected target and target boundary position point data.

[0014] In one of the embodiments, the initial detection target includes a non-boundary position target, and the convolutional neural network also includes a conventional convolution operator;

[0015] The step of inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a heat map, target category, and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected includes:

[0016] Extracting panoramic image features of the panoramic image to be detected by using the conventional convolution operator;

[0017] Based on the panoramic image features, non-boundary position targets in the panoramic image to be detected are determined, and a heat map, target category, and target boundary position point data corresponding to the non-boundary position targets are obtained.

[0018] In one embodiment, the initial detection target includes a boundary position target, and the step of inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a heat map, target category, and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected includes:

[0019] Extracting panoramic image features of the panoramic image to be detected by using a preset target deformation adaptive convolution operator;

[0020] Determine a boundary position target at a boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target;

[0021] identifying target attributes between a first detection target and a second detection target based on the panoramic image feature, wherein the first detection target and the second detection target are initial detection targets at relative positions in the panoramic image;

[0022] When the target attribute indicates that the first detection target and the second detection target are the same detection target, obtaining a heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target;

[0023] According to the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target, the heat map, target category and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected are obtained.

[0024] In one of the embodiments, before performing convolution processing on the panoramic image to be detected by a convolutional neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detected target in the panoramic image to be detected, it also includes:

[0025] Acquire a historical panoramic image with target category annotations and target boundary position point annotations, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0026] Constructing a model training data set based on the historical panoramic images;

[0027] The initial convolutional neural network including the preset target deformation adaptive convolution operator is trained through the model training data group to obtain the convolutional neural network including the preset target deformation adaptive convolution operator.

[0028] In one of the embodiments, obtaining the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detected target includes:

[0029] Generate a data group corresponding to the detected target according to the target category and target boundary position point data of the detected target;

[0030] The data groups corresponding to each detected target are filled into a preset target detection result list to generate the target detection result corresponding to the panoramic image to be detected.

[0031] A device for detecting an object in a panoramic image, the device comprising:

[0032] A data acquisition module, used to acquire the panoramic image to be detected;

[0033] A convolution processing module, configured to perform convolution processing on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator, to obtain target categories and target boundary position point data of the detected targets in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0034] The target detection module is used to obtain the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detection target.

[0035] In one of the embodiments, the convolution processing module is specifically used to: input the panoramic image to be detected into a convolutional neural network containing a preset target deformation adaptive convolution operator to obtain a heat map, target category and target boundary position point data corresponding to an initial detection target in the panoramic image to be detected; filter the initial detection target according to the heat map to obtain a detection target; obtain the target category and target boundary position point data corresponding to the detection target.

[0036] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0037] Acquire the panoramic image to be detected;

[0038] Performing convolution processing on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator to obtain target categories and target boundary position point data of the detected target in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0039] According to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained.

[0040] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0041] Acquire the panoramic image to be detected;

[0042] Performing convolution processing on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator to obtain target categories and target boundary position point data of the detected target in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0043] According to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained.

[0044] The above-mentioned panoramic image target detection method, device, computer equipment and storage medium obtain a panoramic image to be detected; perform convolution processing on the panoramic image to be detected through a convolution neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detection target in the panoramic image to be detected, and the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; according to the target category and target boundary position point data of the detection target, obtain the target detection result corresponding to the panoramic image to be detected. When detecting a panoramic image, the present application uses a preset target deformation adaptive convolution operator to extract the target category and target boundary position point data of the detection target, and then obtains the final target detection result through the target category and target boundary position point data. The target boundary position point is used to represent the position of the detection target, and the area of ​​all detection targets including the detection target at the boundary of the panoramic image can be effectively determined, thereby improving the accuracy of panoramic image target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A diagram showing an application environment of a method for detecting an object in a panoramic image in an embodiment;

[0046] Figure 2 is a schematic diagram of a flow chart of a method for detecting an object in a panoramic image in one embodiment;

[0047] Figure 3 In one embodiment Figure 2 Schematic diagram of the sub-process of step 203;

[0048] Figure 4 In one embodiment Figure 3 Schematic diagram of a sub-process of step 302;

[0049] Figure 5A flowchart of a neural network model training step in one embodiment;

[0050] Figure 6 is a structural block diagram of a panoramic image object detection device in one embodiment;

[0051] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The applicant has found that the existing panoramic images generally have the phenomenon of panoramic distortion. Panoramic distortion refers to the fact that in the scanning and imaging process of the panoramic image, since the image distance remains unchanged, the object distance increases with the increase of the scanning angle, resulting in a gradual reduction in the scale from the center to the sides of the image. Most of the existing target detection algorithms for panoramic images use the target's Bounding-Box (BBox) or Bounding Field-of-View (BFoV). However, in panoramic images, due to the existence of distortion, in the detection method using BBox, the detected rectangular frame cannot reasonably frame the deformed and extended detection target, resulting in detection failure. For BFoV, BFoV regards the panoramic image as a sphere, uses the latitude and longitude coordinates of the target detection mark to represent its center point, and uses its two field of view angles (Field-of-Views) in the horizontal and vertical directions to represent the space it occupies. BFoV is specifically defined as (φ, θ, h, w). φ and θ are the latitude and longitude coordinates of the target on the sphere, respectively; h and w represent the two field of view angles of the target in the horizontal and vertical directions, similar to height and width. The distortion of the upper and lower areas of BFoV can be extended to include the target. At the same time, since BFoV is defined on the sphere, the problem of not being able to judge whether the left and right sides are the same target is avoided. However, due to the unique symmetry of BFoV, when the objects in the panoramic image are asymmetrical at the boundaries at both ends, the area detected by BFoV still cannot always contain the detected target well, thereby affecting the accuracy of target detection. In response to this situation, the applicant proposed the target detection method of the present application.

[0054] The object detection method of the panoramic image provided by this application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. Among them, when the data processing staff of the terminal 102 needs to detect the target in the panoramic image, the panoramic image to be detected can be sent to the server 104, and the server 104 performs target detection on the panoramic image to be detected submitted by the terminal 102. The server 104 obtains the panoramic image to be detected; the panoramic image to be detected is convoluted by a convolutional neural network containing a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detected target in the panoramic image to be detected, and the target boundary position point data includes the target boundary position point whose coordinates exceed the boundary of the panoramic image to be detected; according to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained. Among them, the terminal 102 can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0055] In one embodiment, Figure 2 As shown, a method for detecting a target in a panoramic image is provided. Figure 1 Taking the server 104 in the example as an example, the following steps are included:

[0056] Step 201: Acquire a panoramic image to be detected.

[0057] Step 203, convolution processing is performed on the panoramic image to be detected by a convolutional neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detection target in the panoramic image to be detected, and the target boundary position point data includes the target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected.

[0058] Among them, a panoramic image is a special image with an aspect ratio of generally 2:1, which is composed of multiple images. According to the longitude and latitude expansion method, the width of the image is the latitude 0-2π, and the height of the image is the longitude 0-π. Therefore, it can record all information of 360 degrees horizontally and 180 degrees in pitch. At present, if a panoramic image is used for target detection, some objects in the panoramic image will be divided into the left and right sides of the image in the horizontal direction, which makes it impossible to detect them as the same object. At the same time, due to the existence of panoramic distortion, the detection method of target detection through a rectangular frame cannot effectively frame the detection target, thereby affecting the accuracy of target detection of the panoramic image. Accurate target detection for panoramic images can be achieved through the target detection method of the panoramic image of the present application. Operators are the basic units of neural network calculations, and convolution operations are the main components of neural networks, which are used to extract statistical and semantic features of targets. The preset target deformation adaptive convolution operator refers to the present application modifying the existing convolutional neural network model for target detection by replacing some convolution operators with convolution operators that can adapt to target deformation, such as deformable convolution, equirectangular projection convolution, and spherical convolution. These preset target deformation adaptive convolution operators are obtained by training with panoramic images. The target boundary position point can be a coordinate point at the boundary position of the detection target. For a detection target, there can be multiple boundary coordinate points. In one embodiment, there are 9 boundary coordinate points, that is, the left, middle and right points above, middle and below the target are used, and a total of 9 points represent the location of the target in the panoramic image. They are: upper left, upper middle, upper right, middle left, center, middle right, lower left, lower middle and lower right. The area formed by connecting the 9 points of the target boundary is selected to represent the area where the target is located. Compared with the rectangular frame representation of 4 points, it has better scalability and can represent more complex shapes. It can effectively adapt to the distortion of the target such as tilt and extension, so as to detect the target more accurately. For the target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected, it specifically includes target boundary position points with negative coordinates and target boundary position points greater than the image width, where negative coordinates represent the left side exceeding the boundary, and coordinates greater than the image width represent the right side exceeding the boundary. In the panoramic image, this coordinate beyond the boundary is meaningful, which indicates that there is still a part of the target on the other side of the image. The target category of the detected target is preset data. According to the purpose of target detection, you can set which categories of targets need to be identified when training the convolutional neural network. The convolutional neural network is not limited here, and can be implemented by anchor free target detection neural network models such as CornerNet, CenterNet, and FCOS.

[0059] Specifically, when the terminal 102 needs to perform target detection of a panoramic image, the panoramic image to be detected can be submitted to the server 104 through the terminal 102, so that the target detection corresponding to the panoramic image to be detected can be performed through the server 104 to determine the type of the detected target in the panoramic image to be detected and the position of the detected target. The server 104 receives the panoramic image to be detected. That is, the panoramic image to be detected can be convolved by a convolutional neural network containing a preset target deformation adaptation convolution operator to obtain the target category and target boundary position point data of the detected target in the panoramic image to be detected, and the target boundary position point data includes the target boundary position point whose coordinates exceed the boundary of the panoramic image to be detected. In the process of target detection, the detection model of target detection needs to be able to extract features on the left and right sides of the image and judge that it is the same target. The traditional convolutional neural network has weak processing capabilities for this situation. Therefore, the present application replaces some traditional convolution operators with preset target deformation adaptation convolution operators to construct a convolution model that is more suitable for panoramic images. The preset target deformation adaptation convolution operator has better adaptability to the target deformation of the panoramic image. By using the preset target deformation adaptation convolution operator to perform convolution processing on the target detection candidate area at the border of the panoramic image, it is possible to effectively determine whether the targets on the left and right sides of the panoramic image are the same target, and output the corresponding target category and a set of target boundary position point data for the same target. For targets at non-border locations, other conventional target detection convolution operators of the convolutional neural network can be used for detection.

[0060] Step 205 , obtaining the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detected target.

[0061] Specifically, after obtaining the target categories and target boundary position point data corresponding to all the detected targets in the panoramic image to be detected, this part of the data can be sorted, that is, the target category and target boundary position point data of each detected target in the panoramic image to be detected are merged and sorted, and then the target detection result corresponding to the panoramic image to be detected is output.

[0062] The above-mentioned target detection method for panoramic images obtains a panoramic image to be detected; performs convolution processing on the panoramic image to be detected through a convolution neural network including a preset target deformation adaptive convolution operator, and obtains the target category and target boundary position point data of the detection target in the panoramic image to be detected, and the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; and obtains the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detection target. When detecting a panoramic image, the present application uses a preset target deformation adaptive convolution operator to extract the target category and target boundary position point data of the detection target, and then obtains the final target detection result through the target category and target boundary position point data. The target boundary position point is used to represent the position of the detection target, and the area of ​​all detection targets including the detection target at the boundary of the panoramic image can be effectively determined, thereby improving the accuracy of panoramic image target detection.

[0063] In one embodiment, Figure 3 As shown, step 203 includes:

[0064] Step 302: Input the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a heat map, target category, and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected.

[0065] Step 304: filter the initial detection target according to the heat map to obtain the detection target.

[0066] Step 306, obtaining the target category corresponding to the detected target and the target boundary position point data.

[0067] The heat map is the feature map output by the neural network. Each point in the map represents the confidence that a target exists at that position. Therefore, it is possible to determine whether a target exists at each point in the panoramic image to be detected based on the heat map.

[0068] Specifically, when performing target detection, the obtained panoramic image to be detected can be input into a trained convolutional neural network containing a preset target deformation adaptation convolution operator, and the convolutional neural network outputs multiple branches, namely, a heat map of the detected target, a target category, and target boundary position point data, wherein the target boundary position point data can specifically include an offset of the target boundary position point, and the target boundary position point can be located by the offset of the target boundary position point; then the server 104 uses the heat map to filter out the detection targets below a certain threshold by parsing the output of the convolutional neural network, and retains the detection targets with high confidence; finally, the target category and target boundary position point data corresponding to the required detection target can be obtained to realize target detection of the panoramic image. In this embodiment, the initial detection target is filtered by the heat map, and the initial detection target that does not meet the requirements can be effectively excluded, thereby improving the accuracy of target detection.

[0069] In one embodiment, the initial detection target includes a non-boundary position target, and the convolutional neural network also includes a conventional convolution operator. Step 302 includes: extracting panoramic image features of the panoramic image to be detected by using a conventional convolution operator; determining the non-boundary position target in the panoramic image to be detected based on the panoramic image features, and obtaining a heat map, target category, and target boundary position point data corresponding to the non-boundary position target.

[0070] Among them, the non-boundary position target refers to the initial detection target that is not divided into the two ends of the panoramic image. The non-boundary target is a complete target, generally located in the middle of the panoramic image. The panoramic image feature can include the coordinate position of the currently detected initial detection target in the panoramic image to be detected, so that it can be determined which initial detection targets in the panoramic image to be detected belong to the non-boundary position target based on the panoramic image feature.

[0071] Specifically, for the detection of non-boundary position targets, convolution calculations can be performed using the usual convolution operators in the convolutional neural network to extract the corresponding heat map, target category, and target boundary position point data. When identifying non-boundary position targets, the non-boundary position targets in the panoramic image to be detected can be identified by whether the coordinates of the initial detection target in the panoramic image features include the coordinates at the boundary position of the panoramic image to be detected. When the coordinates of the initial detection target do not include the coordinates at the boundary position of the panoramic image to be detected, that is, when all the coordinates in the initial detection target are within the boundary range of the panoramic image to be detected, the initial detection target is regarded as a non-boundary position target. In this embodiment, by extracting the panoramic image features of the panoramic image to be detected, the initial detection target corresponding to the non-boundary position target can be effectively detected to ensure the detection effect of the target detection.

[0072] In one embodiment, Figure 4As shown, the initial detection target also includes a boundary position target, and step 302 includes:

[0073] Step 401: extract panoramic image features of the panoramic image to be detected by using a preset target deformation adaptive convolution operator.

[0074] Step 403: determine the boundary position target at the boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target.

[0075] Step 405 , identifying target attributes between a first detection target and a second detection target based on the panoramic image feature, wherein the first detection target and the second detection target are initial detection targets at relative positions in the panoramic image.

[0076] Step 407, when the target attribute indicates that the first detection target and the second detection target are the same detection target, obtaining the heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target;

[0077] Step 409, based on the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target, obtain the heat map, target category and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected.

[0078] Among them, the boundary position target refers to the segmented initial detection target, and the different parts of the boundary position target are generally arranged at the left and right ends of the panoramic image. By presetting the target deformation adaptation convolution operator, the target property can be effectively adapted to the deformation, and the panoramic image features can be extracted by presetting the target deformation adaptation convolution operator to extract the panoramic image, which can effectively extract the panoramic image features corresponding to the boundary position target from the panoramic image to be detected. The detection target refers to the target located at the boundary of the panoramic image in the panoramic image to be detected, and these targets are segmented to both sides of the image by the panoramic image. The target attribute is specifically used to determine whether the first detection target and the second detection target, which are two detection targets in relative positions, are the same target. When the two detection targets in relative positions are the same target, the target attributes of the two detection targets are the same. When the two detection targets in relative positions are not the same target, the target attributes of the two detection targets are different.

[0079] Specifically, when identifying the coordinates at the boundary position, since the target may have been segmented into two opposite boundaries in the panoramic image, the target deformation is generated. Therefore, at this time, the panoramic image features corresponding to these targets can be extracted by presetting the target deformation adaptation convolution operator. Based on the extracted panoramic image features, it is further determined which targets belong to the detection targets, and the target attributes corresponding to the two detection targets at relative positions are identified. For example, for a panoramic image to be detected with a width of 0-2π latitude and a height of 0-π longitude. A two-dimensional plane coordinate system can be established with the lower left end point of the image as the origin, the width direction of the image as the X axis, and the height direction of the image as the Y axis. Then the boundary position in the panoramic image to be detected is the left boundary of X=0 and the right boundary of X=2π. For the detection targets at relative positions, it specifically refers to the detection targets containing the same Y-axis coordinates. For example, if the coordinates of a detection target A are identified to include (0, 0.5π), it can be determined that the detection target B containing the coordinates (2π, 0.5π) is the boundary position target of the relative position of the detection target A. Then, the panoramic image features extracted by the convolutional neural network can be used to further identify and judge whether the two detection targets in relative positions are the same detection targets. When the target attribute characterizes that the detection targets in relative positions are the same target, the target boundary position point data corresponding to the detection target can be obtained. That is, according to the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target, the heat map, target category and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected are obtained. Because the first detection target and the second detection target are the same target, when identifying the target, any one of them can be selected as the final boundary position target. Generally, the detection target on either side of the left or right can be fixed as the final boundary position target. In this embodiment, by extracting the panoramic image features of the target detection candidate area, the target boundary position point data corresponding to the detection target can be effectively detected to ensure the detection effect of the target detection.

[0080] In one embodiment, if Figure 5 As shown, before step 203, the following steps are also included:

[0081] Step 502 : Acquire a historical panoramic image with target category annotations and target boundary position point annotations. The target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected.

[0082] Step 504: construct a model training data set based on the historical panoramic images.

[0083] Step 506: Train the initial convolutional neural network including the preset target deformation adaptive convolution operator through the model training data group to obtain the convolutional neural network including the preset target deformation adaptive convolution operator.

[0084] Among them, the historical panoramic images specifically refer to panoramic images of detected targets under various target categories in the historical data. These historical panoramic images can be used to train the initial state of the convolutional neural network to obtain a convolutional neural network that contains a preset target deformation adaptation convolution operator.

[0085] Specifically, when constructing the model training data, the categories and target boundary position points corresponding to each detection target in the historical panoramic image can be annotated manually, and the model training data group can be constructed through the annotated historical panoramic image. Then, the initial convolutional neural network containing the target deformation adaptation convolution operator is trained through the model training data group to obtain the convolutional neural network containing the preset target deformation adaptation convolution operator. In other embodiments, in addition to constructing the model training data group, the historical panoramic image can also construct a model verification group to verify the trained convolutional neural network. Only when the recognition accuracy of the verification group is higher than the preset threshold, the trained model can be used as a convolutional neural network containing the preset target deformation adaptation convolution operator. Otherwise, the model parameters need to be adjusted before training. In this embodiment, by constructing the model training data group, the training of the neural network model can be effectively completed to ensure the accuracy of target detection.

[0086] In one embodiment, step 205 includes: generating a data group corresponding to the detection target based on the target category and target boundary position point data of the detection target; filling the data group corresponding to each detection target into a preset target detection result list to generate a target detection result corresponding to the panoramic image to be detected.

[0087] Specifically, a blank target detection result list can be pre-constructed, and then after obtaining the target category and target boundary position point data of each detection target in the panoramic image through the neural network model, a data group corresponding to the detection target can be generated. The data group can specifically be an array including target category and target boundary position point data. Then, the data group corresponding to each detection target is filled into the blank preset target detection result list, so as to obtain the target detection result corresponding to the panoramic image to be detected in the form of a final list. In this embodiment, by filling the data group corresponding to each detection target into the preset target detection result list, the target detection result corresponding to the panoramic image to be detected is generated, which can effectively improve the intuitiveness of the target detection result.

[0088] It should be understood that although Figure 2-5 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-5 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0089] In one embodiment, Figure 6 As shown, a device for detecting an object in a panoramic image is provided, comprising:

[0090] The data acquisition module 601 is used to acquire the panoramic image to be detected.

[0091] The convolution processing module 603 is used to perform convolution processing on the panoramic image to be detected through a convolution neural network including a preset target deformation adaptive convolution operator, and obtain the target category and target boundary position point data of the detection target in the panoramic image to be detected, wherein the target boundary position point data includes the target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected.

[0092] The target detection module 605 is used to obtain the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detected target.

[0093] In one of the embodiments, the convolution processing module 603 is specifically used to: input the panoramic image to be detected into a convolutional neural network containing a preset target deformation adaptive convolution operator, obtain the heat map, target category and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected; filter the initial detection target according to the heat map to obtain the detection target; obtain the target category and target boundary position point data corresponding to the detection target.

[0094] In one of the embodiments, the initial detection target includes a non-boundary position target, the convolutional neural network also includes a conventional convolution operator, and the convolution processing module 603 is specifically used to: extract the panoramic image features of the panoramic image to be detected by using a conventional convolution operator; determine the non-boundary position targets in the panoramic image to be detected based on the panoramic image features, and obtain the heat map, target category and target boundary position point data corresponding to the non-boundary position targets.

[0095] In one of the embodiments, the convolution processing module 603 is specifically used to: extract panoramic image features of the panoramic image to be detected by a preset target deformation adaptive convolution operator; determine the boundary position target at the boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target; identify the target attributes between the first detection target and the second detection target based on the panoramic image features, the first detection target and the second detection target are the initial detection targets in relative positions in the panoramic image; when the target attributes characterize that the first detection target and the second detection target are the same detection target, obtain the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target, and obtain the heat map, target category and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected according to the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target.

[0096] In one of the embodiments, a model training module is also included, which is used to: obtain historical panoramic images with target category annotations and target boundary position point annotations, the target boundary position point data including target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; construct a model training data group based on the historical panoramic images; and train an initial convolutional neural network including a preset target deformation adaptive convolution operator through the model training data group to obtain a convolutional neural network including a preset target deformation adaptive convolution operator.

[0097] In one of the embodiments, the target detection module 605 is specifically used to: generate a data group corresponding to the detection target based on the target category of the detection target and the target boundary position point data; fill the data group corresponding to each detection target into a preset target detection result list to generate a target detection result corresponding to the panoramic image to be detected.

[0098] For the specific definition of the target detection device for panoramic images, please refer to the definition of the target detection method for panoramic images above, which will not be repeated here. Each module in the above-mentioned target detection device for panoramic images can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0099] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store traffic forwarding data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for detecting a target in a panoramic image is implemented.

[0100] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0101] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0102] Acquire the panoramic image to be detected;

[0103] The panoramic image to be detected is subjected to convolution processing by a convolutional neural network including a preset target deformation adaptive convolution operator, so as to obtain target categories and target boundary position point data of the detected target in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected;

[0104] According to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained.

[0105] In one embodiment, when the processor executes the computer program, the following steps are also implemented: inputting the panoramic image to be detected into a convolutional neural network containing a preset target deformation adaptive convolution operator, obtaining the heat map, target category and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected; filtering the initial detection target according to the heat map to obtain the detection target; obtaining the target category and target boundary position point data corresponding to the detection target.

[0106] In one embodiment, when the processor executes the computer program, the following steps are also implemented: extracting panoramic image features of the panoramic image to be detected by a conventional convolution operator; determining non-boundary position targets in the panoramic image to be detected based on the panoramic image features, and obtaining thermal maps, target categories, and target boundary position point data corresponding to the non-boundary position targets.

[0107] In one embodiment, when the processor executes the computer program, the following steps are also implemented: extracting panoramic image features of the panoramic image to be detected by a preset target deformation adaptive convolution operator; determining a boundary position target at a boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target; identifying target attributes between a first detection target and a second detection target based on the panoramic image features, the first detection target and the second detection target being initial detection targets at relative positions in the panoramic image; when the target attributes characterize that the first detection target and the second detection target are the same detection target, obtaining a heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target; obtaining a heat map, target category, and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected based on the heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target.

[0108] In one embodiment, when the processor executes the computer program, the following steps are also implemented: obtaining historical panoramic images with target category annotations and target boundary position point annotations, the target boundary position point data including target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; constructing a model training data group based on the historical panoramic images; and training an initial convolutional neural network including a preset target deformation adaptive convolution operator through the model training data group to obtain a convolutional neural network including a preset target deformation adaptive convolution operator.

[0109] In one embodiment, when the processor executes the computer program, the following steps are also implemented: generating a data group corresponding to the detection target based on the target category of the detection target and the target boundary position point data; filling the data group corresponding to each detection target into a preset target detection result list to generate a target detection result corresponding to the panoramic image to be detected.

[0110] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0111] Acquire the panoramic image to be detected;

[0112] A convolutional neural network including a preset target deformation adaptive convolution operator is used to perform convolution processing on the panoramic image to be detected, and the target category and target boundary position point data of the detection target in the panoramic image to be detected are obtained. The target boundary position point data include target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected. According to the target category and target boundary position point data of the detection target, the target detection result corresponding to the panoramic image to be detected is obtained.

[0113] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: inputting the panoramic image to be detected into a convolutional neural network containing a preset target deformation adaptive convolution operator to obtain the heat map, target category and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected; filtering the initial detection target according to the heat map to obtain the detection target; obtaining the target category and target boundary position point data corresponding to the detection target.

[0114] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented: extracting panoramic image features of the panoramic image to be detected by a conventional convolution operator; determining non-boundary position targets in the panoramic image to be detected based on the panoramic image features, and obtaining thermal maps, target categories, and target boundary position point data corresponding to the non-boundary position targets.

[0115] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented: extracting panoramic image features of the panoramic image to be detected by a preset target deformation adaptive convolution operator; determining a boundary position target at a boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target; identifying target attributes between a first detection target and a second detection target based on the panoramic image features, the first detection target and the second detection target being initial detection targets at relative positions in the panoramic image; when the target attributes characterize that the first detection target and the second detection target are the same detection target, obtaining a heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target; obtaining a heat map, target category, and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected based on the heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target.

[0116] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented: obtaining historical panoramic images with target category annotations and target boundary position point annotations, the target boundary position point data including target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; constructing a model training data group based on the historical panoramic images; and training an initial convolutional neural network including a preset target deformation adaptive convolution operator through the model training data group to obtain a convolutional neural network including a preset target deformation adaptive convolution operator.

[0117] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: generating a data group corresponding to the detection target based on the target category of the detection target and the target boundary position point data; filling the data group corresponding to each detection target into a preset target detection result list to generate a target detection result corresponding to the panoramic image to be detected.

[0118] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0119] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for detecting an object in a panoramic image, the method comprising: Acquire the panoramic image to be detected; The panoramic image to be detected is convolved by a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a target category and target boundary position point data of the detection target in the panoramic image to be detected, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected, and the preset target deformation adaptive convolution operator is obtained by replacing some convolution operators of the convolutional neural network with convolution operators that can adapt to target deformation, including convolution operators of deformable convolution, equirectangular projection convolution and spherical convolution, and the preset target deformation adaptive convolution operator is used to determine whether the targets on the left and right sides of the panoramic image to be detected are the same target, and output a corresponding target category and a set of target boundary position point data for the same target; According to the target category and target boundary position point data of the detected target, the target detection result corresponding to the panoramic image to be detected is obtained.

2. The method according to claim 1, characterized in that The convolution processing is performed on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detection target in the panoramic image to be detected, including: Inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator, and obtaining a heat map, target category, and target boundary position point data corresponding to an initial detection target in the panoramic image to be detected; Filtering the initial detection target according to the thermal map to obtain the detection target; Obtain the target category corresponding to the detected target and target boundary position point data.

3. The method according to claim 2, characterized in that The initial detection target includes a non-boundary position target, and the convolutional neural network also includes a conventional convolution operator; The step of inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a heat map, target category, and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected includes: Extracting panoramic image features of the panoramic image to be detected by using the conventional convolution operator; Based on the panoramic image features, non-boundary position targets in the panoramic image to be detected are determined, and a heat map, target category, and target boundary position point data corresponding to the non-boundary position targets are obtained.

4. The method according to claim 2, characterized in that: The initial detection target includes a boundary position target; The step of inputting the panoramic image to be detected into a convolutional neural network including a preset target deformation adaptive convolution operator to obtain a heat map, target category, and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected includes: Extracting panoramic image features of the panoramic image to be detected by using a preset target deformation adaptive convolution operator; Determine a boundary position target at a boundary position in the panoramic image to be detected based on the panoramic image features to obtain an initial detection target; identifying target attributes between a first detection target and a second detection target based on the panoramic image feature, wherein the first detection target and the second detection target are initial detection targets at relative positions in the panoramic image; When the target attribute indicates that the first detection target and the second detection target are the same detection target, obtaining a heat map, target category, and target boundary position point data corresponding to the first detection target and the second detection target; According to the heat map, target category and target boundary position point data corresponding to the first detection target and the second detection target, the heat map, target category and target boundary position point data corresponding to the boundary position target in the panoramic image to be detected are obtained.

5. The method according to claim 1, characterized in that Before the convolution processing is performed on the panoramic image to be detected by a convolution neural network including a preset target deformation adaptive convolution operator to obtain the target category and target boundary position point data of the detected target in the panoramic image to be detected, the method further includes: Acquire a historical panoramic image with target category annotations and target boundary position point annotations, wherein the target boundary position point data includes target boundary position points whose coordinates exceed the boundary of the panoramic image to be detected; Constructing a model training data set based on the historical panoramic images; The initial convolutional neural network including the preset target deformation adaptive convolution operator is trained through the model training data group to obtain the convolutional neural network including the preset target deformation adaptive convolution operator.

6. The method according to claim 1, characterized in that The acquiring the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detected target comprises: Generate a data group corresponding to the detected target according to the target category and target boundary position point data of the detected target; The data group corresponding to the detected target is filled into the preset target detection result list to generate the target detection result corresponding to the panoramic image to be detected.

7. A target detection device for panoramic images, characterized in that: The device comprises: A data acquisition module, used to acquire the panoramic image to be detected; A convolution processing module, used to perform convolution processing on the panoramic image to be detected through a convolution neural network including a preset target deformation adaptive convolution operator, to obtain a target category and target boundary position point data of a detection target in the panoramic image to be detected, wherein the target boundary position point data includes a target boundary position point whose coordinates exceed the boundary of the panoramic image to be detected, and the preset target deformation adaptive convolution operator is obtained by replacing some convolution operators of the convolution neural network with convolution operators that can adapt to target deformation, including convolution operators of deformable convolution, equirectangular projection convolution and spherical convolution, and the preset target deformation adaptive convolution operator is used to determine whether the targets on the left and right sides of the panoramic image to be detected are the same target, and output a corresponding target category and a set of target boundary position point data for the same target; The target detection module is used to obtain the target detection result corresponding to the panoramic image to be detected according to the target category and target boundary position point data of the detection target.

8. The device according to claim 7, characterized in that The convolution processing module is specifically used to: input the panoramic image to be detected into a convolutional neural network containing a preset target deformation adaptation convolution operator, obtain the heat map, target category and target boundary position point data corresponding to the initial detection target in the panoramic image to be detected; filter the initial detection target according to the heat map to obtain the detection target; Obtain the target category corresponding to the detected target and target boundary position point data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Panoramic image target detection method based on spherical projection grid and spherical convolution

    CN110163271A

  • Target tracking method of panoramic video, readable storage medium and computer equipment

    CN111242977A

  • Image detection method and device and computer readable storage medium

    CN111402228A

  • Method and device for generating virtual reality data

    US20210150683A1