Image processing method and related device

By using spherical coordinate system and field of view angle to calculate the area of ​​the target intersection area in the panoramic image, the error problem of bounding box intersection ratio calculation in the panoramic image is solved, and the accuracy of detection is improved.

CN114140512BActive Publication Date: 2025-08-29HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110932912.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-13
Publication Date
2025-08-29
Estimated Expiration
2041-08-13

AI Technical Summary

Technical Problem

In panoramic images, it is difficult for the prior art to accurately calculate the intersection ratio of bounding boxes, especially at the poles of the image, which affects the accuracy of object detection.

Method used

By directly calculating the spherical area of ​​the target intersection area of ​​the first bounding box and the second bounding box on the panoramic image, the intersection ratio is determined, and the position information and field angle under the spherical coordinate system are used to improve the calculation accuracy.

Benefits of technology

This greatly reduces errors, improves the accuracy of bounding box interchange ratio, and enhances the accuracy of panoramic image object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140512B_ABST
    Figure CN114140512B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method and related equipment. The method can be used in the field of image processing in the field of artificial intelligence. The method includes: generating a spherical area of ​​a target intersection area between the first and second bounding boxes on a first panoramic image based on first position information of a first bounding box and second position information of a second bounding box, wherein the first position information indicates the spherical coordinates of the vertices of the first bounding box on the first panoramic image, and the second position information indicates the spherical coordinates of the vertices of the second bounding box on the first panoramic image, and the second bounding box is a desired bounding box corresponding to the first bounding box; and determining an intersection-over-union ratio (IoU) of the first and second bounding boxes on the first panoramic image based on the spherical area of ​​the target intersection area. Directly calculating the IoU ratio of the first and second bounding boxes on the sphere significantly reduces error and improves the accuracy of the present solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to an image processing method and related equipment. Background Art

[0002] With the development of science and technology, a large number of panoramic cameras have been developed in recent years, which can capture panoramic images with a 360-degree field of view. Panoramic image technology can be applied to industries such as autonomous driving, virtual navigation or video surveillance.

[0003] In order to determine the accuracy of a first bounding box generated when performing object detection on a panoramic image, it is necessary to calculate the intersection over union (IoU) between the generated first bounding box and a desired second bounding box. Specifically, an unfolded image of the panoramic image is obtained, the position of the first bounding box on the unfolded image is determined, and the position of the second bounding box on the unfolded image is determined. The IoU between the bounding rectangle of the first bounding box and the bounding rectangle of the second bounding box is determined on the unfolded image, and the IoU between the two bounding rectangles is determined as the IoU between the first and second bounding boxes.

[0004] However, in the cylindrical unfolded image, the bounding box's circumscribed rectangle cannot accurately reflect the bounding box's area, and the closer the bounding box is to the two poles of the panoramic image, the greater the error between the bounding box and the bounding box's circumscribed rectangle. Therefore, a more accurate method for calculating the intersection-over-union ratio between the first bounding box and the second bounding box is urgently needed. Summary of the Invention

[0005] The embodiments of the present application provide an image processing method and related equipment for directly calculating the spherical area of ​​a target intersection region of a first bounding box and a second bounding box on a first panoramic image, and calculating the intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection region. That is, directly calculating the intersection-over-union ratio of the first bounding box and the second bounding box on the sphere greatly reduces errors and improves the accuracy of the present solution.

[0006] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0007] In a first aspect, embodiments of the present application provide an image processing method that can be used in the field of artificial intelligence for object detection in panoramic images. The method includes: an electronic device performing object detection on a first panoramic image using a first model to obtain a detection result corresponding to the first panoramic image. The detection result includes obtaining first position information of a first bounding box on the first panoramic image. The first position information indicates the spherical coordinates of the vertices of the first bounding box on the first panoramic image. The spherical coordinates include spherical coordinates in the longitude direction and spherical coordinates in the latitude direction. The electronic device obtains second position information of a second bounding box on the first panoramic image. The second position information indicates the spherical coordinates of the vertices of the second bounding box on the first panoramic image. The second bounding box may be a desired bounding box corresponding to the first bounding box. Based on the first and second position information, the electronic device generates a spherical area of ​​a target intersection region between the first and second bounding boxes on the first panoramic image. The electronic device then determines an intersection-over-union ratio (IoU) of the first and second bounding boxes on the first panoramic image based on the spherical area of ​​the target intersection region, the spherical area of ​​the first and second bounding boxes, and the spherical area of ​​the first and second bounding boxes. This is to generate an IoU ratio of the first and second bounding boxes on the spherical surface.

[0008] In this implementation, the spherical area of ​​the target intersection region of the first bounding box and the second bounding box on the first panoramic image is directly calculated, and the intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image is calculated based on the spherical area of ​​the target intersection region. That is, the intersection-over-union ratio of the first bounding box and the second bounding box on the spherical surface is directly calculated, which greatly reduces the error and improves the accuracy of this solution.

[0009] In a possible implementation of the first aspect, the first bounding box is a rectangle, and the first position information includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, the first field of view angle and the second field of view angle corresponding to the first bounding box, the first field of view angle being the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle being the field of view angle corresponding to the edge of the first bounding box in the latitude direction; wherein the spherical coordinates of the center point of the first bounding box on the first panoramic image, the first field of view angle and the second field of view angle corresponding to the first bounding box are used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image.

[0010] In this implementation, when the first bounding box is limited to a rectangle, the first position information may include the spherical coordinates of the center point of the first bounding box on the first panoramic image, the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the field of view angle corresponding to the edge of the first bounding box in the latitude direction. Compared with directly obtaining the spherical coordinates of the vertices of any polygon, the aforementioned method can reduce the difficulty of the process of obtaining the first position information, which is conducive to improving the accuracy of the obtained first position information, and further conducive to improving the accuracy of the intersection-union ratio finally obtained.

[0011] In a possible implementation of the first aspect, the first position information of the first bounding box includes the spherical coordinates of each vertex of the first bounding box on the first panoramic image, and the second position information of the second bounding box includes the spherical coordinates of each vertex of the second bounding box on the first panoramic image.

[0012] In a possible implementation of the first aspect, the electronic device obtaining first position information of a first bounding box on a first panoramic image may include: the electronic device inputting the first panoramic image into a first model to perform target detection on the first panoramic image using the first model. The electronic device generates target indication information using a first branch of the first model, where the target indication information is used to indicate the spherical coordinates of a center point of the first bounding box on the first panoramic image. The electronic device generates first and second field of view angles corresponding to the center point of the first bounding box using a second branch of the first model, thereby obtaining the first and second field of view angles corresponding to the first bounding box.

[0013] In a possible implementation of the first aspect, if the first bounding box and the second bounding box intersect, the electronic device generates a spherical area of ​​the intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information, which may include: the electronic device generates the spherical coordinates of each vertex of the target intersection area on the first panoramic image based on the spherical coordinates of each vertex of the first bounding box on the first panoramic image and the spherical coordinates of each vertex of the second bounding box on the first panoramic image; and generates the spherical area of ​​the target intersection area based on the spherical coordinates of each vertex of the target intersection area on the first panoramic image and the area calculation principle of a spherical polygon.

[0014] This implementation provides a specific implementation method for calculating the spherical area of ​​the target intersection region when the first bounding box and the second bounding box intersect. The operation is simple and easy to implement.

[0015] In a possible implementation of the first aspect, the electronic device may generate spherical coordinates of vertices of the target intersection area on the first panoramic image based on the first position information and the second position information. This may include: the electronic device determining a first intersection point set based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, based on the principle of calculating the intersection between two edges on a sphere. The first intersection point set includes the spherical coordinates of multiple first intersection points, where a first intersection point is an intersection point between any two edges of the first and second bounding boxes. The electronic device then combines all the first intersection points, all the vertices of the first bounding box, and all the vertices of the second bounding box into a second point set. The electronic device selects multiple second points from the second point set and determines the spherical coordinates of each second point, where the second point is a point in the second point set that is located in the target intersection area. Because a first intersection point may overlap with a vertex of the first bounding box or a vertex of the second bounding box, the electronic device removes duplicate points from the multiple second points to obtain all vertices of the target intersection area, thereby obtaining the spherical coordinates of each vertex of the target intersection area on the first panoramic image.

[0016] In a possible implementation of the first aspect, the electronic device determines, based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, that the intersection between the first bounding box and the second bounding box is empty, and then determines the area of ​​the target intersection region to be 0. Alternatively, if the electronic device determines, based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, that the first bounding box is located inside the second bounding box, then the area of ​​the target intersection region is determined to be the spherical area of ​​the first bounding box. Alternatively, if the electronic device determines, based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, that the second bounding box is located inside the first bounding box, then the area of ​​the target intersection region is determined to be the spherical area of ​​the second bounding box.

[0017] In one possible implementation of the first aspect, the method is applied during the training phase of the first model and / or the method is applied to detect the accuracy of a first bounding box generated by the first model. This implementation provides two specific application scenarios for this solution, increasing its implementation flexibility.

[0018] In a second aspect, an embodiment of the present application provides an image processing method that can be used in the field of artificial intelligence for target detection in panoramic images. The method includes: an electronic device obtaining first position information of a first bounding box on a first panoramic image, the first position information being used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; obtaining second position information of a second bounding box on the first panoramic image, the second position information being used to indicate the spherical coordinates of the vertices of the second bounding box on the first panoramic image; generating a spherical area of ​​a target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; and determining an intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area.

[0019] In the second aspect of the embodiment of the present application, the electronic device can also execute the steps executed by the electronic device in each possible implementation method of the first aspect. For the specific implementation steps of the second aspect of the embodiment of the present application and the various possible implementation methods of the second aspect, as well as the beneficial effects brought about by each possible implementation method, you can refer to the description of the various possible implementation methods in the first aspect, and will not go into details here.

[0020] On the third aspect, an embodiment of the present application provides a model for image processing, which can be used in the field of target detection on panoramic images in the field of artificial intelligence. The model is a second model for target detection of C-type objects in panoramic images, where C is an integer greater than or equal to 1, and the second model includes a first branch and a second branch. The second model is used to receive a second panoramic image to perform target detection on the second panoramic image; the first branch is used to generate C indication information corresponding to category C, and the target indication information is the indication information corresponding to the target category in category C among the C indication information. The target indication information is used to indicate the spherical coordinates of the center point of one or more target bounding boxes on the second panoramic image, and the predicted category of each target bounding box is the target category. The spherical coordinates include spherical coordinates in the longitude direction and spherical coordinates in the latitude direction.

[0021] The second branch is used to generate the first field of view angle and the second field of view angle corresponding to each pixel point of the second panoramic image. The first field of view angle and the second field of view angle corresponding to each pixel point of the second panoramic image may include the first field of view angle and the second field of view angle corresponding to the center point of the target bounding box. The first field of view angle and the second field of view angle corresponding to the center point of the target bounding box can also be determined as the first field of view angle and the second field of view angle corresponding to the target bounding box. The first field of view angle corresponding to the center point of the target bounding box is the field of view angle corresponding to the edge of the target bounding box in the longitude direction, and the second field of view angle corresponding to the center point of the target bounding box is the field of view angle corresponding to the edge of the target bounding box in the latitude direction. The target indication information, the first field of view angle and the second field of view angle corresponding to the center point of the target bounding box are used to indicate the position of the target bounding box on the second panoramic image and the predicted category of the target bounding box.

[0022] In this implementation, the first model can directly obtain the spherical coordinates of the center point of each target bounding box in the second panoramic image, the predicted category of each target bounding box, and the first field of view angle and the second field of view angle corresponding to each target bounding box, providing a simpler model for target detection and improving the efficiency of the target detection process.

[0023] In a possible implementation of the third aspect, the second model is used to perform target detection on C-type objects in the second panoramic image, where C is an integer greater than or equal to 1, the target indication information corresponds to the target category in category C, and the target indication information includes multiple target probability values ​​corresponding one-to-one to multiple pixel points in the first panoramic image. A target probability value is used to indicate the probability that a pixel point in the first panoramic image belongs to the target category, wherein the larger the target probability value corresponding to a pixel point, the greater the probability that the pixel point is determined as the center point of the bounding box.

[0024] In a possible implementation of the third aspect, the target indication information may be specifically expressed as a heat map corresponding to the first panoramic image, and a position (x, y) of the heat map corresponds to a score p. xyc , the fraction p xyc represents the probability that the predicted category of the pixel at position (x, y) is the target category, x represents the coordinate of the pixel in the longitude direction in the spherical coordinate system corresponding to the first panoramic image, and y represents the coordinate of the pixel in the latitude direction in the spherical coordinate system corresponding to the first panoramic image. xyc The value range is 0-1, p xyc The higher the value of , the greater the probability that the pixel at position (x, y) belongs to the target category.

[0025] In a possible implementation of the third aspect, the second model also includes a third branch; the third branch is used to generate a first offset and a second offset corresponding to the center point of the target bounding box; wherein the first offset is a spherical offset in the longitude direction, the second offset is a spherical offset in the latitude direction, the spherical coordinates of the center point of the target bounding box on the second panoramic image, and the first offset and the second offset are used to indicate the spherical coordinates of the updated center point of the target bounding box on the second panoramic image.

[0026] In this implementation, since the electronic device needs to perform a downsampling operation in the process of generating each target indication information through the first model, a discretization error will occur when generating the spherical coordinates of the center point of the bounding box based on the target indication information, that is, there will be an error between the spherical coordinates of the center point obtained based on the target indication information generated by the first branch and the spherical coordinates of the actual center point. Updating the spherical coordinates of the center point generated by the first branch according to the first offset and the second offset generated by the third branch is conducive to obtaining more accurate spherical coordinates of the center point.

[0027] Fourthly, embodiments of the present application provide a model training method that can be used in the field of artificial intelligence for object detection in panoramic images. The method is used to train a second model, the second model including a first branch and a second branch. The method includes: an electronic device inputting a third panoramic image into the second model, and performing object detection on the third panoramic image using the second model; generating first indication information through the first branch, the first indication information being used to indicate the spherical coordinates of the center point of the predicted bounding box on the third panoramic image, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction.

[0028] The electronic device generates a first predicted field of view angle and a second predicted field of view angle corresponding to the center point of the predicted bounding box through the second branch. The first predicted field of view angle is the field of view angle corresponding to the edge of the predicted bounding box in the longitude direction, and the second predicted field of view angle is the field of view angle corresponding to the edge of the predicted bounding box in the dimensional direction. The first indication information, the first predicted field of view angle and the second predicted field of view angle are all included in the position information of the predicted bounding box.

[0029] The electronic device obtains position information of an expected bounding box corresponding to the predicted bounding box on the third panoramic image, the position information of the expected bounding box including second indication information, a first expected field of view angle, and a second expected field of view angle, the second indication information being used to indicate the spherical coordinates of a center point of the expected bounding box on the third panoramic image, the first expected field of view angle being the field of view angle corresponding to an edge of the expected bounding box in the longitude direction, and the second expected field of view angle being the field of view angle corresponding to an edge of the expected bounding box in the dimensional direction.

[0030] The electronic device trains the second model based on the position information of the predicted bounding box, the position information of the expected bounding box and the target loss function to obtain a trained second model; wherein the target loss function includes a first loss term and a second loss term, the first loss term is used to indicate the similarity between the first indication information and the second indication information, the second loss term is used to indicate the similarity between the first predicted field of view angle and the first expected field of view angle, and the second loss term is also used to indicate the similarity between the second predicted field of view angle and the second expected field of view angle.

[0031] In a possible implementation of the fourth aspect, the second model is used to perform target detection on C-type objects in the third panoramic image, where C is an integer greater than or equal to 1, and the first indication information corresponds to the target category in category C. The first indication information includes multiple target probability values ​​corresponding one-to-one to multiple pixel points in the first panoramic image, and a target probability value is used to indicate the probability that a pixel point in the first panoramic image belongs to the target category, wherein the larger the target probability value corresponding to a pixel point, the greater the probability that the pixel point is determined as the center point of the bounding box.

[0032] For the specific implementation steps of the fourth aspect of the embodiment of the present application and the various possible implementation methods of the fourth aspect, as well as the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the third aspect, and will not be repeated here one by one.

[0033] In a fifth aspect, an embodiment of the present application provides an image processing device that can be used in the field of target detection on panoramic images in the field of artificial intelligence. The image processing device includes: a detection module that performs target detection on a first panoramic image through a first model to obtain a detection result corresponding to the first panoramic image, the detection result including first position information of a first bounding box on the first panoramic image, the first position information being used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; an acquisition module that acquires second position information of a second bounding box on the first panoramic image, the second bounding box being a desired bounding box corresponding to the first bounding box, the second position information being used to indicate the spherical coordinates of the vertices of the second bounding box on the first panoramic image; a generation module that generates a spherical area of ​​a target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; and a determination module that determines an intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area.

[0034] The image processing device provided in the fifth aspect of the embodiment of the present application can also execute the steps executed by the electronic device in each possible implementation method of the first aspect. For the specific implementation steps of the fifth aspect of the embodiment of the present application and the various possible implementation methods of the fifth aspect, as well as the beneficial effects brought about by each possible implementation method, you can refer to the description of the various possible implementation methods in the first aspect, and will not repeat them one by one here.

[0035] In a sixth aspect, an embodiment of the present application provides an image processing device that can be used in the field of target detection on panoramic images in the field of artificial intelligence. The image processing device includes: an acquisition module for acquiring first position information of a first bounding box on a first panoramic image, the first position information being used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image, the spherical coordinates including the spherical coordinates in the longitude direction and the spherical coordinates in the latitude direction; the acquisition module is also used to acquire second position information of a second bounding box on the first panoramic image, the second position information being used to indicate the spherical coordinates of the vertices of the second bounding box on the first panoramic image; a generation module for generating the spherical area of ​​a target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; a determination module for determining the intersection-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area.

[0036] The image processing device provided in the sixth aspect of the embodiment of the present application can also execute the steps executed by the electronic device in each possible implementation method of the first aspect. For the specific implementation steps of the sixth aspect of the embodiment of the present application and the various possible implementation methods of the sixth aspect, as well as the beneficial effects brought about by each possible implementation method, you can refer to the description of the various possible implementation methods in the first aspect, and will not repeat them one by one here.

[0037] In a seventh aspect, an embodiment of the present application provides an image processing device that can be used in the field of performing target detection on panoramic images in the field of artificial intelligence. The image processing device is configured with a second model, the second model including a first branch and a second branch, and the image processing device includes an input module and a processing module, wherein the input module is used to input the second panoramic image into the second model; the processing module is used to generate target indication information through the first branch, the target indication information is used to indicate the spherical coordinates of the center point of the target bounding box on the second panoramic image, and the predicted category of the target bounding box, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; the processing module is further used to generate a first field of view angle and a second field of view angle corresponding to the center point of the target bounding box through the second branch, the first field of view angle being the field of view angle corresponding to the edge of the target bounding box in the longitude direction, and the second field of view angle being the field of view angle corresponding to the edge of the target bounding box in the latitude direction; wherein the target indication information and the first and second field of view angles corresponding to the center point of the target bounding box are used to indicate the position of the target bounding box on the second panoramic image and the predicted category of the target bounding box.

[0038] The image processing device provided in the seventh aspect of the embodiment of the present application can also execute the steps executed by the second model in each possible implementation method of the second aspect. For the specific implementation steps of the seventh aspect of the embodiment of the present application and the various possible implementation methods of the seventh aspect, as well as the beneficial effects brought about by each possible implementation method, you can refer to the description of the various possible implementation methods in the third aspect, and will not go into details here.

[0039] In an eighth aspect, an embodiment of the present application provides an electronic device, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the image processing method described in the first or second aspect above is implemented; or, when the program instructions stored in the memory are executed by the processor, the steps performed by the second model in the third aspect above are implemented.

[0040] In the ninth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the program is run on a computer, the computer executes the image processing method described in the first or second aspect above, or the computer executes the steps performed by the second model in the third aspect above.

[0041] In the tenth aspect, an embodiment of the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the image processing method described in the first or second aspect above, or the processing circuit is configured to execute the steps performed by the second model in the third aspect above.

[0042] In the eleventh aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, enables the computer to execute the image processing method described in the first or second aspect above, or enables the computer to execute the steps performed by the second model in the third aspect above.

[0043] In a twelfth aspect, an embodiment of the present application provides a chip system, which includes a processor for implementing the functions involved in the above-mentioned various aspects, for example, sending or processing the data and / or information involved in the above-mentioned method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of a chip, or it can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1a A schematic diagram of the structure of the artificial intelligence main framework provided in the embodiment of the present application;

[0045] Figure 1b An application scenario diagram of the image processing method provided in an embodiment of the present application;

[0046] Figure 1c Another application scenario diagram of the image processing method provided in the embodiment of the present application;

[0047] Figure 2 A schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the longitude and latitude directions in the image processing method provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the first field of view angle and the second field of view angle corresponding to the first bounding box in the image processing method provided in an embodiment of the present application;

[0051] Figure 6 A schematic diagram of the first model provided in an embodiment of the present application;

[0052] Figure 7 Another schematic diagram of the first model provided in an embodiment of the present application;

[0053] Figure 8 Three schematic diagrams of the intersection of a first bounding box and a second bounding box provided in an embodiment of the present application;

[0054] Figure 9 A system architecture diagram of an image processing system provided in an embodiment of the present application;

[0055] Figure 10 A schematic structural diagram of the second model provided in an embodiment of the present application;

[0056] Figure 11 Another structural diagram of the second model provided in an embodiment of the present application;

[0057] Figure 12 A flowchart of a neural network training method provided in an embodiment of the present application;

[0058] Figure 13 A schematic diagram of generating a Gaussian kernel radius in the neural network training method provided in an embodiment of the present application;

[0059] Figure 14 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;

[0060] Figure 15 Another structural diagram of the image processing device provided in an embodiment of the present application;

[0061] Figure 16 Another structural diagram of the image processing device provided in an embodiment of the present application;

[0062] Figure 17 This is a structural diagram of an electronic device provided in an embodiment of the present application;

[0063] Figure 18 Another structural diagram of an electronic device provided in an embodiment of the present application;

[0064] Figure 19 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0066] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0067] First, the overall workflow of the artificial intelligence system is described. Figure 1a , Figure 1a The following diagram illustrates a structural diagram of the AI ​​framework. This framework is explained below from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it encompasses the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed progression from "data-information-knowledge-wisdom." The "IT value chain," encompassing the entire process from the underlying infrastructure of human intelligence, information (provided and processed by technology), to the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0068] (1) Infrastructure

[0069] The infrastructure provides computing power support for artificial intelligence systems, enabling communication with the outside world and providing support through the basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically adopt hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs); the basic platform includes related platform guarantees and support such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to obtain data, and this data is provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.

[0070] (2) Data

[0071] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0072] (3) Data processing

[0073] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0074] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0075] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0076] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0077] (4) General ability

[0078] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0079] (5) Smart products and industry applications

[0080] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart manufacturing, smart transportation, smart homes, smart medical care, autonomous driving, smart cities, etc.

[0081] The present application can be applied to various application fields that use panoramic image technology. For example, panoramic images can be used in smart manufacturing, intelligent transportation, smart cities, smart homes, autonomous driving and other fields.

[0082] As an example, to more intuitively understand the application field of the embodiment of the present application, please refer to Figure 1b , Figure 1b An application scenario diagram of the image processing method provided in an embodiment of the present application is shown. Figure 1b is an expanded view of a panoramic image of a factory building. In the field of intelligent manufacturing, a panoramic image of a factory building can be obtained to monitor the operating status of each device in the factory building in real time, so as to locate the faulty equipment in a timely manner.

[0083] As another example, see Figure 1c , Figure 1cAn application scenario diagram of the image processing method provided in an embodiment of the present application is shown. Figure 1c , which is an expanded view of a panoramic image of a traffic road. In the field of intelligent transportation, a panoramic image of a traffic road can be obtained, and target detection can be performed on the aforementioned panoramic image to facilitate real-time understanding of vehicle conditions on the traffic road.

[0084] As another example, in the field of autonomous driving, an autonomous vehicle obtains a panoramic image of the surrounding environment through a panoramic camera and performs target detection on the aforementioned panoramic image, so that it can detect surrounding objects in 360° without blind spots, so as to improve the safety and reliability of the autonomous driving process. The application scenarios of panoramic images are not exhaustively listed here.

[0085] As another example, in the field of smart terminals, a smart robot can be equipped with a panoramic camera to obtain a panoramic image of the surrounding environment through the panoramic camera, and perform target detection on the aforementioned panoramic image to facilitate understanding of objects in the robot's surrounding environment, etc. The application fields of the embodiments of the present application are not exhaustively listed here.

[0086] In the above-mentioned application fields, after obtaining a panoramic image, the model can be used to perform target detection on the panoramic image to obtain a detection result. The detection result includes the position information of one or more bounding boxes and the predicted category corresponding to each bounding box. The position information of a bounding box is used to reflect the position of a target object in the panoramic image.

[0087] During the training phase of the above-mentioned model, or when testing the accuracy of the detection results output by the above-mentioned model, it may be necessary to calculate the intersection-over-union ratio between two bounding boxes, where the two bounding boxes include a first bounding box and a second bounding box. The first bounding box is a predicted bounding box generated by the model, and the second bounding box is an expected bounding box corresponding to the predicted bounding box. The expected bounding box may also be referred to as a correct bounding box or a labeled bounding box, etc.

[0088] In order to more accurately calculate the intersection-over-union ratio between two bounding boxes on a panoramic image, the present application embodiment provides an image processing method, see Figure 2 , Figure 2A flowchart of an image processing method provided by an embodiment of the present application is provided. A1: An electronic device performs target detection on a first panoramic image using a first model to obtain a detection result corresponding to the first panoramic image. The detection result includes first position information of a first bounding box on the first panoramic image. The first position information indicates the spherical coordinates of each vertex of the first bounding box on the first panoramic image. The spherical coordinates include spherical coordinates in the longitude and latitude directions. A2: The electronic device obtains second position information of a second bounding box on the first panoramic image. The second bounding box is a desired bounding box corresponding to the first bounding box. The second position information indicates the spherical coordinates of each vertex of the second bounding box on the first panoramic image. A3: The electronic device generates the spherical area of ​​a target intersection region between the first and second bounding boxes on the first panoramic image based on the first and second position information. A4: The electronic device determines an intersection-over-union ratio (IoU) of the first and second bounding boxes on the first panoramic image based on the spherical area of ​​the target intersection region. Specifically, the IoU ratio of the first and second bounding boxes on the first panoramic image is determined based on the spherical area of ​​the first bounding box, the spherical area of ​​the second bounding box, and the spherical area of ​​the target intersection region.

[0089] In the embodiment of the present application, the spherical area of ​​the target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image is calculated based on the spherical coordinates of each vertex of the first bounding box on the first panoramic image and the spherical coordinates of each vertex of the second bounding box on the first panoramic image, and then the intersection-and-union ratio of the first bounding box and the second bounding box on the first panoramic image is calculated, that is, the intersection-and-union ratio of the first bounding box and the second bounding box on the spherical surface is directly calculated, which greatly reduces the error and improves the accuracy of the present solution.

[0090] For details, please refer to Figure 3 , Figure 3 This is a flow chart of an image processing method provided in an embodiment of the present application. The image processing method provided in an embodiment of the present application may include:

[0091] 301. The electronic device obtains first position information of a first bounding box on a first panoramic image, where the first position information is used to indicate the spherical coordinates of each vertex of the first bounding box on the first panoramic image.

[0092] In an embodiment of the present application, after the electronic device obtains the first panoramic image, it can input the first panoramic image into the first model to perform target detection on the first panoramic image through the first model to obtain a detection result output by the first model, which includes at least one first position information corresponding one-to-one to at least one first bounding box on the first panoramic image, and a predicted category corresponding to each first bounding box.

[0093] Among them, the first model can be specifically expressed as a neural network for performing target detection on panoramic images. As an example, the first model can adopt convolutional neural networks (CNN); the first model can also be expressed as a model that is not a neural network, which is not limited here.

[0094] The first panoramic image is presented as a spherical image; a first bounding box on the first panoramic image is used to indicate the position of a target object in the first panoramic image; the first position information is used to indicate the spherical coordinates of each vertex of the first bounding box on the first panoramic image. Spherical coordinates are coordinates established based on a spherical coordinate system, the origin of which is the center of the spherical first panoramic image, and the spherical coordinates include spherical coordinates in the longitude direction and spherical coordinates in the latitude direction. For a more intuitive understanding of this solution, please refer to Figure 4 , Figure 4 A schematic diagram of the longitude and latitude directions in the image processing method provided in an embodiment of the present application, Figure 4 Including two sub-schematic diagrams (a) and (b), Figure 4 The schematic diagram (a) shows an example of a panoramic image. Figure 4 The schematic diagram (b) shows the origin of the spherical coordinate system, the longitude direction under the spherical coordinate system, and the latitude direction under the spherical coordinate system. It should be understood that Figure 4 This is only for facilitating understanding of this solution and is not intended to limit this solution.

[0095] Furthermore, in one implementation, the first position information output by the first model includes the spherical coordinates of each vertex of the first bounding box on the first panoramic image, so there is no need to define the shape of the first bounding box. After obtaining the at least one piece of first position information output by the first model that corresponds one-to-one with at least one first bounding box, the electronic device can directly obtain the spherical coordinates of each vertex of each first bounding box on the first panoramic image.

[0096] In another implementation, the first position information output by the first model includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, the first field of view angle corresponding to the first bounding box, and the second field of view angle. The first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction. Each first bounding box needs to be rectangular. For a more intuitive understanding of this solution, please refer to Figure 5 , Figure 5 A schematic diagram of the first field of view angle and the second field of view angle corresponding to the first bounding box in the image processing method provided in an embodiment of the present application, wherein B1 represents the first field of view angle corresponding to the first bounding box, and B2 represents the second field of view angle corresponding to the first bounding box. It should be understood that Figure 5 This is only for facilitating understanding of this solution and is not intended to limit this solution.

[0097] After the electronic device obtains at least one first position information output by the first model that corresponds one-to-one to at least one first bounding box, for any one of the at least one first bounding box, the electronic device can generate the spherical coordinates of each vertex of the first bounding box on the first panoramic image based on the spherical coordinates of the center point of the first bounding box on the first panoramic image, the above-mentioned first field of view angle and the above-mentioned second field of view angle.

[0098] In an embodiment of the present application, when the first bounding box is limited to a rectangle, the first position information may include the spherical coordinates of the center point of the first bounding box on the first panoramic image, the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the field of view angle corresponding to the edge of the first bounding box in the latitude direction. Compared with directly obtaining the spherical coordinates of the vertices of any polygon, the aforementioned method can reduce the difficulty of the process of obtaining the first position information, which is conducive to improving the accuracy of the obtained first position information, and further helps to improve the accuracy of the intersection-union ratio finally obtained.

[0099] The function of the first model in this implementation is to perform target detection on C types of objects in a panoramic image, where C is a positive integer greater than or equal to 1, and the first model may include a first branch and a second branch. Specifically, the electronic device inputs the first panoramic image into the first model to perform target detection on the first panoramic image through the first model. The electronic device can generate C indication information corresponding one-to-one to C types of objects through the first branch in the first model, wherein the C indication information is used to indicate the spherical coordinates of the center point of each first bounding box in at least one first bounding box corresponding to the first panoramic image on the first panoramic image, and the predicted category of each first bounding box.

[0100] Furthermore, for any one of the C indication information (hereinafter referred to as "target indication information" for the convenience of description), the target indication information corresponds to a target category in class C, and the target indication information includes multiple target probability values ​​corresponding one-to-one to multiple pixel points in the first panoramic image. One target probability value in the target indication information is used to indicate the probability that a pixel point in the first panoramic image belongs to a target category in class C. The C target indication information is used to obtain the spherical coordinates of the center point of each first bounding box corresponding to the first panoramic image.

[0101] Furthermore, the target indication information can be specifically expressed as a heat map corresponding to the first panoramic image, and a position (x, y) of a heat map corresponds to a score p. xyc , the fraction p xycrepresents the probability that the predicted category of the pixel at position (x, y) is the target category, x represents the coordinate of the pixel in the longitude direction in the spherical coordinate system corresponding to the first panoramic image, and y represents the coordinate of the pixel in the latitude direction in the spherical coordinate system corresponding to the first panoramic image. xyc The value range is 0-1, p xyc The higher the value of , the greater the probability that the pixel at position (x, y) belongs to the target category. In a heat map, the larger the target probability value corresponding to a pixel, the greater the probability that the pixel is determined to be the center point of the bounding box.

[0102] The electronic device generates a first field of view angle and a second field of view angle corresponding to each pixel point in the first panoramic image through the second branch in the first model. After generating the spherical coordinates of the center point of each first bounding box in multiple first bounding boxes corresponding to the first panoramic image according to the first branch in the first model, the electronic device can obtain the first field of view angle and the second field of view angle corresponding to each center point, and determine the first field of view angle and the second field of view angle corresponding to each center point as the first field of view angle and the second field of view angle corresponding to each first bounding box; the first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction.

[0103] The electronic device can output, using the first model, at least one piece of first position information corresponding to at least one first bounding box, and a predicted category for each first bounding box based on the target indication information, the first field of view angle, and the second field of view angle. Specifically, the target indication information, the first field of view angle, and the second field of view angle are used to determine the position of the first bounding box on the first panoramic image. Furthermore, a pixel with a greater probability value in the target indication information has a greater probability of being determined as the center point of the target category, and therefore, a greater probability of being the center point of the bounding box.

[0104] For a more intuitive understanding of this solution, please refer to Figure 6 , Figure 6 A schematic diagram of a first model provided in an embodiment of the present application, wherein the first model is used to detect C-type objects in a first panoramic image. The first model may include a backbone network, a first branch, and a second branch. Figure 6In this example, the backbone network using a stacked hourglass structure is used. Furthermore, the electronic device can perform convolution processing on the input first panoramic image through the backbone network in the first model to extract a feature map of the first panoramic image. The extracted feature map carries the semantic features of the first panoramic image, and gradually reduces the resolution of the extracted feature map through a pooling layer. The feature map with reduced resolution is then passed to the first branch and the second branch, respectively.

[0105] As shown in the figure, the electronic device can further convolute the feature map generated by the backbone network through the convolution layer in the first branch of the first model, and perform pooling processing through the pooling layer in the first branch of the first model, and finally generate C heat maps corresponding to C categories one by one, that is, C target indication information is generated. For the interpretation of the heat map, please refer to the above description and will not be repeated here.

[0106] The electronic device can perform convolution processing on the feature map generated by the backbone network again through the second branch in the first model to generate the first field of view angle and the second field of view angle corresponding to each pixel point in the first panoramic image. The concepts of the first field of view angle and the second field of view angle can be found in the above description and will not be repeated here.

[0107] The electronic device can determine the spherical coordinates of the center point corresponding to each of the multiple first bounding boxes based on the C heat maps through the first model; the electronic device generates the first field of view angle and the second field of view angle corresponding to the center point of each first bounding box based on the spherical coordinates of the center point corresponding to each first bounding box and the first field of view angle and the second field of view angle corresponding to each pixel in the first panoramic image, and then outputs the first position information corresponding to each of the multiple first bounding boxes. It should be understood that Figure 6 This is only for facilitating understanding of this solution and is not intended to limit this solution.

[0108] In an embodiment of the present application, the spherical coordinates of the center point of each target bounding box in the second panoramic image, the predicted category of each target bounding box, and the first field of view angle and the second field of view angle corresponding to each target bounding box can be directly obtained through the first model, providing a simpler model for target detection and improving the efficiency of the target detection process.

[0109] Optionally, the first model may further include a third branch. Because the electronic device needs to perform a downsampling operation when generating each target indication information through the first model, discretization error may occur when generating the spherical coordinates of the center point of the bounding box based on the target indication information. The electronic device may then generate a first offset and a second offset corresponding to each pixel in the first panoramic image through the third branch in the first model, and then generate updated spherical coordinates of the center point of each first bounding box based on the spherical coordinates of the center point of each first bounding box and the first offset and second offset corresponding to each pixel in the first panoramic image.

[0110] The first offset is a spherical offset in the longitude direction, the second offset is a spherical offset in the latitude direction, and the target indication information, the first offset, and the second offset are used to indicate the spherical coordinates of the updated center point of the first bounding box on the first panoramic image.

[0111] For a more intuitive understanding of this solution, please refer to Figure 7 , Figure 7 Another schematic diagram of the first model provided in the embodiment of the present application, Figure 7 The above needs to be combined Figure 6 , such as Figure 7 As shown, the first model may include a backbone network, a first branch, a second branch, and a third branch. For an understanding of the backbone network, the first branch, and the second branch in the first model, please refer to the above description of the backbone network, the first branch, and the second branch. Figure 6 The description is not repeated here.

[0112] The electronic device can also perform convolution processing on the feature map generated by the backbone network again through the third branch in the first model to generate a first offset and a second offset corresponding to each pixel point in the first panoramic image.

[0113] Correspondingly, the electronic device determines the spherical coordinates of the center point of each first bounding box according to C heat maps (generated by the first branch of the first model) through the first model, and then generates the updated spherical coordinates of the center point of each first bounding box according to the spherical coordinates of the center point of each first bounding box and the first offset and second offset corresponding to each pixel point in the first panoramic image (generated by the third branch of the first model).

[0114] The electronic device generates the first field of view angle and the second field of view angle corresponding to the center point of each first bounding box according to the spherical coordinates of the center point corresponding to each first bounding box and the first field of view angle and the second field of view angle corresponding to each pixel in the first panoramic image (generated by the second branch of the first model), and then outputs the first position information corresponding to each first bounding box in the multiple first bounding boxes. The first position information corresponding to each first bounding box includes the spherical coordinates of the updated center point of each first bounding box and the first field of view angle and the second field of view angle corresponding to the center point of each first bounding box. It should be understood that Figure 7 This is only for facilitating understanding of this solution and is not intended to limit this solution.

[0115] In an embodiment of the present application, since the electronic device needs to perform a downsampling operation in the process of generating each target indication information through the first model, a discretization error will be generated when generating the spherical coordinates of the center point of the bounding box based on the target indication information, that is, there will be an error between the spherical coordinates of the center point obtained based on the target indication information generated by the first branch and the spherical coordinates of the actual center point. Updating the spherical coordinates of the center point generated by the first branch according to the first offset and the second offset generated by the third branch is conducive to obtaining more accurate spherical coordinates of the center point.

[0116] 302. The electronic device obtains second position information of a second bounding box on the first panoramic image, where the second position information is used to indicate the spherical coordinates of each vertex of the second bounding box on the first panoramic image.

[0117] In an embodiment of the present application, second position information of a second bounding box on the first panoramic image may be pre-stored on the electronic device. The second bounding box may be an expected bounding box corresponding to the first bounding box, and the second position information is used to indicate the spherical coordinates of each vertex of the second bounding box on the first panoramic image.

[0118] In one implementation, the second position information of the second bounding box includes the spherical coordinates of each vertex of the first bounding box on the first panoramic image.

[0119] In another implementation, if the second bounding box is rectangular, the second position information of the second bounding box may also include the spherical coordinates of the center point of the second bounding box on the first panoramic image, and the first and second field of view angles corresponding to the second bounding box, where the first field of view angle is the field of view angle corresponding to the longitudinal edge of the second bounding box, and the second field of view angle is the field of view angle corresponding to the latitudinal edge of the second bounding box. In step 302, the electronic device needs to generate the spherical coordinates of each vertex of the second bounding box on the first panoramic image based on the second position information of the second bounding box on the first panoramic image.

[0120] It should be noted that the embodiment of the present application does not limit the execution order of steps 301 and 302. Step 301 may be executed first, and then step 302; or step 302 may be executed first, and then step 301.

[0121] 303. The electronic device generates a spherical area of ​​a target intersection region of the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information.

[0122] In an embodiment of the present application, after the electronic device obtains the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box based on the first position information and the second position information, in one case, the electronic device determines that the intersection between the first bounding box and the second bounding box is empty based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, and then determines the area of ​​the target intersection area to be 0. If each vertex of the first bounding box is not located within the second bounding box, and each vertex of the second bounding box is not located within the first bounding box, then the area of ​​the target intersection area is determined to be 0.

[0123] In another embodiment, the electronic device determines that the first bounding box is located inside the second bounding box based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, and then determines the area of ​​the target intersection region to be the spherical area of ​​the first bounding box. If each vertex of the first bounding box is located inside the second bounding box, then the first bounding box is determined to be inside the second bounding box.

[0124] Specifically, regarding the method for calculating the spherical area of ​​the first bounding box, the electronic device can generate the spherical area of ​​the first bounding box based on the first position information of the first bounding box and the principle of calculating the area of ​​a spherical polygon. To further understand this solution, an example of a formula for calculating the spherical area of ​​the first bounding box based on the principle of calculating the area of ​​a spherical polygon is disclosed below:

[0125]

[0126] Wherein, A(b1) represents the spherical area of ​​the first bounding box, α1 represents the first field of view angle corresponding to the first bounding box, and β1 represents the second field of view angle corresponding to the first bounding box. It should be noted that in formula (1), the radius of the first panoramic image is 1 as an example, that is, the first panoramic image is expressed as a unit sphere as an example. The example in formula (1) is only for the convenience of understanding the calculation method of the spherical area and is not used to limit this solution.

[0127] In another case, the electronic device determines that the second bounding box is located inside the first bounding box based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, and then determines the area of ​​the target intersection region as the spherical area of ​​the second bounding box. If each vertex of the second bounding box is located inside the first bounding box, then the second bounding box is determined to be located inside the first bounding box. The calculation method for the spherical area of ​​the second bounding box is the same as the calculation method for the spherical area of ​​the first bounding box, and can be directly understood by reference, and is not further described here.

[0128] In another embodiment, the electronic device determines that the first bounding box and the second bounding box intersect based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box. If at least one vertex of the second bounding box is located within the first bounding box and at least one vertex of the first bounding box is located within the second bounding box, then the first bounding box and the second bounding box are determined to intersect. For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 Three schematic diagrams of the intersection of the first bounding box and the second bounding box provided in the embodiments of the present application, Figure 8 Including three sub-schematic diagrams (a), (b) and (c), the three sub-schematic diagrams are schematic diagrams after the first and second bounding boxes of the spherical shape are expanded. It should be understood that Figure 8 This is only for facilitating understanding of this solution and is not intended to limit this solution.

[0129] Specifically, step 303 may include: the electronic device generating spherical coordinates of each vertex of the target intersection area on the first panoramic image based on the spherical coordinates of the vertices of the first bounding box on the first panoramic image and the spherical coordinates of the vertices of the second bounding box on the first panoramic image. The electronic device generates the spherical area of ​​the target intersection area based on the spherical coordinates of the vertices of the target intersection area on the first panoramic image and based on a spherical polygon area calculation principle.

[0130] In an embodiment of the present application, a specific implementation method for calculating the spherical area of ​​a target intersection region when a first bounding box and a second bounding box intersect is provided, which is simple to operate and easy to implement.

[0131] More specifically, for the process of generating spherical coordinates of each vertex of the target intersection area on the first panoramic image, the electronic device can determine a first intersection point set based on the spherical coordinates of each vertex of the first bounding box and the spherical coordinates of each vertex of the second bounding box, based on the principle of calculating the intersection between two edges on a sphere. The first intersection point set includes the spherical coordinates of multiple first intersection points, where a first intersection point is an intersection point between any two edges of the first bounding box and the second bounding box.

[0132] The electronic device forms a second point set from all the first intersection points, all the vertices of the first bounding box, and all the vertices of the second bounding box, selects multiple second points from the second point set, and determines the spherical coordinates of each second point, where the second point is a point in the second point set located in the target intersection area.

[0133] Since the first intersection point may coincide with a vertex of the first bounding box, or the first intersection point may coincide with a vertex of the second bounding box, the electronic device removes duplicate points from the multiple second points to obtain all vertices of the target intersection area to obtain the spherical coordinates of each vertex of the target intersection area on the first panoramic image.

[0134] Regarding the process of calculating the spherical area of ​​the target intersection area. The electronic device can calculate the dihedral angle corresponding to each vertex of the target intersection area, and then generate the spherical area of ​​the target intersection area based on the spherical polygon area calculation principle. Among them, the dihedral angle refers to a figure composed of two half-planes starting from a straight line. This straight line is called the edge of the dihedral angle, and the two half-planes are called the faces of the dihedral angle. To further understand this solution, an example of a formula for calculating the spherical area of ​​the target intersection area based on the spherical polygon area calculation principle is disclosed below:

[0135]

[0136] Among them, A(b1∩b2) represents the spherical area of ​​the target intersection area, n represents the number of vertices in the target intersection area, ω i The dihedral angle corresponding to the i-th vertex among the n vertices representing the target intersection area. It should be noted that in formula (1), the radius of the first panoramic image is 1 as an example, that is, the first panoramic image is expressed as a unit sphere as an example. The example in formula (1) is only for the convenience of understanding the calculation method of the spherical area and is not used to limit this solution.

[0137] To further understand this solution, the following uses the example where both the first bounding box and the second bounding box are rectangles to disclose the code used by the electronic device when executing step 303:

[0138] Input:Two spherical rectangles b1 and b2 denoted as and / / Input: first position information of the first bounding box and second position information of the second bounding box

[0139] Output: the area of ​​intersection A(b1∩b2); / / Output: the spherical area of ​​the target intersection area

[0140]

[0141] return 0;

[0142] end / / If the intersection of the first bounding box and the second bounding box is empty, the spherical area of ​​the target intersection area is 0

[0143]

[0144] return min(A(b1),A(b2));

[0145] end / / If the first bounding box is included in the second bounding box, or the second bounding box is included in the first bounding box, the spherical area of ​​the target intersection area is the smaller spherical area of ​​the first bounding box and the second bounding box

[0146] compute the vertices V i of spherical rectangle b i ; / / Determine all first intersections between the edges of the first bounding box and the edges of the second bounding box

[0147] compute the set P of intersection points between boundaries of b1 and those of b2; / / Filter out the second intersection point located in the target intersection area from all the intersection points between the edges of the first bounding box and the edges of the second bounding box

[0148] P←P∪V1∪V2; / / Get all second intersection points, all vertices of the first bounding box and all vertices of the second bounding box

[0149] remove the points p in P such that or

[0150] remove duplicated points in P via loop detection;

[0151] for p i ∈P do; / / Remove duplicate points to obtain the spherical coordinates of each vertex in the target intersection area on the first panoramic image

[0152] compute the angle ω i ; / / Get the dihedral angle corresponding to each vertex in the target intersection area

[0153] end;

[0154] return A(b1∩b2)computed via Equation 3; / / Calculate the spherical area of ​​the target intersection area

[0155] It should be understood that the above codes are shown only to facilitate understanding of the present solution and are not intended to limit the present solution.

[0156] 304. The electronic device generates a spherical area of ​​the first bounding box on the first panoramic image according to the first position information.

[0157] In the embodiment of the present application, step 304 is an optional step. If the spherical area of ​​the first bounding box on the first panoramic image has been generated in step 303, step 304 does not need to be executed; if the spherical area of ​​the first bounding box on the first panoramic image has not been calculated in step 304, step 304 needs to be executed. The specific implementation method of step 304 can be found in the description of step 303 and is not repeated here.

[0158] 305. The electronic device generates a spherical area of ​​a second bounding box on the first panoramic image according to the second position information.

[0159] In the embodiment of the present application, step 305 is an optional step. If the spherical area of ​​the second bounding box on the first panoramic image has been generated in step 303, step 305 does not need to be executed; if the spherical area of ​​the second bounding box on the first panoramic image has not been calculated in step 305, step 305 needs to be executed. The specific implementation method of step 305 can be found in the description of step 303 and is not repeated here.

[0160] It should be noted that if both steps 304 and 305 are executed, the execution order of steps 304 and 305 is not limited in the embodiment of the present application. Step 304 can be executed first and then step 305; or step 305 can be executed first and then step 304.

[0161] 306. The electronic device determines an intersection-over-union ratio between the first bounding box and the second bounding box on the first panoramic image according to the spherical area of ​​the target intersection region.

[0162] In an embodiment of the present application, the electronic device generates an intersection-and-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area, the spherical area of ​​the first bounding box, and the spherical area of ​​the second bounding box, that is, generates an intersection-and-union ratio of the first bounding box and the second bounding box on the sphere.

[0163] To further understand the present solution, an example of calculating the intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image is disclosed as follows:

[0164]

[0165] Among them, IoU(b1,b2) represents the intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image, A(b1∩b2) represents the spherical area of ​​the target intersection area (that is, the intersection area of ​​the first bounding box and the second bounding box on the sphere), A(b1∪b2) represents the spherical area of ​​the union area of ​​the first bounding box and the second bounding box on the sphere, A(b1) represents the spherical area of ​​the first bounding box, and A(b2) represents the spherical area of ​​the second bounding box. It should be understood that the example in formula (3) is only for the convenience of understanding the calculation method of the spherical area and is not used to limit this solution.

[0166] In an embodiment of the present application, the spherical area of ​​the target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image is directly calculated, and the intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image is calculated based on the spherical area of ​​the target intersection area. That is, the intersection-over-union ratio of the first bounding box and the second bounding box on the spherical surface is directly calculated, which greatly reduces the error and improves the accuracy of the present solution.

[0167] The embodiment of the present application also provides a second model for image processing, which is used to detect C-type objects in a panoramic image to generate position information of at least one target bounding box corresponding to at least one object. The position information of each target bounding box includes the spherical coordinates of the center point of the target bounding box on the panoramic image and the first field of view angle corresponding to the target bounding box and the second field of view angle corresponding to the target bounding box. Before introducing the second model provided in the embodiment of the present application in detail, let's first combine Figure 9 The image processing system provided by the embodiment of the present application is introduced. Figure 9 , Figure 9 A system architecture diagram of the image processing system provided in the embodiment of the present application, Figure 9 In the figure, the image processing system 900 includes an execution device 910, a training device 920, a database 930 and a data storage system 940, and the execution device 910 includes a computing module 911.

[0168] Database 930 stores a set of training images, and training device 920 generates a second model 901, which is used to detect objects in an input panoramic image. Training device 920 iteratively trains second model 901 using the set of training images in database 930, resulting in a mature second model 901. Furthermore, second model 901 can be implemented using a neural network or a non-neural network model.

[0169] The mature second model 901 obtained by the training device 920 can be applied to different systems or devices, such as mobile phones, tablets, laptops, VR devices, monitoring systems, radar data processing systems, etc. Among them, the execution device 910 can call data, codes, etc. in the data storage system 940, and can also store data, instructions, etc. in the data storage system 940. The data storage system 940 can be placed in the execution device 910, or the data storage system 940 can be an external memory relative to the execution device 910. The calculation module 911 can perform target detection on the input panoramic image through the second model 901 to output the detection results, the position information of at least one bounding box of the detection result, and the predicted category of each bounding box. The position information of the bounding box is used to indicate the position of the object in the input panoramic image.

[0170] In some embodiments of this application, please refer to Figure 9 The "user" can directly interact with the execution device 910, that is, the execution device 910 can directly display the predicted image output by the second model 901 to the "user". It is worth noting that Figure 9 This is merely a schematic diagram of the architecture of the image processing system provided by an embodiment of the present invention, and the positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in other embodiments of the present application, the execution device 910 and the client device may be separate devices, and the execution device 910 may be configured with an input / output (I / O) interface, through which the execution device 910 exchanges data with the client device.

[0171] In combination with the above description, it can be seen that the specific implementation process of the reasoning stage and the training stage of the second model provided in the embodiment of the present application will be described below.

[0172] 1. Reasoning Stage

[0173] In the embodiment of the present application, the inference stage describes how the execution device 910 uses the second model 901 to perform image processing to generate a predicted image. For details, please refer to Figure 10 , Figure 10A structural schematic diagram of a second model provided in an embodiment of the present application, wherein the second model 1000 provided in an embodiment of the present application may include a first branch 1001 and a second branch 1002; the second model 1000 is used to receive a second panoramic image to perform target detection on the second panoramic image; the first branch 1001 is used to generate target indication information, where the target indication information is used to indicate the spherical coordinates of the center point of a target bounding box in the second panoramic image on the second panoramic image, and the predicted category of the target bounding box, where the spherical coordinates include spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; the second branch 1002 is used to generate a first field of view angle and a second field of view angle corresponding to the center point of the target bounding box, where the first field of view angle is the field of view angle corresponding to the edge of the target bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the target bounding box in the latitude direction; wherein the spherical coordinates of the center point of the target bounding box on the second panoramic image and the first and second field of view angles corresponding to the center point of the target bounding box are used to determine the position of the target bounding box on the second panoramic image.

[0174] Optionally, see Figure 11 , Figure 11 This is another structural schematic diagram of the second model provided in an embodiment of the present application, where the second model 1000 also includes a third branch 1003; the third branch 1003 is used to generate a first offset and a second offset corresponding to the center point of the target bounding box, wherein the first offset is a spherical offset in the longitude direction, the second offset is a spherical offset in the latitude direction, the spherical coordinates of the center point of the target bounding box on the second panoramic image, and the first offset and the second offset are used to indicate the spherical coordinates of the updated center point of the target bounding box on the second panoramic image.

[0175] Furthermore, the second model 1000 is used to perform target detection on Class C objects in the second panoramic image, and the target indication information corresponds to the target category in Class C. The target indication information includes multiple target probability values ​​corresponding one-to-one to multiple pixel points in the first panoramic image, and a target probability value is used to indicate the probability that a pixel point in the first panoramic image belongs to the target category in Class C.

[0176] In the embodiment of the present application, the information interaction, execution process, etc. between the modules / units in the execution of the second model 1000 are the same as those in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 8 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0177] 2. Training Phase

[0178] In the embodiment of the present application, the training phase describes how the training device 920 generates a mature second model using the image data set in the database 930. For details, please refer to Figure 12 , Figure 12 A flowchart of a neural network training method provided in an embodiment of the present application is provided. The neural network training method provided in an embodiment of the present application may include:

[0179] 1201. The electronic device inputs the third panoramic image into the second model, and performs target detection on the third panoramic image through the second model.

[0180] 1202. The electronic device generates first indication information through the first branch of the second model, where the first indication information is used to indicate the spherical coordinates of a center point of a predicted bounding box corresponding to the third panoramic image on the third panoramic image.

[0181] 1203. The electronic device generates a first predicted field of view angle and a second predicted field of view angle corresponding to the center point of the predicted bounding box through the second branch of the second model, where the first predicted field of view angle is the field of view angle corresponding to the edge of the predicted bounding box in the longitude direction, and the second predicted field of view angle is the field of view angle corresponding to the edge of the predicted bounding box in the latitude direction.

[0182] 1204. The electronic device generates a first prediction offset and a second prediction offset corresponding to the center point of the predicted bounding box through the third branch of the second model, where the first prediction offset is a spherical offset in the longitude direction, and the second prediction offset is a spherical offset in the latitude direction. The first indication information, the first prediction offset, and the second prediction offset are used to indicate the spherical coordinates of the updated center point of the predicted bounding box on the third panoramic image.

[0183] In an embodiment of the present application, a training data set is configured on the electronic device, and the training data set includes multiple panoramic images, position information of at least one expected bounding box corresponding to each panoramic image, and a predicted category corresponding to each expected bounding box. During each training process, the electronic device can obtain a panoramic image for training from the training data set (hereinafter referred to as the "third panoramic image" for convenience of description), input the third panoramic image into the second model, and generate the position information of at least one predicted bounding box corresponding to the third panoramic image through steps 1201 to 1204. The specific implementation method of the electronic device performing steps 1201 to 1204 can be found in Figure 3 The description of step 301 in the corresponding embodiment is omitted here.

[0184] The meaning of "first indication information" is similar to that of "target indication information", the meaning of "first predicted field of view angle" is similar to that of "first field of view angle", the meaning of "second predicted field of view angle" is similar to that of "second field of view angle", the meaning of "first predicted offset" is similar to that of "first offset", and the meaning of "second predicted offset" is similar to that of "second offset", and all of them can be referred to above. Figure 3 Please understand the description in the corresponding embodiment and do not elaborate on it here.

[0185] 1205. The electronic device obtains first position information of an expected bounding box corresponding to the predicted bounding box on the third panoramic image, where the first position information of the expected bounding box includes second indication information, a first expected field of view angle, and a second expected field of view angle. The second indication information is used to indicate the spherical coordinates of a center point of the expected bounding box on the third panoramic image. The first expected field of view angle is the field of view angle corresponding to an edge of the expected bounding box in the longitude direction, and the second expected field of view angle is the field of view angle corresponding to an edge of the expected bounding box in the latitude direction.

[0186] In an embodiment of the present application, the electronic device also obtains at least one expected bounding box corresponding to the third panoramic image to obtain second position information of at least one expected bounding box corresponding to the third panoramic image. The second position information of an expected bounding box may include the spherical coordinates of the center point of the expected bounding box on the third panoramic image, the first expected field of view angle of the expected bounding box, and the second expected field of view angle of the expected bounding box.

[0187] The electronic device needs to generate second indication information corresponding to at least one expected bounding box corresponding to the third panoramic image based on the Gaussian kernel function principle according to the second position information of at least one expected bounding box corresponding to the third panoramic image, the second indication information including the expected probability value of each pixel point in the third panoramic image, and the meaning of the "second indication information" is the same as Figure 3 The meaning of the "target indication information" in the corresponding embodiment is similar, please refer to the above description, and will not be repeated here.

[0188] Specifically, since there may be multiple expected bounding boxes on the third panoramic image, the second indication information includes multiple groups of indication information corresponding to the multiple expected bounding boxes. For any expected bounding box in at least one expected bounding box corresponding to the third panoramic image. A target threshold may be pre-configured on the electronic device, and the Gaussian kernel radius is calculated based on the Gaussian kernel function principle according to the spherical coordinates of the center point of the expected bounding box, the target threshold, the first expected field of view angle of the expected bounding box, and the second expected field of view angle of the expected bounding box. The spherical coordinates of the center point of the expected bounding box are the actual spherical coordinates of the center point of the expected bounding box obtained from the training data set, and the intersection-over-union ratio between the spherical area of ​​the third bounding box corresponding to the Gaussian kernel radius and the expected bounding box is less than or equal to the target threshold.

[0189] Furthermore, the larger the target threshold value, the fewer pixels with a probability value other than 0 are marked in the second indication information; the smaller the target threshold value, the more pixels with a probability value other than 0 are marked in the second indication information. As an example, the target threshold value may be 0.6, 0.7, 0.8, or other values, etc., which are not limited here.

[0190] Furthermore, in order to understand this solution more intuitively, Figure 13 The process of generating the Gaussian kernel radius in the embodiment of the present application is introduced. Figure 13 A schematic diagram of generating a Gaussian kernel radius in the training method of a neural network provided in an embodiment of the present application. After determining the spherical coordinates of the center point of the desired bounding box, the first desired field of view angle and the second desired field of view angle corresponding to the desired bounding box, the training device can calculate Figure 13 The distance r corresponding to the three cases (a), (b) and (c) is Figure 13 In the (a) sub-diagram, assuming that the desired bounding box is outside the third bounding box, calculate the value of the distance r when the spherical area of ​​the desired bounding box and the third bounding box is approximately equal to the target threshold. Figure 13 The distance r in the sub-schematic diagram (a) represents the distance between the edge of the desired bounding box and the edge of the third bounding box in the longitude direction, and also represents the distance between the edge of the desired bounding box and the edge of the third bounding box in the latitude direction.

[0191] exist Figure 13 In the (b) sub-diagram, assuming that the desired bounding box is inside the third bounding box, calculate the value of the distance r when the spherical area of ​​the desired bounding box and the third bounding box is approximately equal to the target threshold. Figure 13 The distance r in the sub-schematic diagram (b) represents the distance between the edge of the desired bounding box and the edge of the third bounding box in the longitude direction, and also represents the distance between the edge of the desired bounding box and the edge of the third bounding box in the latitude direction.

[0192] exist Figure 13 In the (c) sub-diagram, assuming that the desired bounding box and the third bounding box intersect, the value of the distance r is calculated when the spherical area of ​​the desired bounding box and the third bounding box is approximately equal to the target threshold. Figure 13 The distance r in the (c) sub-diagram represents the distance between the center point of the desired bounding box and the center point of the third bounding box in the longitude direction, and also represents the distance between the center point of the desired bounding box and the center point of the third bounding box in the latitude direction. Figure 13 The target distance r with the smallest value is selected from the three distances r corresponding to the three sub-schematic diagrams (a), (b) and (c), and the target distance r is determined as the Gaussian kernel radius. It should be understood that Figure 13 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.

[0193] The electronic device performs a linear transformation on the Gaussian kernel radius to obtain a standard deviation of the Gaussian kernel function, and generates a set of indication information in the second indication information based on the standard deviation of the Gaussian kernel function and the spherical coordinates of the center point of the expected bounding box. The electronic device performs the above operation for each expected bounding box in the third panoramic image, thereby obtaining multiple sets of indication information in the second indication information. Optionally, the electronic device marks the expected probability values ​​of the remaining pixels in the third panoramic image as 0, thereby obtaining complete second indication information.

[0194] 1206. The electronic device obtains a first expected offset and a second expected offset corresponding to the expected bounding box, where the first expected offset is a spherical offset in a longitude direction, and the second expected offset is a spherical offset in a latitude direction.

[0195] In some embodiments of the present application, after the electronic device obtains a second panoramic image marked with a desired bounding box from a training data set, it can obtain the actual spherical coordinates of the center point of the desired bounding box (hereinafter referred to as the first spherical coordinates of the center point of the desired bounding box for convenience of description), where the first spherical coordinates include the actual spherical coordinates of the center point of the desired bounding box in the longitude direction and the actual spherical coordinates of the center point of the desired bounding box in the latitude direction.

[0196] The electronic device may further downsample the second panoramic image marked with the desired bounding box, and obtain second spherical coordinates of the center point of the desired bounding box based on the downsampled second panoramic image. The second spherical coordinates have lower precision than the first spherical coordinates, that is, the second spherical coordinates are obtained by quantizing the first spherical coordinates. The first spherical coordinates include the quantized spherical coordinates of the center point of the desired bounding box in the longitude direction and the quantized spherical coordinates of the center point of the desired bounding box in the latitude direction.

[0197] The electronic device calculates the difference between the actual spherical coordinates and the quantized spherical coordinates of the center point of the desired bounding box in the longitude direction, and determines the difference as a first desired offset; calculates the difference between the actual spherical coordinates and the quantized spherical coordinates of the center point of the desired bounding box in the latitude direction, and determines the difference as a second desired offset.

[0198] 1207. The electronic device trains the second model according to the target loss function.

[0199] In an embodiment of the present application, the electronic device can calculate the function value of the target loss function based on the position information of the predicted bounding box generated by the second model and the position information of the expected bounding box, and reversely update the parameters of the second model based on the function value of the target loss function to complete one training of the second model.

[0200] Steps 1204 and 1206 are optional. If steps 1204 and 1206 are not performed, the target loss function may include a first loss term and a second loss term, and the target loss function is obtained by weighted summing the first loss term and the second loss term. The first loss term is used to indicate the similarity between the first indication information and the second indication information; the second loss term is used to indicate the similarity between the first predicted field of view angle and the first expected field of view angle, and is also used to indicate the similarity between the second predicted field of view angle and the second expected field of view angle.

[0201] Furthermore, the first loss term can adopt a focal loss function (Focal Loss), a cross entropy loss function, a 0-1 loss function, or other types of loss functions. Since the first indication information includes the predicted probability value of each pixel point of the third panoramic image, and the second indication information includes the expected probability value of each pixel point of the third panoramic image, the first loss term is used to indicate the similarity between the predicted probability value and the expected probability value of each pixel point of the third panoramic image.

[0202] Optionally, since the third panoramic image is a sphere, in the expanded image corresponding to the third panoramic image, the area of ​​objects located near the equator of the third panoramic image is smaller, and the area of ​​objects located near the poles of the third panoramic image is larger. Different weights can be set for different pixel points in the third panoramic image, so that the weights of the pixel points near the equator of the third panoramic image are higher, and the weights of the pixel points near the poles of the third panoramic image are lower, so as to increase the attention to objects near the equator of the panoramic image during the training stage. Since objects with smaller areas are more difficult to learn, the above method can be beneficial to improving the overall accuracy of the second model after training.

[0203] To further understand this solution, an example of the first loss term is disclosed below:

[0204]

[0205] Among them, L cls represents the first loss term, N represents the number of pixels in the third panoramic image, and w xy Represents the weight of each pixel in the third panoramic image, p xyc represents the predicted probability value corresponding to a pixel point in the cth first indication information, represents the expected probability value corresponding to a pixel point in the c-th first indication information, y represents the coordinate of a pixel point in the latitude direction, H represents the length of the expanded image of the third panoramic image (corresponding to the latitude direction of the third panoramic image), and W represents the width of the expanded image of the third panoramic image (corresponding to the longitude direction of the third panoramic image). It should be understood that the examples in formula (4) and formula (5) are only for the convenience of understanding this solution, and formula (4) and formula (5) are both presented with the third panoramic image as a unit sphere as an example. The specific first loss term can be determined in combination with actual conditions and is not limited here.

[0206] The second loss term can adopt L1 loss function, MSE loss function or other types of loss functions. To further understand this solution, an example of the second loss term is disclosed below:

[0207]

[0208] Among them, L fov represents the second loss term, N represents the number of pixels in the third panoramic image, s i represents the first predicted field of view angle and the second predicted field of view angle corresponding to a pixel point in the third panoramic image, represents the first expected field of view angle and the second expected field of view angle corresponding to a pixel point in the third panoramic image. It should be understood that the example in formula (6) is only for the convenience of understanding this solution and is not used to limit this solution.

[0209] If steps 1204 and 1206 are executed, the target loss function may further include a third loss term, which is used to indicate the similarity between the first predicted offset and the first expected offset, and is also used to indicate the similarity between the second predicted offset and the second expected offset. The target loss function is obtained by weighted summing the first loss term, the second loss term, and the third loss term.

[0210] The third loss term can adopt L1 loss function, MSE loss function or other types of loss functions. To further understand this solution, an example of the second loss term is disclosed below:

[0211]

[0212] Among them, L off represents the third loss term, Ci represents the coordinates of the center point of the predicted bounding box obtained by the first branch of the second model, o i Represents the first prediction offset and the second prediction offset corresponding to the center point of the predicted bounding box, Represents the first expected offset and the second expected offset corresponding to the center point of the predicted bounding box, T(.) is the transformation of the azimuth angle into a unit vector, <.,.> represents the calculation of the vector inner product. It should be understood that the example in formula (7) is only for the convenience of understanding this scheme and is not used to limit this scheme.

[0213] In the embodiment of the present application, the electronic device repeatedly performs steps 1201 to 1207 to implement iterative training of the second model until a convergence condition is met, thereby obtaining a trained second model, i.e., a mature second model. The convergence condition may be a convergence condition of reaching a target loss function, or the number of iterations of steps 1201 to 1207 may reach a preset number.

[0214] In order to gain a more direct understanding of the beneficial effects of the embodiments of this application, experiments were conducted using the public datasets 360-Indoor and 360-VOC, respectively. The evaluation metric for the experiments was the average precision (AP) of object detection in panoramic images. The experimental results are shown in Table 1.

[0215]

[0216] Table 1

[0217] Among them, Center Net and Sphere-SSD are two neural networks used for object detection in panoramic images. Before determining whether the predicted bounding box output by the model is correct, an IoU threshold between the predicted bounding box and the expected bounding box is pre-set. That is, when the IoU between the predicted bounding box and the expected bounding box is greater than or equal to the IoU threshold, the predicted bounding box generated by the model is determined to be an accurate bounding box. AP represents the average of the average accuracy rates obtained with a step size of 0.05 and an IoU threshold from 0.5 to 0.95. AP 50 Represents the average accuracy when the IoU threshold is set to 0.5, AP 75 It represents the average accuracy when the IoU threshold is set to 0.75. From the data shown in Table 1, it can be seen that the average accuracy of target detection in panoramic images using the second model provided in the embodiment of the present application is the highest.

[0218] In Figures 1 to Figure 13 On the basis of the corresponding embodiment, in order to better implement the above solution of the embodiment of the present application, the following also provides related equipment for implementing the above solution. Figure 14, Figure 14 A structural schematic diagram of an image processing device provided in an embodiment of the present application, the image processing device 1400 includes: a detection module 1401, which is used to perform target detection on a first panoramic image using a first model to obtain a detection result corresponding to the first panoramic image, the detection result including first position information of a first bounding box on the first panoramic image, the first position information being used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; an acquisition module 1402, which is also used to obtain second position information of a second bounding box on the first panoramic image, the second position information being used to indicate the spherical coordinates of the vertices of the second bounding box on the first panoramic image; a generation module 1403, which is used to generate a spherical area of ​​a target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; and a determination module 1404, which is used to determine an intersection-over-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area.

[0219] In one possible design, the first bounding box is a rectangle; the first position information of the first bounding box includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, the first field of view angle and the second field of view angle corresponding to the first bounding box, the first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction.

[0220] In one possible design, if the first bounding box and the second bounding box intersect, the generation module 1403 is specifically used to: generate the spherical coordinates of the vertices of the target intersection area on the first panoramic image based on the spherical coordinates of the vertices of the first bounding box on the first panoramic image and the spherical coordinates of the vertices of the second bounding box on the first panoramic image; and generate the spherical area of ​​the target intersection area based on the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

[0221] In one possible design, the image processing apparatus 1400 is applied to a training phase of a first model, and / or the apparatus is applied to detect the accuracy of a first bounding box generated by the first model.

[0222] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1400 are the same as those in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0223] See also Figure 15 , Figure 15Another structural schematic diagram of an image processing device provided in an embodiment of the present application, the image processing device 1500 includes: an acquisition module 1501, used to obtain first position information of a first bounding box on a first panoramic image, the first position information is used to indicate the spherical coordinates of the vertices of the first bounding box on the first panoramic image, and the spherical coordinates include spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; the acquisition module 1501 is also used to obtain second position information of a second bounding box on the first panoramic image, the second position information is used to indicate the spherical coordinates of the vertices of the second bounding box on the first panoramic image; a generation module 1502 is used to generate a spherical area of ​​a target intersection area of ​​the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; a determination module 1503 is used to determine the intersection-union ratio of the first bounding box and the second bounding box on the first panoramic image based on the spherical area of ​​the target intersection area.

[0224] In one possible design, the first bounding box is a rectangle; the first position information of the first bounding box includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, the first field of view angle and the second field of view angle corresponding to the first bounding box, the first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction.

[0225] In one possible design, if the first bounding box and the second bounding box intersect, the generation module 1503 is specifically used to: generate the spherical coordinates of the vertices of the target intersection area on the first panoramic image based on the spherical coordinates of the vertices of the first bounding box on the first panoramic image and the spherical coordinates of the vertices of the second bounding box on the first panoramic image; and generate the spherical area of ​​the target intersection area based on the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

[0226] In one possible design, the image processing apparatus 1500 is applied to a training phase of a first model, and / or the apparatus is applied to detect the accuracy of a first bounding box generated by the first model.

[0227] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1500 are the same as those in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0228] See also Figure 16 , Figure 16Another structural schematic diagram of an image processing device provided in an embodiment of the present application, wherein image processing device 1600 is configured with a second model, the second model including a first branch and a second branch, and image processing device 1600 includes an input module 1601 and a processing module 1602, wherein input module 1601 is configured to input a second panoramic image into the second model; processing module 1602 is configured to generate target indication information through the first branch, the target indication information being used to indicate the spherical coordinates of the center point of a target bounding box on the second panoramic image, and the predicted category of the target bounding box, the spherical coordinates including spherical coordinates in the longitude direction and spherical coordinates in the latitude direction; processing module 1602 is further configured to generate a first field of view angle and a second field of view angle corresponding to the center point of the target bounding box through the second branch, the first field of view angle being the field of view angle corresponding to the edge of the target bounding box in the longitude direction, and the second field of view angle being the field of view angle corresponding to the edge of the target bounding box in the latitude direction; wherein the target indication information and the first and second field of view angles corresponding to the center point of the target bounding box are used to indicate the position of the target bounding box on the second panoramic image and the predicted category of the target bounding box.

[0229] In one possible design, the second model is used to perform target detection on C-type objects in the second panoramic image, where C is an integer greater than or equal to 1, the target indication information corresponds to the target category in category C, and the target indication information includes multiple target probability values ​​corresponding to multiple pixel points in the first panoramic image, and a target probability value is used to indicate the probability that a pixel point in the first panoramic image belongs to the target category.

[0230] In one possible design, the second model also includes a third branch; the processing module 1602 is further used to generate a first offset and a second offset corresponding to the center point of the target bounding box through the third branch of the second model, wherein the first offset is a spherical offset in the longitude direction, the second offset is a spherical offset in the latitude direction, the spherical coordinates of the center point of the target bounding box on the second panoramic image, and the first offset and the second offset are used to indicate the spherical coordinates of the updated center point of the target bounding box on the second panoramic image.

[0231] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1600 are the same as those in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0232] Next, an electronic device provided by an embodiment of the present application is introduced, which is used to implement Figures 3 to 8 In the corresponding embodiment, when the electronic device is specifically a server, please refer to Figure 17 , Figure 17 It is a structural diagram of an electronic device provided in an embodiment of the present application. Specifically, the electronic device 1700 is implemented by one or more servers. The electronic device 1700 may have relatively large differences due to different configurations or performances. It may include one or more central processing units (CPU) 1722 (for example, one or more processors) and memory 1732, and one or more storage media 1730 (for example, one or more mass storage devices) for storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage medium 1730 can be short-term storage or persistent storage. The program stored in the storage medium 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the electronic device. Furthermore, the central processing unit 1722 can be configured to communicate with the storage medium 1730 to execute a series of instruction operations in the storage medium 1730 on the electronic device 1700.

[0233] The electronic device 1700 may also include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input and output interfaces 1758, and / or one or more operating systems 1741, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0234] In the embodiment of the present application, the central processing unit 1722 is used to execute Figures 3 to 8 The specific manner in which the central processing unit 1722 performs the above steps is the same as that in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 8 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0235] The present application also provides another electronic device. Figure 18 , Figure 18 Another structural diagram of an electronic device provided in an embodiment of the present application. In one case, the electronic device 1800 may be deployed with Figure 14 The image processing device 1400 described in the corresponding embodiment, or, Figure 15 The image processing device 1500 described in the corresponding embodiment is used to implement Figures 3 to 8 In another embodiment, the electronic device 1800 may be provided with Figure 16The image processing apparatus 1600 described in the corresponding embodiment is used to perform Figure 10 or Figure 11 The steps of the second model shown are performed. In this case, the electronic device 1800 can be a virtual reality (VR) device, an intelligent robot, or a monitoring data processing device, etc., which is not limited here. Specifically, the electronic device 1800 includes: a receiver 1801, a transmitter 1802, a processor 1803, and a memory 1804 (wherein the number of processors 1803 in the electronic device 1800 can be one or more, Figure 18 (taking one processor as an example), the processor 1803 may include an application processor 18031 and a communication processor 18032. In some embodiments of the present application, the receiver 1801, the transmitter 1802, the processor 1803 and the memory 1804 may be connected via a bus or other means.

[0236] Memory 1804 may include read-only memory and random access memory, and provides instructions and data to processor 1803. A portion of memory 1804 may also include non-volatile random access memory (NVRAM). Memory 1804 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0237] Processor 1803 controls the operation of the electronic device. In specific applications, the various components of the electronic device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0238] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1803. Processor 1803 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 1803 or software instructions. The above processor 1803 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1803 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1804, and processor 1803 reads information from memory 1804 and, in conjunction with its hardware, completes the steps of the above method.

[0239] Receiver 1801 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the electronic device. Transmitter 1802 can be used to output digital or character information through the first interface. Transmitter 1802 can also be used to send instructions to the disk pack through the first interface to modify the data in the disk pack. Transmitter 1802 can also include a display device such as a display screen.

[0240] In one embodiment of the present application, the application processor 18031 is used to execute Figures 3 to 8 The image processing method executed by the electronic device in the corresponding embodiment. The specific manner in which the application processor 18031 executes each step is the same as that in the present application. Figures 3 to 8 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 8 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0241] In another embodiment, the application processor 18031 is used to execute Figure 10or Figure 11 The specific manner in which the application processor 18031 performs each step is the same as that in 10 or 11 in this application. Figure 11 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figure 10 or Figure 11 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0242] The present application also provides a computer program product which, when executed on a computer, enables the computer to execute the aforementioned Figure 10 or Figure 11 The steps performed by the second model in the embodiment shown, or the steps of making the computer perform the above Figures 3 to 8 The illustrated embodiment describes the steps performed by the electronic device in the method.

[0243] The present application also provides a computer-readable storage medium in which a program for signal processing is stored. When the program is run on a computer, the computer executes the above-mentioned Figure 10 or Figure 11 The steps performed by the second model in the embodiment shown, or the steps of making the computer perform the above Figures 3 to 8 The illustrated embodiment describes the steps performed by the electronic device in the method.

[0244] The image processing device, neural network training device, execution device and electronic device provided in the embodiments of the present application can be specifically a chip, which includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip to execute the above Figure 10 or Figure 11 The steps performed by the second model in the embodiment shown, or, to enable the chip to perform the above Figures 3 to 8 The image processing method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), etc.

[0245] For details, please refer to Figure 19 , Figure 19This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 190. NPU 190 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1903, which is controlled by controller 1904 to extract matrix data from memory and perform multiplication operations.

[0246] In some implementations, the arithmetic circuit 1903 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1903 is a two-dimensional systolic array. The arithmetic circuit 1903 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1903 is a general-purpose matrix processor.

[0247] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1902 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1901 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1908.

[0248] Unified memory 1906 is used to store input and output data. Weight data is directly transferred to weight memory 1902 through the Direct Memory Access Controller (DMAC) 1905. Input data is also transferred to unified memory 1906 through the DMAC.

[0249] BIU stands for Bus Interface Unit 1910 , which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1909 .

[0250] The bus interface unit 1910 (BIU) is used for the instruction fetch memory 1909 to obtain instructions from the external memory, and is also used for the storage unit access controller 1905 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0251] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1906 or to transfer weight data to the weight memory 1902 or to transfer input data to the input memory 1901.

[0252] The vector calculation unit 1907 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0253] In some implementations, the vector calculation unit 1907 can store the processed output vector to the unified memory 1906. For example, the vector calculation unit 1907 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1903, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1907 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1903, for example, for use in a subsequent layer in a neural network.

[0254] An instruction fetch buffer 1909 connected to the controller 1904 is used to store instructions used by the controller 1904;

[0255] Unified memory 1906, input memory 1901, weight memory 1902, and instruction fetch memory 1909 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0256] In which, the operations of each layer in the first model or the second model shown in the above embodiments can be performed by the operation circuit 1903 or the vector calculation unit 1907.

[0257] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.

[0258] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0259] Through the description of the above embodiments, it is clear to those skilled in the art that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course it can also be implemented by means of dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, electronic device, or electronic device, etc.) to execute the methods described in each embodiment of the present application.

[0260] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0261] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, an electronic device, or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, electronic device, or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as an electronic device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

Claims

1. An image processing method, characterized in that: The method comprises: performing object detection on the first panoramic image using the first model to obtain a detection result corresponding to the first panoramic image, the detection result including first position information of a first bounding box on the first panoramic image, the first position information being used to indicate spherical coordinates of vertices of the first bounding box on the first panoramic image, the spherical coordinates including spherical coordinates in a longitude direction and spherical coordinates in a latitude direction; Acquire second position information of a second bounding box on the first panoramic image, where the second bounding box is an expected bounding box corresponding to the first bounding box, and the second position information is used to indicate spherical coordinates of vertices of the second bounding box on the first panoramic image; generating a spherical area of ​​a target intersection region of the first bounding box and the second bounding box on the first panoramic image according to the first position information and the second position information; An intersection-over-union ratio (IoU) between the first bounding box and the second bounding box on the first panoramic image is determined according to a spherical area of ​​the target intersection region.

2. The method according to claim 1, characterized in that The first bounding box is a rectangle; The first position information of the first bounding box includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, a first field of view angle corresponding to the first bounding box, and a second field of view angle, where the first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction.

3. The method according to claim 1 or 2, characterized in that If the first bounding box and the second bounding box intersect, generating a spherical area of ​​an intersection region of the first bounding box and the second bounding box on the first panoramic image according to the first position information and the second position information includes: generating spherical coordinates of vertices of the target intersection area on the first panoramic image according to the spherical coordinates of vertices of the first bounding box on the first panoramic image and the spherical coordinates of vertices of the second bounding box on the first panoramic image; The spherical area of ​​the target intersection area is generated according to the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

4. The method according to claim 1 or 2, characterized in that The method is applied to the training phase of the first model, and / or the method is applied to detecting the accuracy of a first bounding box generated by the first model.

5. An image processing method, characterized in that: The method comprises: Acquire first position information of a first bounding box on the first panoramic image, where the first position information is used to indicate spherical coordinates of vertices of the first bounding box on the first panoramic image, where the spherical coordinates include spherical coordinates in a longitude direction and spherical coordinates in a latitude direction; Acquire second position information of a second bounding box on the first panoramic image, where the second position information is used to indicate spherical coordinates of vertices of the second bounding box on the first panoramic image; generating a spherical area of ​​a target intersection region of the first bounding box and the second bounding box on the first panoramic image according to the first position information and the second position information; An intersection-over-union ratio (IoU) between the first bounding box and the second bounding box on the first panoramic image is determined according to a spherical area of ​​the target intersection region.

6. The method according to claim 5, characterized in that If the first bounding box and the second bounding box intersect, generating a spherical area of ​​an intersection region of the first bounding box and the second bounding box on the first panoramic image according to the first position information and the second position information includes: generating spherical coordinates of vertices of the target intersection area on the first panoramic image according to the spherical coordinates of vertices of the first bounding box on the first panoramic image and the spherical coordinates of vertices of the second bounding box on the first panoramic image; The spherical area of ​​the target intersection area is generated according to the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

7. An image processing device, characterized in that The image processing device includes a second model for performing target detection on a panoramic image, wherein the second model includes a first branch and a second branch; wherein The second model is used to receive a second panoramic image to perform target detection on the second panoramic image; The first branch is configured to generate target indication information, where the target indication information is configured to indicate the spherical coordinates of a center point of a target bounding box on the second panoramic image and a predicted category of the target bounding box, where the spherical coordinates include spherical coordinates in a longitude direction and spherical coordinates in a latitude direction; The second branch is configured to generate a first field of view angle and a second field of view angle corresponding to the center point of the target bounding box, wherein the first field of view angle is the field of view angle corresponding to the edge of the target bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the target bounding box in the latitude direction; The target indication information, the first field of view angle and the second field of view angle corresponding to the center point of the target bounding box are used to indicate the position of the target bounding box on the second panoramic image and the predicted category of the target bounding box.

8. The device according to claim 7, characterized in that The second model is used to perform target detection on C-type objects in the second panoramic image, where C is an integer greater than or equal to 1, and the target indication information corresponds to the target category in the C category. The target indication information includes multiple target probability values ​​corresponding to multiple pixel points in the first panoramic image, and a target probability value is used to indicate the probability that a pixel point in the first panoramic image belongs to the target category.

9. The device according to claim 7 or 8, characterized in that The second model also includes a third branch; The third branch is used to generate a first offset and a second offset corresponding to the center point of the target bounding box, wherein the first offset is a spherical offset in the longitude direction, the second offset is a spherical offset in the latitude direction, and the spherical coordinates of the center point of the target bounding box on the second panoramic image, the first offset and the second offset are used to indicate the spherical coordinates of the updated center point of the target bounding box on the second panoramic image.

10. An image processing device, characterized in that: The device comprises: a detection module that performs object detection on the first panoramic image using a first model to obtain a detection result corresponding to the first panoramic image, the detection result including first position information of a first bounding box on the first panoramic image, the first position information being used to indicate spherical coordinates of vertices of the first bounding box on the first panoramic image, the spherical coordinates including spherical coordinates in a longitude direction and spherical coordinates in a latitude direction; an acquisition module, configured to acquire second position information of a second bounding box on the first panoramic image, where the second bounding box is a desired bounding box corresponding to the first bounding box, and the second position information is used to indicate spherical coordinates of vertices of the second bounding box on the first panoramic image; a generating module, configured to generate a spherical area of ​​a target intersection region of the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; A determination module is configured to determine an intersection-over-union ratio (IoU) of the first bounding box and the second bounding box on the first panoramic image based on a spherical area of ​​the target intersection region.

11. The device according to claim 10, characterized in that The first bounding box is a rectangle; The first position information of the first bounding box includes the spherical coordinates of the center point of the first bounding box on the first panoramic image, a first field of view angle corresponding to the first bounding box, and a second field of view angle, where the first field of view angle is the field of view angle corresponding to the edge of the first bounding box in the longitude direction, and the second field of view angle is the field of view angle corresponding to the edge of the first bounding box in the latitude direction.

12. The device according to claim 10 or 11, characterized in that If the first bounding box and the second bounding box intersect, the generating module is specifically configured to: generating spherical coordinates of vertices of the target intersection area on the first panoramic image according to the spherical coordinates of vertices of the first bounding box on the first panoramic image and the spherical coordinates of vertices of the second bounding box on the first panoramic image; The spherical area of ​​the target intersection area is generated according to the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

13. The device according to claim 10 or 11, characterized in that The device is applied to a training phase of the first model, and / or the device is applied to detect the accuracy of a first bounding box generated by the first model.

14. An image processing device, characterized in that: The device comprises: an acquisition module, configured to acquire first position information of a first bounding box on the first panoramic image, where the first position information is used to indicate spherical coordinates of vertices of the first bounding box on the first panoramic image, where the spherical coordinates include spherical coordinates in a longitude direction and spherical coordinates in a latitude direction; The acquisition module is further configured to acquire second position information of a second bounding box on the first panoramic image, where the second position information is used to indicate spherical coordinates of vertices of the second bounding box on the first panoramic image; a generating module, configured to generate a spherical area of ​​a target intersection region of the first bounding box and the second bounding box on the first panoramic image based on the first position information and the second position information; A determination module is configured to determine an intersection-over-union ratio (IoU) of the first bounding box and the second bounding box on the first panoramic image based on a spherical area of ​​the target intersection region.

15. The device according to claim 14, characterized in that If the first bounding box and the second bounding box intersect, the generating module is specifically configured to: generating spherical coordinates of vertices of the target intersection area on the first panoramic image according to the spherical coordinates of vertices of the first bounding box on the first panoramic image and the spherical coordinates of vertices of the second bounding box on the first panoramic image; The spherical area of ​​the target intersection area is generated according to the spherical coordinates of the vertices of the target intersection area on the first panoramic image.

16. A computer program product, characterized in that When the computer program runs on a computer, it enables the computer to execute the method according to any one of claims 1 to 4, or enables the computer to execute the method according to claim 5 or 6, or enables the computer to execute the steps executed by the second model according to any one of claims 7 to 9.

17. A computer-readable storage medium, characterized in that The method comprises a program which, when run on a computer, causes the computer to execute the method according to any one of claims 1 to 4, or causes the computer to execute the method according to claim 5 or 6, or causes the computer to execute the steps executed by the second model according to any one of claims 7 to 9.

18. An electronic device, characterized in that: comprising a processor and a memory, the processor being coupled to the memory, The memory is used to store programs; The processor is used to execute the program in the memory so that the electronic device performs the method as described in any one of claims 1 to 4, or the electronic device performs the method as described in claim 5 or 6, or the electronic device performs the steps performed by the second model in any one of claims 7 to 9.

Citation Information

Patent Citations

  • Hybrid model traffic signal detection method and system based on intersection information

    CN110728170A

  • Multi-sensor image fusion method and system

    CN112233079A