Orbit-based image set classification
Image set classification using metadata to identify orbits and classify scenes into distinct site types addresses the issue of irrelevant data in 3D reconstruction, ensuring accurate and efficient 3D pointcloud generation.
Patent Information
- Application Number
- PCT/EP2024/073937
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-05
AI Technical Summary
Existing 3D reconstruction techniques for image sets captured by UAVs often result in pointclouds with excess data points that are irrelevant to the scene of interest, leading to decreased resolution or loss of important details, as they fail to balance between cropping too much or too little.
Implement image set classification using metadata to identify orbits and classify scenes into distinct site types, allowing for targeted cropping of 3D pointclouds to include only relevant data points.
Enables the generation of 3D pointclouds that are accurately confined to areas of interest, improving resolution and detail preservation by allocating datapoints only to necessary regions.
Smart Images

Figure EP2024073937_05032026_PF_FP_ABST
Abstract
Description
[0001]ORBIT-BASED IMAGE SET CLASSIFICATION TECHNICAL FIELD Embodiments presented herein relate to a method, an image processing device, a computer program, and a computer program product for image set classification. BACKGROUND As an introductory example, digital twin (DT) portals can be used to collect raw scans (in terms of two-dimensional (2D) images), of real-world sites and generate different representations of them. Many scans are obtained by guiding an unmanned aerial vehicle (UAV) to capture a plurality of images, defining an image set, along one or more orbits around a scene of interest. For example, in the telecommunication industry, cell site installations ranging from large radio towers to telecommunication equipment mounted on building rooftops and / or walls can be scanned and the resulting images can be uploaded to a DT portal for a three-dimensional (3D) pointcloud, spin and / or SkyBox views, business information modelling (BIM) models, etc. to be generated of the scene of interest. Other examples of scenes of interest are power equipment sites, industry sites, cultural heritage sites, just to mention a few. UAVs can in this way be used for scanning a variety of different types of sites. For the creation of the DT, a 3D reconstruction process has to be performed on the collected image set. In this respect, 3D reconstruction can be facilitated using different techniques, for example by employing neural volumetric representations, such as Neural Radiance Fields (NeRF) modelling. NeRF modelling inter alia address the problem of learning a 3D scene representation Ω from sparse input views (defined by the image set ^ and the camera poses ^ from which the images in the image set were captured). Here, the camera poses can be provided by a prior Structure-from-Motion (SfM) process applied to the images ^ ∈ ^. In further detail, the NeRFmodelling yields a scene representation in the form of a continuous five-dimensional (5D) function that outputs radiance (i.e., color) emitted in each direction ^^,^^, ateach point in 3D space ^^, ^, ^^. The pointcloud is extracted through sampling thelearned function at arbitrary pixels from various camera views. In particular, each sampled pixel, that has assigned color and depth values rendered from the NeRF model, directly corresponds to a point in 3D space; these sampled points are aggregated into a single pointcloud Ω. The 3D reconstruction process can then be triggered once the image set has been collected and uploaded to a server, database, or the like. In other words, NeRF can be used to learn a 5D vector-valued function whose input is a 3D location (x,y,z) and a two-dimensional (2D) viewing direction (θ, φ), and whose output is emitted color (defined by the radiance) r and a volume density (structure) σ, also referred to as structure. Reference is here made to the block diagram 100 in Fig.1. According to the block diagram 100, camera poses C corresponding to a set of input images I are estimated by a Structure from Motion (SfM) processing block 110. A continuous 3D representation Ω of the scene is, based on the camera poses C and the set of input images I, learned by neural network (NN) training in a NeRF processing block 120. The 3D representation Ω can then be exposed to a user, or some other application, by a NeRF viewer block 130. After a NeRF model is learned, it can be used to render scene views from either user- controlled or some predefined input camera poses. Taking the telecommunication industry as a non-limiting and illustrative example, different types of cell site installations can be scanned by UAVs in a variety of ways, using different combinations of flight paths. The images would then typically capture a scene covering a large physical area, whilst the 3D reconstruction might be required to model only part of the captured scene. However, these requirements can vary. For some cell site installations the scene of interest includes only the telecommunication equipment itself, whilst for other cell site installations sites the scene of interest includes not only the telecommunication equipment itself but also its surroundings, such as the building, tower, etc. at which the telecommunication equipment is installed. For yet further cell site installations sites, even part of the environment is of interest in order for an additional architectural model of the surrounding environment to be generated. Therefore, for at least some scenes of interest, unbounded pointcloud sampling of the modelled 3D space might include a significant amount of excess area, representing part of the scene that is of no interest for the scenario at hand. Assuming a given requirement on the total number of datapoints in the 3D pointcloud, this could lead to decreased resolution for the objects in the scene of interest. However, on the other hand, too much cropping of the captured scene (i.e., of the 3D pointcloud) can lead to removing important details of the surrounding that, for example, might be required in BIM model generation. Hence, there is a need for techniques that achieve a balance between the listed above extremes. SUMMARY An object of embodiments herein is to address the above issues in order to enable the 3D pointcloud to only be composed of datapoints that are of relevance for the scenario at hand. A particular object is to use image set classification for enabling a 3D pointcloud with a requirement on the total number of datapoints to be composed only of datapoints that are of relevance for the scenario at hand. A particular object is to perform image set classification using metadata as available for the images in the image set. According to a first aspect there is presented a method for image set classification. The method is performed by an image processing device. The method comprises obtaining an image set comprising a plurality of images. Each of the plurality of images has its own metadata. The method comprises identifying orbits along which the plurality images were captured, using the metadata of each of the plurality of images. The method comprises classifying the image set as depicting a scene of one distinct site type. The site type is selected from a set of candidate site types based on the identified orbits, including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits, including their relative positions, shapes and sizes. According to a second aspect there is presented an image processing device for image set classification. The image processing device comprises processing circuitry. The processing circuitry is configured to cause the image processing device to obtain an image set comprising a plurality of images. Each of the plurality of images has its own metadata. The processing circuitry is configured to cause the image processing device to identify orbits along which the plurality images were captured, using the metadata of each of the plurality of images. The processing circuitry is configured to cause the image processing device to classify the image set as depicting a scene of one distinct site type. The site type is selected from a set of candidate site types based on the identified orbits, including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits, including their relative positions, shapes and sizes. According to a third aspect there is presented a computer program for image set classification. The computer program comprises computer code which, when run on processing circuitry of an image processing device, causes the image processing device to perform actions. One action comprises the image processing device to obtain an image set comprising a plurality of images. Each of the plurality of images has its own metadata. One action comprises the image processing device to identify orbits along which the plurality images were captured, using the metadata of each of the plurality of images. One action comprises the image processing device to classify the image set as depicting a scene of one distinct site type. The site type is selected from a set of candidate site types based on the identified orbits, including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits, including their relative positions, shapes and sizes. According to a fourth aspect there is presented a computer program product comprising a computer program according to the third aspect and a computer readable storage medium on which the computer program is stored. The computer readable storage medium could be a non-transitory computer readable storage medium. Advantageously, these aspects enable the 3D pointcloud to only be composed of datapoints that are of relevance for the scenario at hand. Advantageously, by enabling more accurate isolation of the area(s) of interest within the 3D space, and constraining the resulting pointcloud to that area(s) of interest, these aspects can be used to improves 3D pointcloud generation with 3D reconstruction from image sets with available metadata. Advantageously, this means that, given a requirement on e.g., number of datapoints in the 3D pointcloud, all those datapoints can be allocated to the area of interest and not wasted on surrounding regions for those scenarios which do not require it. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dependent claims as well as from the drawings. Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to "a / an / the element, apparatus, component, means, module, step, etc." are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, module, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS The inventive concept is now described, by way of example, with reference to the accompanying drawings, in which: Fig.1 is a block diagram comprising an SfM processing block, a NeRF processing block , and a NeRF viewer block according to an example; Figs.2 and 3 are block diagrams of image processing devices according to embodiments; Fig.4 is a flowchart of methods according to embodiments; Fig.5 schematically illustrates different types of scenes and corresponding types of orbits according to embodiments; Fig.6 is a schematic diagram showing structural units of an image processing device according to an embodiment; Fig.7 shows one example of a computer program product comprising computer readable storage medium according to an embodiment. DETAILED DESCRIPTION The inventive concept will now be described more fully hereinafter with reference to the accompanying drawings, in which certain embodiments of the inventive concept are shown. This inventive concept may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concept to those skilled in the art. Like numbers refer to like elements throughout the description. Any step or feature illustrated by dashed lines should be regarded as optional. As noted above, there is a need for techniques that achieve a balance between the extremes of cropping too little and cropping too much. The embodiments disclosed herein therefore relate to techniques for image set classification. In order to obtain such techniques there is provided an image processing device, a method performed by the image processing, a computer program product comprising code, for example in the form of a computer program, that when run on an image processing device, causes the image processing device to perform the method. As an introductory example, in the context of telecommunication sites, a distinction needs sometimes to be made between tower sites, and rooftop sites. This is because 3D pointclouds of different dimensions are needed for rooftop sites compared to tower sites. For rooftop sites, the expected pointcloud size might be in excess of the position span of the UAV, as the surrounding architecture is also part of the area of interest, whereas for tower sites, the only area of interest is the tower itself and not the surrounding area. In this respect, hereinafter are disclosed techniques for identifying distinct flight path segments, hereinafter referred to as orbits, and classifying image sets based on the identified orbits. In general terms, the classification is based on analyzing metadata of the images in the image set to detect the orbits. This orbit-analysis enables site classification. This information can be used for subsequent processing of the image set. Fig. 2 is a block diagram of an image processing device 200 configured for image set classification according to an embodiment. The image processing device 200 is configured to receive as input an image set 204 as, for example, provided via user file upload 202. An image set classification block 206 is then applied to classify the image set as representing images captures of a scene of one distinct site type. A 3D reconstruction block 208 is configured to generate a 3D pointcloud , as explained above with reference to Fig.1. Based on the site type the image set was classified as, the 3D pointcloud is then subjected to further processing, as in Fig.2 represented by cropping. In this respect, whether the 3D pointcloud is provided to a wide crop block 212 or a narrow crop block 214 is dependent on the site type, and this is facilitated by the decision block 210 that uses the information as provided by the image set classification block 206 in order to direct the 3D pointcloud to either the wide crop block 212 or the narrow crop block 214. The thus cropped 3D pointcloud (possibly also information) is saved in a database (DB) 216. A pointcloud 218 with appropriate cropping can then be provided to one or more poincloud consumers 220a:220N. Further details of the image set classification block 206 will be provided next. Fig. 3 is a block diagram of an image processing device 300 according to an embodiment. This block diagram illustrates an implementation of site classification based on orbit analysis. To avoid clutter, the notation has been very slightly simplified compared to the description below (using “ Δ^” instead of “ ^Δ^^௧,Δ^^^,Δ^^௧^ of orbit / subset ^^”). The block diagram in Fig.3 provides an example implementation of the image set classification block 206 in Fig.2. The image processing device 300 comprises a number of blocks 302:314. The function of each block will be disclosed below in conjunction with the description of the method in the flowchart of Fig.4. Fig. 4 is a flowchart illustrating embodiments of methods for image set classification. The methods are performed by the image processing device 200, 300. The methods are advantageously provided as computer programs. S102: The image processing device 200, 300 obtains an image set ^ comprising aplurality of images ^ ∈ ^. In some examples, the plurality of images are any of: 2Dred-green-blue (RGB) images, RGB images with depth information, thermal images. Each of the plurality of images has its own metadata. The metadata might be provided as exchangeable image file (EXIF) data. In some embodiments, the metadata comprises timestamps ^ூand 3D coordinates describing times and places of capture of the plurality of images. Here, the 3D coordinates could be provided in terms of any type of Global Navigation Satellite System (GNSS) coordinates. For example, the 3D coordinates for image ^ might berepresented by a global positioning system (GPS) vector ^^^^^^^^^^ூ ൌ ^latூ , lonூ , altூ^ whereeach image ^ ∈ ^ thus has a latitude, a longitude, and an altitude value. In someexamples, the metadata specifies that the image set was captured by a UAV. S104: The image processing device 200, 300 identifies orbits 410a-1, ..., 410d-2 along which the plurality images were captured. The orbits are by the image processing device 200, 300 identified using the metadata of each of the plurality of images. S106: The image processing device 200, 300 classifies the image set as depicting a scene 420a:420d of one distinct site type. Here, the site type is selected from a set of candidate site types based on the identified orbits 410a-1, ..., 410d-2, including their relative positions, shapes and sizes. Further, each of the distinct site types is associated with its own set of orbits 410a-1, ..., 410d-2, including their relative positions, shapes and sizes. In some examples, the set of candidate site types pertain to different types of telecommunication sites, or different types of power equipment sites, or different types of industry sites, or different types of cultural heritage sites. In some examples, the set of candidate site types is provided as input to the image processing device 200, 300. For example, for telecommunication sites, then the input could be {tower site, rooftop site}, for power equipment sites, then the input could be {substation, transmission tower, overhead power line}, or just different types of transmission towers, such as {suspension tower, dead-end terminal tower, tension tower, transposition tower} etc. Embodiments relating to further details of image set classification as performed by the image processing device 200, 300 will now be disclosed with continued reference to Figs.2, 3, and 4, as well as to Fig.5. Further aspects of the image processing device 200, 300 identifying the orbits 410a-1, ..., 410d-2 along which the plurality images were captured will be disclosed next. Images ^ in the image set ^ can be sorted based on their timestamp ^ூsuch that forthe ^-th image, ^^ ^ ^ ^^ା^. This gives a sorted set of images^^∗. Hence, in some embodiments, the orbits 410a-1, ..., 410d-2 are identified using a set of sorted images ^∗as obtained by sorting the plurality of images according to their timestamps. The sorting can be implemented by the sort block 302.The sorted set^^∗ can be split into continuous subsets^^^, ^ଶ, … . One reason forsplitting the set of sorted images into subsets is for each subset to represent exactly one orbit, which could simplify the processing. Hence, in some embodiments, the orbits 410a-1, ..., 410d-2 are identified using the set of sorted images as split intosubsets ^^, ^ଶ, … of images based on the timestamps. The splitting can beimplemented by the split into orbits block 304. In this respect, a subset boundary can, for example, be placed between two adjacentimages ^^ , ^^ା^ wherever their timestamp difference exceeds a predefined thresholdvalue ^. That is, such that ^^ା^ െ ^^ ^ ^^. That is, in some embodiments, as part ofsplitting the set of sorted images into the subsets of images, a subset boundary is placed between two images with adjacent timestamps whenever their timestamp difference is larger than a predefined threshold value ^. In some non-limiting examples, the predefined threshold value is larger than 5 seconds, preferably larger than 10 seconds, and even more preferably larger than 15 seconds.An example algorithm for grouping a sorted image set^^∗ into subsets^^^, ^ଶ, … ofimages is disclosed next. Given: set^^∗of ^ images ^ with timestamps ^; a threshold ^;^ _list_of_subsets = ∅ _current_subset = ∅ _for ^ in range(1,n): ____current_subset.append(^^)____if( ^ ൌൌ ^ or ^^ା^ െ ^^ ^ ^):________list_of_subsets.append(current_subset) ________current_subset = ∅; end if;^ _end for; _idx = 1 _for ^ in range(1, sizeof(list_of_subsets)):____if sizeof(list_of_subsets[^]) ^ ^2: ________^^ௗ௫= list_of_subsets[^] ________idx=idx+1; end if; _end for For each subset ^^, the range (e.g., in metric scale or some other absolute scale metric) for latitude, longitude, and altitude are calculated, as given by the 3D coordinates of the images. For each of the latitude, longitude, and altitude, the range can be calculated as the difference between the maximum and the minimum of these values. Hence, in some embodiments, the shape of the orbit for a given subset of images is defined by differences Δ between the 3D coordinates for the images in this given subset of images. Here, the latitude range can be calculated as Δ୪ୟ^= and the altitude range In some examples, the orbit center position^ ^^^^^^^^^^#^ೕis calculated as: ^ ^^^^^^^^^^#^ೕ ൌ^ ^lat#^ೕ , lon#^ೕ , alt#^ೕ ൧ where sizeof^^^^ is the size of ^^, i.e., the number of elements in ^^. In other examples, the orbit center ^^^^^^^^^^#^ೕis the point closest to the focus of all images in ^^, or the centroid of a circle or oval fitted to the positions of all images in ^^. Calculation of the ranges and the centers of the orbits can be implemented by the calculate range and center for each orbit blocks 306, 310. The initial subgroup splitting based on timestamps may be excessive, and some subsets may in fact describe different parts of a single orbit. If any two subsets have approximately the same mean latitude and longitude (representing the “horizontal center”) and approximately the same Δ^^௧and Δ^^^(representing the “horizontal range”), these subsets are merged into a single subset (and the mean position and the range are updated accordingly). Hence, in some embodiments, any two subsets of images are merged into one subset of images in case these any two subsets of images have same horizontal center and same horizontal range, as derivable from the metadata of the images in these any two subsets of images. The merging can be implemented by the merge orbits block 308.As an illustrative example, for two subsets ^^, ^^ and a merge distance threshold ^,and it is true that ^^lat#^^ , lon#^^ ൧ฮ ^ ^, then ^^ ^← ^ ^^ ∪ ^^ and ^^ is deleted. An example threshold value is^ ൌ 3 meters.Each of the non-empty subsets^^^, ^ଶ, … can then be classified as one of several orbitsbased on the orbit span in latitude and longitude. In this respect, in some non- limiting examples, the shapes of the identified orbits 410a-1, ..., 410d-2 are any of: planar orbits, stack orbits, column orbits, spot orbits, or any combination thereof. For example, a planar orbit is defined as an orbit that is more than twice as wide than tall and has a horizontal diameter of at least ^ (where the value of the ^ depends on the scenario); a stack orbit is defined as an orbit that is at least half as tall as it is wide, with a horizontal diameter of at least ^; a column orbit is defined as an orbit that is more tall than wide with a horizontal diameter of less than ^; and a spot orbit is an orbit with a horizontal diameter of less than ^, and more wide than tall. The above conditions for orbit class assignment in stricter form can be summarized as follows:If Δୟ୪^ ^ ^ ∙ ^diam௫௬ and diam௫௬ ^ ^ then ^^ is a planar orbit,else if Δୟ୪^ ^ ^ ∙ ^diam௫௬ and diam௫௬ ^ ^ then ^^ is a stack orbit,else if Δୟ୪^ ^ ^^^^^௫௬ and diam௫௬ ^ ^ then ^^ is a column orbit, andelse if Δୟ୪^ ^ ^diam௫௬ and diam௫௬ ^ ^ then ^^ is a spot orbit.Here diam௫௬ ൌThe value of ^ suggested depends on the scenario. In some examples, ^ ൌ 10 meters, and ^ ൌ 0.4.The subset ^^can then be assigned an orbit class based on the orbit shape. Hence, in some embodiments, the image processing device 200, 300 is configured to perform (optional) step S104-2 as part of identifying the orbits 410a-1, ..., 410d-2 in step S104. S104-2: The image processing device 200, 300 assigns an orbit category to each given subsets of images using the metadata of per each given subset of images. The orbit category is determined using the differences between the 3D coordinates for the images per given subset of images. The actions in step S104 and step S104-2 can be implemented by the classify orbits block 312. Further aspects of the image processing device 200, 300 classifying the image set as depicting a scene 420a:420d of one distinct site type will be disclosed next. In some aspects, the dataset ओ is classified using the orbit categories of all the orbits 410a-1, ..., 410d-2 as obtained in step S104-2. That is, in some embodiments, the image processing device 200, 300 is configured to perform (optional) step S106-2 as part of classifying the image set in step S106. S106-2: The image processing device 200, 300 assigns one distinct site type to the set of sorted images by comparing the orbit categories assigned to the subsets of images with a set of conditions. Here, a distinct set of the conditions are fulfilled for the orbit categories of each of the site types. The actions in step S106 and step S106-2 can be implemented by the classify site block 314. As a non-limiting and illustrative example, in the context of different types of telecommunication sites, a “tower site” is identifiable as having a combination of concentric planar and / or stack orbits, with optional nearby column and spot orbits that are within the range of the concentric planar and stack orbits. Conversely, a “rooftop site” is identifiable as having non-concentric planar and / or stack orbits, or column / spot orbits outside the range of planar / stack orbits. An example algorithm for classifying an image set as depicting either a “tower site” or a “rooftop site” in the context of different types of telecommunication sites is disclosed next. A similar algorithm can be utilized for other types of sites. _If ^ does not have any planar or stack orbits: ____Classify ^ as “rooftop site” _else:____If exists at least 1 pair ^^ , ^^ of planar / stack orbits where ฮ^^^^#^ೖ , ^^^#^ೖ ൧ െ ^ ^ / / not concentric; threshold ^=5 meters for example / / :________Classify ^ as “rooftop site”; ____else: / / all planar & stack orbits were concentric / / ________If no column / spot orbits: ____________Classify ^ as “tower site” ________else: / / has spot / column orbits / / ____________If every spot / column is “within range” of at least 1 planar / stack orbit: ________________ Classify ^ as “tower site” ____________else: ________________ Classify ^ as “rooftop site” ____________end if; ________end if; ____ end if; _end if; In the above, an orbit ^^being “within range” of an orbit ^ଶmeans that the horizontal center point of the first orbit, ^^^^#^భ, ^^^#^భ൧, is within the area bounded from point ^^^^#^మ , ^^^#^మ ൧ െ ^Δ^^௧,^మ ,Δ^^^,^మ൧ / 2 to point ^lat# #^మ , lon^మ ൧ ^ ^Δ^^௧,^మ ,Δ^^^,^మ൧ / 2.Further aspects of how the image set, or the 3D pointcloud generated from the image set, might be processed will be disclosed next. At least part of the processing is dependent on which of the distinct site types the image set has been classified as. Particularly, in some embodiments, the image processing device 200, 300 is configured to perform (optional) step S108. S108: The image processing device 200, 300 performs an action on the image set. At least part of the action depends on which of the distinct site types the image set has been classified as. There could be different types of actions. In some embodiments, the action involves cropping the image set or a 3D pointcloud generated from the image set. The amount of cropping is defined by which of the distinct site types the image set has been classified as. For example, a 3D pointcloud Ω could be generated from dataset ^ using e.g. SfM and NeRF as described above. Then, the 3D pointcloud could be cropped differently depending on which distinct site type the image set was classified as. This cropping can be performed after the 3D pointcloud Ω has been generated, or the cropping bounds can be part of the 3D pointcloud generation parameters. As an illustrative example, in the context of telecommunication sites, for a telecommunication site classified as a “rooftop site”, it is desirable to have a wider horizontal crop area to include more of the surrounding buildings’ architecture. In contrast, for a telecommunication site classified as a “tower site” a narrow horizontal crop area and a larger vertical area are desired, in order to only include the tower and the immediate ground vicinity, but not accidentally cropping out the top or bottom of the tower. In some examples, the 3D pointcloud comprises 3D points, and cropping the 3D pointcloud comprises applying a 3D box volume to the 3D pointcloud, where the size of the 3D box volume depends on which of the distinct site types the image set has been classified as. Those 3D points 3D pointcloud that are outside the 3D box volume are excluded from the 3D pointcloud. In further detail, one example of cropping is to use a 3D box volume centered on the image position center ( ^^^^^^^^^^^#) with box edgelengths of either ^2Δ୪ୟ^,^ , 2Δ୪୭୬,^ depending on whetherthe cropping is to be wide or narrow. In some embodiments, the action involves other type of processing than cropping and / or cropping in combination with another type of processing. This processing, and hence, the action, might involve at least one of: adjusting size of a neural radiance field obtained from a 3D pointcloud generated from the image set, adjusting resolution of the 3D pointcloud, automatic annotation of the image set, selecting a trained model for subsequent processing of the image set. In Fig.5 four scenes 420a:420d with two different types of telecommunication sites are schematically illustrated. In Figs.4(a) and 4(b) are illustrated scenes 420a, 420b representing a rooftop site. In Figs.4(c) and 4(d) are illustrated scenes 420c, 420d representing a tower site. In each of Figs.5(a)-(d) is further illustrated typical orbits 410a-1, ..., 410d-2 for each type of telecommunication site. It can be seen that the types of orbits differ for the different types of telecommunication sites. In Figs.5(a) and 5(b) several non-concentric planar orbits are visible, whereas in Figs.5(c) and 5(d) there are several concentric planar orbits, or stack orbits. In this way, if orbits as in Fig.5(a) and / or Fig.5(b) are identified, then the telecommunication site can be classified as a rooftop site. Likewise, if orbits as in Fig.5(c) and / or Fig.5(d) are identified, then the telecommunication site can be classified as a tower site. Fig. 6 schematically illustrates, in terms of a number of structural units, the components of an image processing device 600 according to an embodiment. Processing circuitry 610 is provided using any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller, digital signal processor (DSP), etc., capable of executing software instructions stored in a computer program product 710 (as in Fig.7), e.g. in the form of a storage medium 630. The processing circuitry 610 may further be provided as at least one application specific integrated circuit (ASIC), or field programmable gate array (FPGA). Particularly, the processing circuitry 610 is configured to cause the image processing device 600 to perform a set of operations, or steps, as disclosed above. For example, the storage medium 630 may store the set of operations, and the processing circuitry 610 may be configured to retrieve the set of operations from the storage medium 630 to cause the image processing device 600 to perform the set of operations. The set of operations may be provided as a set of executable instructions. Thus the processing circuitry 610 is thereby arranged to execute methods as herein disclosed. The storage medium 630 may also comprise persistent storage, which, for example, can be any single one or combination of magnetic memory, optical memory, solid state memory or even remotely mounted memory. The image processing device 600 may further comprise a communications (comm.) interface 620 at least configured for communications with other entities, functions, nodes, and devices, e.g., for receiving input to the image processing device 600 from another entity, function, node, or device, and for providing output to another entity, function, node, or device. As such the communications interface 620 may comprise one or more transmitters and receivers, comprising analogue and digital components. The processing circuitry 610 controls the general operation of the image processing device 600 e.g. by sending data and control signals to the communications interface 620 and the storage medium 630, by receiving data and reports from the communications interface 620, and by retrieving data and instructions from the storage medium 630. Other components, as well as the related functionality, of the image processing device 600 are omitted in order not to obscure the concepts presented herein. The image processing device 600 may be provided as a standalone device or as a part of at least one further device. Thus, a first portion of the instructions performed by the image processing device 600 may be executed in a first device, and a second portion of the of the instructions performed by the image processing device 600 may be executed in a second device. For example, user-input and image-rendering aspects of the NeRF viewer (including selection of measurement points, selection of label position, etc.) could be performed on a user device (such as a desktop computer, a laptop computer, a tablet computer, or a smartphone), whilst aspects of classification, view generation, etc. could be performed by at least one GPU enabled virtual machine in a localized or distributed cloud computational environment. The prerequisite steps of learning the 3D NeRF model parametrization and scaling to absolute scale could performed by (GPU enabled) virtual machines in a localized or distributed cloud computational environment. The herein disclosed embodiments are not limited to any particular number of devices on which the instructions performed by the image processing device 600 may be executed. Hence, the methods according to the herein disclosed embodiments are suitable to be performed by an image processing device 600 residing in a cloud computational environment. Therefore, although a single processing circuitry 610 is illustrated in Fig.6 the processing circuitry 610 may be distributed among a plurality of devices, or nodes. The same applies to the computer program 720 of Fig.7. Fig. 7 shows one example of a computer program product 710 comprising computer readable storage medium 730. On this computer readable storage medium 730, a computer program 720 can be stored, which computer program 720 can cause the processing circuitry 610 and thereto operatively coupled entities and devices, such as the communications interface 620 and the storage medium 630, to execute methods according to embodiments described herein. The computer program 720 and / or computer program product 710 may thus provide means for performing any steps as herein disclosed. In the example of Fig.7, the computer program product 710 is illustrated as an optical disc, such as a CD (compact disc) or a DVD (digital versatile disc) or a Blu-Ray disc. The computer program product 710 could also be embodied as a memory, such as a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or an electrically erasable programmable read-only memory (EEPROM) and more particularly as a non-volatile storage medium of a device in an external memory such as a USB (Universal Serial Bus) memory or a Flash memory, such as a compact Flash memory. Thus, while the computer program 720 is here schematically shown as a track on the depicted optical disk, the computer program 720 can be stored in any way which is suitable for the computer program product 710. The inventive concept has mainly been described above with reference to a few embodiments. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the inventive concept, as defined by the appended patent claims.
Claims
CLAIMS 1. A method for image set classification, wherein the method is performed by an image processing device (200, 300, 600), and wherein the method comprises: obtaining (S102) an image set ^ comprising a plurality of images ^ ∈ ^, whereineach of the plurality of images has its own metadata; identifying (S104) orbits (410a-1, ..., 410d-2) along which the plurality images were captured, using the metadata of each of the plurality of images; and classifying (S106) the image set as depicting a scene (420a:420d) of one distinct site type, wherein the site type is selected from a set of candidate site types based on the identified orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes.
2. The method according to claim 1, wherein the set of candidate site types is provided as input to the image processing device (200, 300, 600).
3. The method according to claim 1 or 2, wherein the set of candidate site types pertain to different types of telecommunication sites, or different types of power equipment sites, or different types of industry sites, or different types of cultural heritage sites.
4. The method according to any preceding claim, wherein the shapes of the identified orbits (410a-1, ..., 410d-2) are any of: planar orbits, stack orbits, column orbits, spot orbits, or any combination thereof.
5. The method according to any preceding claim, wherein the metadata comprises timestamps ^ூand three-dimensional coordinates ^^^^^^^^^^ூdescribing times and places of capture of the plurality of images.
6. The method according to claim 5, wherein the orbits (410a-1, ..., 410d-2) are identified using a set of sorted images ^∗as obtained by sorting the plurality of images according to their timestamps.
7. The method according to claim 6, wherein the orbits (410a-1, ..., 410d-2) areidentified using the set of sorted images as split into subsets ^^, ^ଶ, … of images basedon the timestamps.
8. The method according to claim 7, wherein, as part of splitting the set of sorted images into the subsets of images, a subset boundary is placed between two images with adjacent timestamps whenever their timestamp difference is larger than a predefined threshold value ^.
9. The method according to claim 8, wherein the predefined threshold value is larger than 5 seconds, preferably larger than 10 seconds, and even more preferably larger than 15 seconds.
10. The method according to any of claims 7 to 9, wherein any two subsets of images are merged into one subset of images in case said any two subsets of images have same horizontal center and same horizontal range, as derivable from the metadata of the images in said any two subsets of images.
11. The method according to any of claims 5 to 10, wherein the shape of the orbit for a given subset of images is defined by differences Δ between the three- dimensional coordinates for the images in said given subset of images.
12. The method according to claim 11, wherein identifying the orbits (410a-1, ..., 410d-2) comprises: assigning (S104-2) an orbit category to each given subsets of images using the metadata of said each given subset of images, wherein the orbit category is determined using the differences between the three-dimensional coordinates for the images in said given subset of images.
13. The method according to claim 12, wherein classifying the image set comprises: assigning (S106-2) one distinct site type to the set of sorted images by comparing the orbit categories assigned to the subsets of images with a set of conditions, where a distinct set of the conditions are fulfilled for the orbit categories of each of the site types.
14. The method according to any preceding claim, wherein the method further comprises: performing (S108) an action on the image set, wherein at least part of the action depends on which of the distinct site types the image set has been classified as.
15. The method according to claim 14, wherein the action involves cropping the image set or a three-dimensional, 3D, pointcloud generated from the image set, and wherein an amount of cropping is defined by which of the distinct site types the image set has been classified as.
16. The method according to claim 15, wherein the 3D pointcloud comprises 3D points, and wherein cropping the 3D pointcloud comprises applying a 3D box volume to the 3D pointcloud, where a size of the 3D box volume depends on which of the distinct site types the image set has been classified as and where those 3D points 3D pointcloud that are outside the 3D box volume are excluded from the 3D pointcloud.
17. The method according to claim 14, wherein the action involves at least one of: adjusting size of a neural radiance field obtained from a three-dimensional, 3D, pointcloud generated from the image set, adjusting resolution of the 3D pointcloud, automatic annotation of the image set, selecting a trained model for subsequent processing of the image set.
18. The method according to any preceding claim, wherein the metadata specifies that the image set was captured by an unmanned aerial vehicle, and wherein the plurality of images are any of: two-dimensional, 2D, red green blue, RGB, images, RGB images with depth information, thermal images.
19. An image processing device (200, 300, 600) for image set classification, the image processing device (200, 300, 600) comprising processing circuitry (610), the processing circuitry being configured to cause the image processing device (200, 300, 600) to: obtain an image set ^ comprising a plurality of images ^ ∈ ^, wherein each of theplurality of images has its own metadata;identify orbits (410a-1, ..., 410d-2) along which the plurality images were captured, using the metadata of each of the plurality of images; and classify the image set as depicting a scene (420a:420d) of one distinct site type, wherein the site type is selected from a set of candidate site types based on the identified orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes.
20. The image processing device (200, 300, 600) according to claim 19, further being configured to perform the method according to any of claims 2 to 18.
21. A computer program (720) for image set classification, the computer program comprising computer code which, when run on processing circuitry (610) of an image processing device (200, 300, 600), causes the image processing device (200, 300, 600) to: obtain (S102) an image set ^ comprising a plurality of images ^ ∈ ^, whereineach of the plurality of images has its own metadata; identify (S104) orbits (410a-1, ..., 410d-2) along which the plurality images were captured, using the metadata of each of the plurality of images; and classify (S106) the image set as depicting a scene (420a:420d) of one distinct site type, wherein the site type is selected from a set of candidate site types based on the identified orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes, where each of the distinct site types is associated with its own set of orbits (410a-1, ..., 410d-2), including their relative positions, shapes and sizes.
22. A computer program product (710) comprising a computer program (720) according to claim 21, and a computer readable storage medium (730) on which the computer program is stored.
Citation Information
Patent Citations
Acquisition mode recognition method and system for aerial data of unmanned aerial vehicle
CN113867410A