Method and system for detecting objects in physical environment
By combining 3D point cloud images and 2D pixel images, using color profiles and computer vision technology, the efficiency and accuracy of object detection and positioning in complex factory environments are solved, and fast and accurate object recognition and positioning are achieved.
Patent Information
- Application Number
- CN202280101403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-06-10
AI Technical Summary
In complex factory environments, processing high-resolution point cloud data to detect and locate objects in the physical environment is a heavy and slow process and prone to false positives, requiring manual analysis to filter results.
By receiving or obtaining a first image (3D point cloud image) and a second image (2D pixel image) representing the physical environment, the objects in the second image are detected and the corresponding object area is positioned in the first image by color profiles or other computer vision techniques.
This method can quickly identify objects in the factory environment, reduce resource consumption for processing point cloud data, reduce the number of false positive matches, and improve the efficiency of detection and positioning.
Smart Images

Figure CN120129928A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to computer-aided design, visualization, and manufacturing (“CAD”) systems, product lifecycle management (“PLM”) systems, product data management (“PDM”) systems, production environment simulation, and similar systems for managing data of products and other items (collectively referred to as “product data management” systems or PDM systems). More specifically, the present disclosure relates to digital representations of physical environments. Background Art
[0002] Three-dimensional (“3D”) digital models of physical environments are used for various tasks and purposes. For example, the usage of 3D representations of factories or manufacturing assets can include, but is not limited to, manufacturing process analysis, manufacturing process simulation, equipment collision checking, and virtual commissioning.
[0003] As used herein, the terms manufacturing asset and device denote any resource, machine, part, and / or any other object such as a machine present in a manufacturing line (or more generally, in a physical environment). Examples of devices included in a physical or real manufacturing environment include, but are not limited to: industrial robots and their tools; transportation assets such as conveyors, turntables; safety assets such as fences, gates; automation assets such as fixtures, grippers, fixtures for grasping parts, etc.
[0004] In such a physical environment, one problem remains the detection and positioning of assets. In fact, it may happen that in a factory, a device moves from one location to another or new equipment is installed. These changes (especially those regarding the positioning of equipment or new installations) must be entered into the factory's IT system. The simplest way is for an operator to physically tour the factory to identify the assets, determine their current location, and update the database of the IT system with the corresponding data. However, such manual identification is time-consuming and inefficient.
[0005] Today, digital solutions facilitate such tasks. For example, a scanner can automatically scan the current layout of a physical environment (such as a production line in a factory) and automatically identify different assets using image processing techniques known in the art. In particular, point clouds (i.e., digital representations of physical objects or environments by a set of data points in space) are becoming increasingly relevant to applications in the industrial world. For example, a 3D scanning camera device can create a point cloud by determining a large number of points on the surface of a physical environment, and such point cloud technology can then be used for complex analysis and design of various factories, automotive manufacturing lines, microcircuit manufacturing centers, or any other industrial environment. Thus, obtaining a point cloud using a 3D scanner enables a 3D image of a scene (such as a production line in a workshop) to be quickly obtained, the 3D image including positioning information of each acquired point relative to the surrounding space. The ability of point cloud technology to quickly provide a current and correct representation of the object of interest is of great significance for decision-making and task planning, as it shows the latest and accurate state of the workshop. Additionally, based on the point cloud, image reconstruction techniques such as 2D or 3D images of the environment or object can then be used. Meshing techniques are configured to create a 3D mesh from the points of the cloud, thereby converting the point cloud into a 3D surface. Today, it is possible to have a meshing tool automatically create such a mesh, or even directly create a CAD model from the entire point cloud scene.
[0006] Thus, point cloud data can be used to detect and locate objects in a physical environment such as a factory. However, for a complex factory, processing point clouds is a burdensome and slow process, especially in the case of high-resolution scans that produce point clouds with millions of points. Additionally, point cloud-based techniques may produce many false positives, which also require manual analysis to filter the results.
[0007] Therefore, improved techniques for detecting and locating objects in a physical 3D environment are desired. SUMMARY OF THE INVENTION
[0008] Various disclosed embodiments include methods, systems, and computer-readable media for detecting and locating objects in a physical environment. The methods include: i) receiving or obtaining a first image representing the physical environment, where the first image is a 3D point cloud image that includes positioning data for points in the point cloud image, ii) receiving or obtaining a second image representing the physical environment, where the second image is a 2D pixel image of the physical environment, iii) detecting the object in one or more regions of the second image, iv) for each region in the second image where the object has been detected, finding the corresponding region in the first image, and v) providing, via an interface, the corresponding region in the first image as the location where the object has been detected, and optionally, extracting from the first image the location of the object associated with positioning data from at least one point in the corresponding region in the first image.
[0009] A computing system including a processor and an accessible memory or database is also disclosed, where the data processing system is configured to perform the previously described method.
[0010] The present invention also provides a non-transitory computer-readable medium encoded with executable instructions that, when executed, cause one or more data processing systems to perform the previously described method.
[0011] The foregoing has outlined rather broadly the features and technical advantages of the present disclosure so that those skilled in the art may better understand the detailed description below. In the following, additional features and advantages of the subject matter forming the claims of the present disclosure will be described. Those skilled in the art will understand that they can readily use the disclosed concepts and specific embodiments as a basis for modifying or designing other structures for the same purpose of implementing the present disclosure. Those skilled in the art will also recognize that such equivalent constructs do not depart from the spirit and scope of the present disclosure in its broadest form.
[0012] Before proceeding with the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document: The terms "include" and "comprise" and their derivatives mean including but not limited to; the term "or" is inclusive and means and / or; the phrases "associated with" and "associated therewith" and their derivatives may mean including, included within, interconnected with, contain, contained within, connected to or coupled with, coupled to or coupled with, communicable with, cooperate with, interleave, juxtapose, adjacent, bound to or bound with, have, having the nature of, etc.; and the term "controller" means any device, system or part thereof that controls at least one operation, whether such device is implemented in hardware, firmware, software or some combination of at least two of hardware, firmware, software. It should be noted that the functions associated with any particular controller may be centralized or distributed, whether local or remote. Definitions for particular words and phrases are provided throughout this patent document, and those of ordinary skill in the art will understand that such definitions apply to many, if not most, existing instances and future uses of such defined words and phrases. Although some terms may encompass a variety of embodiments, the appended claims may expressly limit these terms to specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals designate like elements and in which:
[0014] Figure 1 A block diagram of a computing system in which embodiments may be implemented is shown.
[0015] Figure 2 A flowchart depicting a preferred embodiment of a method for detecting and locating an object in a physical environment in accordance with the present invention is shown.
[0016] Figure 3A A first image in accordance with the present invention is schematically shown.
[0017] Figure 3B A second image in accordance with the present invention is schematically shown.
[0018] Figure 3C A third image in accordance with the present invention is schematically shown.
[0019] Figure 4 A flowchart depicting a preferred first embodiment of detecting an object in accordance with the present invention is shown.
[0020] Figure 5 Shows a flowchart depicting a preferred second embodiment of detecting an object according to the present invention.
[0021] Figure 6 Shows an example of clustering an object into primary colors.
[0022] Figure 7 Schematically shows the detection of an object in a second image according to the present invention.
[0023] Figure 8 Shows the conversion of a panoramic pixel image to a cube map. Detailed Description
[0024] The following discussion Figures 1 to 8 and the various embodiments used in this patent document to describe the principles of the present disclosure are for illustrative purposes only and should not be construed in any way as limiting the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any suitably arranged device. Many innovative teachings of the present application will be described with reference to exemplary non-limiting embodiments.
[0025] Today, 3D mapping of a physical environment (such as an industrial indoor manufacturing line) can be performed by a laser scanner (such as a three-axis laser scanner), where, by scanning the physical environment, the laser scanner generates a cloud of points, where each point is characterized by coordinates defined in a reference frame, which is typically associated with the position of the laser scanner. Additionally, the laser scanner can integrate a camera device such as a panoramic camera device for acquiring a pixel image of the physical environment such as a panoramic image, especially simultaneously with the acquisition of the point cloud. In other words, from the same viewpoint corresponding to the position of the laser scanner in the physical environment, two images can be acquired, namely a first image representing the point cloud of the physical environment and a second image being the pixel image of the physical environment, preferably an equirectangular panoramic image of the physical environment. Preferably, the first image and the second image are acquired simultaneously. The present invention proposes to use the two images to improve the detection and positioning of objects (such as equipment) in the physical environment. Although preferably, the first image and the second image according to the present invention can be captured by a single device including both a laser scanner and a camera device, and thus captured from the same viewpoint, it is also contemplated within the present invention to acquire the first image from a first viewpoint and the second image from a second viewpoint different from the first viewpoint, and then orient the point cloud according to techniques known in the art to match the viewpoint of the second image. It is still crucial that the first image and the second image are images of the same physical environment, i.e., images of the same real scene.
[0026] Figure 1 FIG. 1 shows a block diagram of a computing system 100 (e.g., a data processing system) in which embodiments may be implemented as, for example, a PDM system that is software or otherwise specially configured to perform the processes described herein, and more particularly implemented as each of the various interconnected and communication systems described herein. The illustrated computing system 100 may include a processor 102 connected to a secondary cache / bridge 104, which in turn is connected to a local system bus 106. The local system bus 106 may be, for example, a Peripheral Component Interconnect (PCI) architecture bus. In the illustrated example, a main memory 108 and a graphics adapter 110 are also connected to the local system bus. The graphics adapter 110 may be connected to a display 111.
[0027] Other peripheral devices, such as a local area network (LAN) / wide area network / wireless (e.g., WiFi) adapter 112, may also be connected to the local system bus 106. An expansion bus interface 114 connects the local system bus 106 to an input / output (I / O) bus 116. The I / O bus 116 is connected to a keyboard / mouse adapter 118, a disk controller 120, and an I / O adapter 122. The disk controller 120 may be connected to a storage device 126, which may be any suitable machine-usable storage medium or machine-readable storage medium, including but not limited to: non-volatile hard-coded type media such as read-only memory (ROM) or electrically erasable programmable read-only memory (EEPROM), magnetic tape storage devices, and user-recordable type media such as floppy disks, hard disk drives, and compact disc read-only memory (CD-ROM) or digital versatile disc (DVD), as well as other known optical, electrical, or magnetic storage devices.
[0028] In the illustrated example, an audio adapter 124 is also connected to the I / O bus 116, and speakers (not shown) may be connected to the audio adapter 124 to play sound. The keyboard / mouse adapter 118 provides a connection for pointing devices (not shown) such as a mouse, trackball, trackpointer, touch screen, etc.
[0029] Optionally, an imaging device including a laser scanner and a camera device is part of or connected to the computing system 100 to provide the first image and the second image of the physical environment to the computing system 100, wherein the laser scanner and the camera device are configured to preferably image the same scene of the physical environment simultaneously. Finally, the first image and the second image may be displayed on the display 111 successively or simultaneously.
[0030] Those of ordinary skill in the art will understand that Figure 1The hardware shown may vary for a particular implementation. For example, other peripheral devices such as an optical disk drive may be used in addition to or in place of the shown hardware. The shown examples are provided for illustrative purposes only and are not meant to imply architectural limitations of the present disclosure.
[0031] The computing system 100 according to an embodiment of the present disclosure may include an operating system employing a graphical user interface. The operating system allows multiple display windows to be presented simultaneously in the graphical user interface, where each display window provides an interface to a different application or to different instances of the same application. The cursor in the graphical user interface may be manipulated by a user via a pointing device. The position of the cursor may be changed and / or events such as clicking a mouse button may be generated to drive a desired response.
[0032] One of various commercial operating systems such as the product Microsoft Windows of Microsoft Corporation located in Redmond, Washington may be employed with suitable modifications. The operating system is modified or created according to the present disclosure as described. TM versions.
[0033] The LAN / WAN / wireless adapter 112 may be connected to a network 130 (not part of the computing system 100), which may be any public or private data processing system network or a combination of networks including the Internet as known to those skilled in the art. The computing system 100 may communicate with a server system 140 via the network 130, which is also not part of the computing system 100 but may be implemented as, for example, a separate data processing system.
[0034] Figure 2 A flowchart of a method for detecting and locating an object in a physical environment is shown. The method will be described in detail below in conjunction with Figure 3A and Figure 3B and Figure 3A and Figure 3BThe first image 301 and the second image 302 of the exemplary and non-limiting physical environment 300 according to the present invention are presented respectively. The first image 301 is a point cloud image of the physical environment. The second image 302 is a pixel image of the physical environment 300. The first image and the second image can be acquired simultaneously or one after another in a short time (e.g., the time interval between the end of the acquisition of one of the images and the start of the acquisition of the other image is less than 30 seconds). The physical environment 300 is, for example, a manufacturing line or a packaging line or any other environment including one or several objects, where at least one object must be positioned and thus also detected. The object can be some furniture such as a chair or a table 310, or equipment of a production line such as robots 320, 330, or a specific component or tool of a robot such as a jaw or a wrench 321, or a robot arm 322, 332, 333, or any other object or equipment as part of the physical environment 300. For the purpose of illustrating the method according to the present invention, the table 310 will be the object to be detected and positioned.
[0035] At step 210, the computing system 100 according to the present invention especially acquires or receives the first image 301 as the point cloud of the physical environment 300 from a 3D laser scanner. The first image 301 can be acquired by a laser scanner that is part of the computing system according to the present invention or connected to the computing system, and the laser scanner is configured to scan the physical environment 300 (e.g., a manufacturing production line), and collect point cloud data from the scan, that is, a set or several sets of data points in space, where each point position is characterized by a set of position coordinates. The points represent the outer surfaces of the objects in the physical environment, and thus the laser scanner records in the point cloud data information about the positions of a plurality of points on the outer surfaces of the objects surrounding the laser scanner in the space, and thus a 2D or 3D image of the physical environment (the physical environment around which the points have been collected) surrounding it can be reconstructed from the point cloud data. Of course, the present invention is not limited to this specific type of scanner, and the first image can be received or acquired from any other type of scanner configured to output such point cloud data when acquiring the point cloud image of the physical environment. The computing system 100 can also acquire or receive the first image 301 from another computing system, from a database, from a memory (e.g., a memory stick).
[0036] At step 220, computing system 100 receives or obtains the second image 302 of the physical environment 300. The second image 302 may be obtained by an imaging device system of a laser scanner. Preferably, the first and second images according to the present invention are obtained using the same viewing point relative to the physical environment. Preferably, the second image 302 is a panoramic image of the physical environment 300. The second image 302 may be, for example, a 360° panoramic image of the physical environment 300. If the first and second images do not share the same viewing point, the computing system according to the present invention may be configured to automatically determine, within the 3D space defined by the point cloud, a viewing point that matches the viewing point used to obtain the second image 302, and optionally configured to orient the first image accordingly (e.g., automatically orient the point cloud so as to be able to display the physical environment from the same viewing point), such that the first image 301 and the second image 302 may represent the same scene when, for example, simultaneously or successively displayed on a display system. Computing system 100 may also obtain or receive the second image 302 from another computing system, from a database, from a memory (e.g., a memory stick), or from another imaging device system.
[0037] At step 230, computing system 100 is configured to detect 230 the object 310 in one or several regions of the second image 302. In other words, the present invention proposes to detect an object in a pixel image. Thus, for this detection, the cloud point image (first image) is not used. This offers the advantage of reducing the resources required to perform this task (e.g., the amount of memory and time for this task). In fact, as already explained, object detection using a cloud point image can be very resource-consuming, especially in high-resolution scans. According to the present invention, different techniques may be used to detect the object 310 in the second image 302.
[0038] According to a first embodiment 400, a color profile associated with the object 310 may be used to detect the object. In this case, the detection according to the present invention comprises the following steps shown by Figure 4 the following steps:
[0039] At step 231, computing system 100 is configured to find pixels in the second image 302 whose color values fall within a range of color values defined by a clustering of color values representing the dominant color of object 310. According to the present invention, object 310 may be associated with one or several clusters of color values, where each cluster represents the most dominant color or one of the most dominant colors of object 310, i.e., each object to be detected and located may be associated with a subset of colors representing the dominant color of the object, and then the subset of colors is used to detect the object in the pixel image. To this end, computing system 100 particularly includes a database or library, which is configured to store a color profile for each of one or several objects that may be part of the physical environment 300, where the color profile of an object defines one or several ranges of color values, and each range of color values represents one of the dominant colors of the object. Preferably, RGB values are used for the color values of the objects.
[0040] Preferably, for each of the relevant objects 310, 61 for which the computing system has determined one or several ranges of color values, the computing system also calculates a reference histogram, which is a color histogram of the pixels of the relevant object 310 (i.e., the image of the relevant object 310, 61). Preferably, such a reference histogram is saved in the database or library for each relevant object.
[0041] Therefore, detecting object 310 using color profile technology requires pre - building a database or library that includes or stores the color profiles of the dominant colors of each object that may be relevant to object detection and location according to the present invention, e.g., for each object, storing its dominant color values. This method takes advantage of the fact that industrial equipment is typically characterized by a specific color scheme according to the equipment type / supplier, thus allowing for rapid identification based on color profiles. In particular, since the number of different types of assets / objects in a physical environment (e.g., an industrial or manufacturing environment) is limited, color profiles can be easily captured for each type of asset / object, especially for important types of equipment in the physical environment (e.g., robots, cranes, etc.), and then the color profiles can be stored in the library or database.
[0042] To populate the library or database with the color profiles of each relevant object of the physical environment (i.e., each object that may be relevant to object detection and location according to the present invention), one or several images of the relevant object are analyzed and / or processed by computing system 100, where the analysis and / or processing includes, in particular, using the k - means algorithm to cluster the pixel colors belonging to the object into groups. In Figure 6The clustering process is schematically shown for a relevant object 60, where clustering 6A results in four different dominant colors represented by four different clusters of pixel color values. The k-means clustering technique enables, for example, the creation of k groups (or clusters) of pixel colors (e.g., 4 groups of pixel colors according to Figure 6 ), where each group is characterized by a mean color value representing a dominant color, and where each pixel of the object is classified into the group characterized by the mean color value that is closest (especially in terms of RGB values) to its color value. In Figure 6 , the colors of bars 61 each represent the mean color value (i.e., the dominant color) of the relevant object, and the size (length) of color bar 61 is proportional to the number of such pixels of the relevant object that are characterized by color values "within the range" of the mean color value defined for the relevant group (or bar), i.e., by the color value that is closest (e.g., in terms of the difference in RGB values) to the mean color value defined for the relevant group (or bar). Preferably, a threshold is used to automatically discard pixel groups that include a lower number (e.g., less than 20% or 10%) of pixels relative to the total number of pixels of the relevant object, so that only the "most" dominant colors (i.e., the remaining clusters after the discarding step) are retained. For example, in Figure 6Among them, groups including less than 10% of the pixels of the relevant object are automatically discarded 6B, resulting in two dominant colors D1 and D2 in this particular case. Thus, for each relevant object, one or several clusters of pixel color values can be determined. For the same object, each cluster of pixel color values represents a different dominant color of the relevant object. Dominant colors represented by a low number of pixels compared to the total number of pixels of the object may be discarded to retain only the dominant colors. Hereinafter, no distinction will be made between the dominant colors and the main colors, and the term "main color" also includes "dominant color". Then, the detection objective is to find in the second image the pixels belonging to one of the clusters (or, in the case where there is only one main color, belonging to the cluster), and these clusters (or this cluster, in the case of only one main color) have been determined for the object to be detected and / or located. To this end, a range of color values is defined for each color cluster representing the main color of the object. Thus, one or several ranges of color values can be defined in the color profile of the object, and then one or several ranges of color values are used to determine whether the pixels of the second image belong to the said color profile. For example, the range of color values can be defined according to the mean color value obtained for the cluster. Alternatively, the range of color values can be defined according to the lowest color value and the highest color value of the cluster (i.e., from the darkest color value and the brightest color value of the cluster). Thus, depending on the cluster of color values representing the main color, a person skilled in the art can use different ways to define the range of said values. For example, consider color values defined by an RGB triple (R, G, B), where R, G, B have values between 0 and 255. Also consider that the main color of the relevant object is "red", and the mean color value of said main color is the RGB value (160, 25, 25). According to the first embodiment, the range of color values for a given main color may include all color values (Ri, Gi, Bi), where Ri_min < Ri < Ri_max, Gi_min < Gi < Gi_max, and Bi_min < Bi < Bi_max, where in the cluster represented by the said main color, K_min and K_max respectively represent the lowest color value and the highest color value for the color K being considered, where K = Ri, Bi, Gi. For example, in the case of a red main color, it may be 110 < Ri < 200, 0 < Gi < 50, and 0 < Bi < 25. In this case, the range of color values can be considered or regarded as a bounding box surrounding or at least partially enclosing the cluster of color values representing the main color. Alternatively, if the mean color value of the said main color is the RGB value (160, 25, 25), the range can be defined according to the said mean color value, for example, by defining intervals according to each of the RGB values Rl = 160, Gl = 25, and B1 = 25.More generally, if the mean color value is defined by an RGB triple (R1, G1, B1), where R1, G1, B1 have values between 0 and 255, the range can be defined as ([Rl - dl (where Rl - dl is positive; otherwise 0), Rl + d2 (where it is less than 255; otherwise 255)], [Gl - d3 (where Gl - d3 is positive; otherwise 0), Gl + d4 (where it is less than 255; otherwise 255)], [Bl - d5 (where Bl - d5 is positive; otherwise 0), Bl + d6 (where it is less than 255; otherwise 255)]), where, for example, d1, ……, d6 are integers included between 5 and 10. In particular, d1 = d2 = …… = d6. Preferably, the number of dominant colors for each object in the database or library is at most two, for example, two clusters including the most pixels. Alternatively, additional (i.e., more than two) dominant colors can be considered. Keeping the number small ensures fast and efficient detection of objects in the second image. For example, the computing system can be configured to select only those associated with such groups of pixels in the mean color value, where the number of pixels in the group is higher than a predefined ratio when compared to the total number of pixels of the associated object. The mean color value of the selected group then becomes one of the dominant colors, for which the range of color values is then defined and used for detection purposes. Regardless of the technique used to select or determine a set of dominant colors for the associated object, the dominant colors are and remain a subset of the color values / colors of the associated object, which subset is configured to enable identification or detection of the associated object in the pixel image of the physical environment. At the end of the analysis or processing of the one or several images for each associated object, the computing system is configured to save in the library or database the range of color values that has been determined or defined for each dominant color of the associated object.
[0043] Preferably, the range of the color values can be automatically defined during the clustering process for each mean color value or each cluster, for example, by determining for each cluster the minimum color value and the maximum color value among the color values of the pixels grouped in the cluster. Then, the range provides one or several color value intervals extending, for example, from the minimum color value to the maximum color value.
[0044] During the detection process, the computing system 100 automatically determines for each pixel of the second image whether the color value of the pixel falls within at least one of the ranges of color values defined for the main color of the object to be detected. If so, the system considers the pixel to belong to the object to be detected; otherwise, the pixel is discarded or ignored. Preferably, the computing system leaves only the pixels in the second image that have been determined to belong to the object (i.e., belong to one of the ranges), and removes the other pixels from the second image. Optionally, and particularly, after the removal of the pixels that are considered not to belong to the object to be detected, the computing system may run morphological transformations of erosion followed by dilation in order to remove noise from the second image. Preferably, the computing system 100 is then configured to convert the second image (e.g., the remaining pixels) into a binary image. In particular, at step 232, the computing system may, for example, convert the detected pixels (i.e., the pixels that are considered to belong to the object to be detected) to white and convert all other pixels of the second image 302 to black. The result of such a conversion is shown in Figure 7 where only the pixels of the table 310 and the robot 330 that have colors falling within the range of the main color of the object to be detected remain in the image.
[0045] At step 233, the computing system 100 is configured to identify one or more shapes formed by one or several groups / clusters of detected pixels. Optionally, at step 234, the computing system is configured to surround each identified shape with bounding boxes 71, 72, as shown in Figure 7 .
[0046] To remove false positives, at step 235, the computing system 100 is configured to: for each identified shape, compare the color histogram of the pixels of the second image 302 that belong to the identified shape with the reference histogram (i.e., the color histogram of the pixels of the object 310 for which the range of color values has been determined), and if the comparison results in a color histogram difference higher than a predefined threshold (i.e., if the difference between the reference histogram and the color histogram of the considered shape is higher than the predefined threshold), the computing system is configured to discard the identified shape; otherwise, the computing system is configured to identify and / or store, at step 236, the region in the second image 302 where the shape corresponding to the object 310 to be detected and located has been identified. In particular, to identify the region, the computing system may be configured to surround the identified shape with a bounding box 71 (if not already implemented). For example, in the case shown in Figure 7 , the shape surrounded by the bounding box 72 will be discarded, while the shape surrounded by the bounding box 71 will be identified as the region where the object has been detected within the second image 302. Finally, and particularly, the identified region may be stored in the computing system for further processing.
[0047] Thus, using the color profile, it is possible to identify or detect the area of the second image where the object to be detected and located is present. Of course, other techniques can be used to detect and locate the object in the second image. This will be described in more detail below for one of the other techniques shown by Figure 5 and Figure 3C together, and one of the other techniques enables the detection of an object that is missing or newly present somewhere in the physical environment. In particular, one of the other techniques enables the detection of the presence or absence of an object in the second image.
[0048] According to this second preferred embodiment 500, the computing system 100 is configured to receive or acquire a third image 303 and optionally a fourth image at step 231'. The third image 303 is a 2D pixel image of the physical environment 300, but acquired at a different time T and preferably from the same viewpoint as the second image 302. The fourth image is, for example, a point cloud image of the physical environment, acquired at the different time T compared to the first point cloud image. Preferably, the second image and the third image are equirectangular panoramic images. In other words, the computing system 100 is configured to receive pictures of the physical environment, where the pictures are acquired from the same viewpoint but at different times. This enables a temporal comparison of images (representing the same physical environment) acquired at different times in order to identify changes that have occurred in the physical environment, where the changes can represent the new presence or absence of one or several objects in the physical environment.
[0049] At step 232', the computing system is configured to compare the second image 302 with the third image 303 in order to identify one or several regions A1, A2 within the second image 302 or the third image 303, wherein the second image 302 and the third image 303 are different. The comparison thus makes it possible to identify regions that are significantly different between successive acquisitions in pixel images that are temporally continuous in the physical environment, wherein the comparison can be based on computer vision techniques for object detection, such as, for example, those described in the paper by Neelam Dwivedi et al. ("An Approach for Unattended Object Detection through Contour Formation using Background Subtraction", Procedia Computer Science (171): p. 1979-1988 (2020)). Before starting the comparison step 232', the method according to the invention may include additional steps for processing the second image and the third image. Such image processing techniques are shown by Figure 8 illustrated. In this example, the second image and the third image are equirectangular panoramic images, and the computing system 100 is configured to convert the second image and the third image into a cube map to minimize distortion. In Figure 8 , the equirectangular panoramic image of the physical environment corresponds to image 801, which is then converted into a cube map 802. Of course, other image processing techniques may be applied before the comparison step 232' to facilitate the comparison between the first image and the second image.
[0050] At step 233', the computing system is configured to discard identified regions whose size is smaller than a predefined region threshold (e.g., defined according to the number of pixels), and for each identified region whose size is greater than the predefined region threshold, at step 234', the computing system is further configured to identify and / or store the regions in the second image 302 in which the regions that have been identified as corresponding to the missing or newly present objects to be detected and located are present. To identify the regions in the second image 302, the computing system 100 can, for example, surround the relevant regions by a bounding box B1 that can be easily detected by the user in the case where the second image is displayed. For example, the computing system can be configured to display (e.g., side by side) the second image and the third image simultaneously, where each region in which a change has been detected (or identified) is highlighted in both images by a bounding box for easy visualization by the user. If needed, the computing system can be configured to automatically orient the first image and the second image to the same viewing point with respect to the regions in which a change has been detected or identified. Preferably, only the regions corresponding to the newly present objects are identified in the second image 302. Due to this discarding step, the computing system ignores small changes between temporally consecutive images, thus reducing false positive results.
[0051] After the detection step 230, the computing system 100 is configured to, at step 240 and for each region in which the object 310 (e.g., missing object or newly present object) has been detected in the second image, find the corresponding region in the first image 301. To this end, the computing system 100 may be configured to automatically determine the corresponding 3D coordinates in the point cloud image for each pixel of the second image 302. In particular, finding the corresponding region may be achieved by the computing system through the following process: projecting the second image 302 onto a sphere using, for example, an equirectangular projection, where the center of the sphere corresponds to the viewpoint from which the second image 302 has been acquired; then, identifying or characterizing each pixel of the second image 302 by two angles that correspond to the spherical coordinates of the pixel relative to a spherical coordinate system centered at the center of the sphere; and projecting rays from the same viewpoint in the 3D point cloud according to the identified spherical coordinates until they intersect the (tessellated) surface defined or reconstructed from the point cloud. Then, the coordinates of the intersection of the ray with the surface are the corresponding 3D coordinates of the pixel in the point cloud relative to the spherical coordinate system. Additionally or alternatively, the computing system may be configured to automatically match the features and / or objects of the physical environment 300 that are present in both the first image 301 and the second image 302 to determine the position of the pixels of the second image in the first image such that the corresponding regions are found. If needed, the computing system may be configured to automatically match the scale used to represent the physical environment 300 in the first image 301 with the scale used to represent the physical environment 300 in the second image 302 such that the objects present in the physical environment 300 are characterized by the same magnification in the first and second images. Of course, other techniques known in the art may be used to determine the corresponding 3D coordinates in the point cloud image for each pixel of the second image 302.
[0052] Then, at step 250, the computing system 100 is configured to provide, via an interface, the corresponding region in the first image 301 as the location where the object 310 has been detected, and optionally, is configured to extract from the first image 301 the position of the object whose location data is associated with at least one point of the corresponding region in the first image 301. In particular, at least one point of the corresponding region may be the center of the (bounding) box B1, 71 that is used in the second image 302 to identify the region in which the shape of the object or the presence / absence of the object has been recognized, and the bounding box is configured to surround the recognized shape / region in the second image. Preferably, the point cloud of the first image may be presented to the user, for example, via a display, and the point cloud of the first image is oriented according to the same viewpoint as the second image, where the corresponding regions including the detected missing or present objects are surrounded by bounding boxes.
[0053] Advantageously, the detection and localization of the object according to the invention can be used to automatically trigger an action. For example, based on the detection and localization, the computing system 100 according to the invention can automatically update a database with the position information of the detected object, and / or generate an alarm signal if the object that previously existed is absent in the physical environment, or can automatically display a first image and highlight the corresponding area in the first image, so that the user can easily identify in the point cloud content where the object has been found or where the object is missing.
[0054] Advantageously, the first preferred embodiment 400 for detecting an object enables a significant reduction in the number of matching candidates when searching for an object in a physical environment by using the color profile of the object / device to identify relevant regions in a second image. This method allows for the rapid identification of relevant objects, especially in industrial environments where objects can be considered color-coded (i.e., one set of primary colors for robots, another set of primary colors for furniture, etc.), and minimizes the number of false positive matches.
[0055] The second preferred embodiment 500 for detection proposes analyzing pixel images of the physical environment obtained from two scans or pixel images taken at different times of the same scene. Thus, the same scene of the physical environment is observed at two different times, which enables the efficient and rapid detection of missing or newly installed objects. Once a change is detected in such a scene, these changes can be presented to the user together with the corresponding regions or areas in the point cloud of the first image. To detect the changes, computer vision techniques such as color profiles or histograms, shape detection, contour detection, etc. can be used. Preferably, the computing system extracts a set of metrics (e.g., the shape of the object, and / or color profile, and / or size, and / or localization, and / or any other metric that can be obtained from the point cloud data) from the second image, the set of metrics being extracted from the point cloud data (especially the positions of the points), and can be used to identify and localize high-level changes between the point cloud scans within the first image, such as a device changing position or the addition / removal of a device. Using the metrics enables the rapid and efficient extraction of information characterizing the object to be detected / localized from the second image, rather than from the point cloud. Compared with methods based on distance calculations or 3D mesh generation, the advantage is faster object detection. Once a change is identified in one or several regions of the second image 302 or the third image 303, the corresponding "area of interest" can be identified in the point cloud (corresponding to the first image or the fourth image), thus enabling the user to quickly evaluate potential changes in the physical environment when displaying the point cloud and pointing out the corresponding areas. Instead, by focusing on the corresponding areas, this particularly eliminates the need to compare all points in the point cloud.
[0056] Thus, the method according to the present invention provides for the interactive recognition of high-level changes in an industrial environment for decision-making and the use of computer vision techniques to detect changes between scans. Instead of 3D data, pixel panoramic images are preferably used for recognizing changes (especially changes between scans). Using pixel images to detect objects enables a "coarse" analysis for detecting relevant image regions and discarding irrelevant small changes, rather than performing time-consuming precise point cloud analysis when processing point cloud data. Thus, relevant regions in the point cloud can be quickly identified from the relevant regions that can be detected in the pixel image, which enables the very quick identification of (missing / present) objects in the point cloud image while reducing the number of false positives. Considering that there are more than a thousand panoramic images and laser scans covering all areas in a typical automotive factory, the method claimed in the present invention allows for the very quick identification of objects with a small number of false positives compared to other computationally intensive methods. This helps, for example, an operator to detect suspicious changes in areas where no changes should be made, or to ensure that changes such as equipment addition or removal have indeed occurred.
[0057] In an embodiment, the term "receive" as used herein may include obtaining from a storage device, receiving from another device or process, receiving via interaction with a user, or otherwise receiving.
[0058] Those skilled in the art will recognize that, for simplicity and clarity, not all structures and operations of all data processing systems suitable for use with the present disclosure are shown or described herein. Instead, only so much of the data processing systems as are unique to the present disclosure or necessary for understanding the present disclosure are shown and described. The remainder of the construction and operation of data processing system 100 may conform to any of a variety of current implementations and practices known in the art.
[0059] It is important to note that while the present disclosure includes a description in the context of a full-featured system, those skilled in the art will understand that at least portions of the present disclosure are capable of being distributed in the form of instructions in any form of machine-usable, computer-usable, or computer-readable medium, and the present disclosure applies equally regardless of the specific type of instruction or signal-bearing medium or storage medium used for actual execution of the distribution. Examples of machine-usable / readable or computer-usable / readable media include: non-volatile hard-coded type media such as read-only memory (ROM) or electrically erasable programmable read-only memory (EEPROM), and user-recordable type media such as floppy disks, hard disk drives, and compact disc read-only memory (CD-ROM) or digital versatile disc (DVD).
[0060] Although the exemplary embodiments of the present disclosure have been described in detail, those skilled in the art will understand that various changes, substitutions, variations, and improvements disclosed herein can be made without departing from the spirit and scope of the present disclosure in its broadest form.
[0061] No description in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims: the scope of the patent subject matter is defined only by the allowed claims.
Claims
1. A method for detecting and locating an object (310) in a physical environment (300), the method comprises: by a computing system (100): - receiving (210) a first image (301) representing the physical environment (300), wherein the first image is a 3D point cloud image, and the 3D point cloud image includes positioning data for points in the point cloud image; - receiving (220) a second image (302) representing the physical environment (300), wherein the second image (302) is a 2D pixel image of the physical environment (300); - detecting (230) the object (310) in one or several regions in the second image (302); - for each region in the second image (302) where the object (310) has been detected, finding (240) a corresponding region in the first image (301); - providing (250) the corresponding region in the first image (301) via an interface as the location where the object (310) has been detected.
2. The method according to claim 1, wherein, the providing step (250) includes: extracting from the first image (301) the position of the object of the positioning data associated with at least one point in the corresponding region in the first image (301).
3. The method according to claim 1 or 2, wherein, finding (240) the corresponding region includes: automatically orienting the point cloud to the same viewpoint as the viewpoint used to acquire the second image (302).
4. The method according to one of claims 1 or 3, wherein, finding (240) the corresponding region includes: automatically determining the corresponding 3D coordinates in the point cloud image for each pixel of the second image (302).
5. The method according to one of claims 1 to 4, wherein, finding (240) the corresponding region includes: projecting the second image (302) onto a sphere, wherein the center of the sphere corresponds to the viewpoint that has been used to acquire the second image (302); identifying each pixel of the second image (302) by two angles, the two angles corresponding to the spherical coordinates of the pixel relative to a spherical coordinate system centered on the center of the sphere; and projecting rays from the same viewpoint in the 3D point cloud according to the identified spherical coordinates until they intersect the surface defined by or reconstructed from the point cloud.
6. The method according to one of claims 1 to 5, comprises: automatically matching the scale used to represent the physical environment (300) in the first image (301) with the scale used to represent the physical environment (300) in the second image (302).
7. The method according to one of claims 1 to 6, wherein, the detecting (230) step includes: - finding (231) in the second image pixels whose color values fall within the range of color values defined by a clustering of color values representing the main color of the object (310); - Identify (233) one or more shapes formed by one or more groups of the found pixels; - For each identified shape, compare (235) the color histogram of the pixels of the second image (302) that belong to the identified shape with the color histogram of the pixels of the object (310) for which the range of the color values has been determined, and if the comparison results in a color histogram difference higher than a predefined threshold, discard the identified shape, otherwise identify (236) and / or store the area in the second image (302) where the shape corresponding to the object (310) to be detected and located has been identified.
8. The method according to claim 7, wherein, at least one point of the corresponding area corresponds to the center of the bounding box for identifying (236) the area of the shape that has been identified in the second image (302), and the bounding box is configured to surround the identified shape in the second image (302).
9. The method according to any one of claims 1 to 6, wherein, the detecting (230) step is configured to detect an object that is missing or newly present at a certain place in the physical environment (300), and the detecting (230) step includes: - Receiving (231') a third image (303), which is a 2D pixel image of the physical environment (300) acquired according to the same viewing point as the second image (302) but at a different time T; - Comparing (232') the second image (302) with the third image (303) to identify one or more areas in which the second image (302) and the third image (303) are different; - Discarding (233') areas smaller than a predefined area threshold, and for each area larger than the predefined area threshold, identifying (234') and / or storing the area in the second image (302) where the area corresponding to the missing or newly present object to be detected and located has been identified.
10. A computing system (100), comprising: a processor; and an accessible memory, and the computing system (100) is configured to: - Obtain or receive (210) a first image (301) representing the physical environment (300), wherein the first image is a 3D point cloud image, and the 3D point cloud image includes positioning data for the points in the point cloud image; - Obtain or receive (220) a second image (302) representing the physical environment (300), wherein the second image (302) is a 2D pixel image of the physical environment (300); - Detect (230) the object (310) in one or more areas in the second image (302); - For each area in the second image (302) where the object (310) has been detected, find (240) the corresponding area in the first image (301); and - Provide (250) the corresponding region in the first image (301) via the interface as the location where the object (310) has been detected.
11. The computing system (100) according to claim 11, wherein, To detect (230) the object (310) in one or several regions of the second image (302), the computing system is configured to: - Find (231) pixels in the second image whose color values fall within the range of color values defined by a clustering of color values representing the main color of the object (310); - Identify (233) one or several shapes formed by one or several clusters of the found pixels; - For each identified shape, compare (235) the color histogram of the pixels in the second image (302) belonging to the identified shape with the color histogram of the pixels of the object (310) for which the range of color values has been determined, and if the comparison yields a color histogram difference higher than a predefined threshold, discard the identified shape, otherwise identify (236) and / or store the region in the second image (302) where the shape corresponding to the object (310) to be detected and located has been identified.
12. The computing system (100) according to claim 11, wherein, To detect (230) the object (310) in one or several regions of the second image (302), the computing system is configured to: - Receive (231') a third image (303), which is a 2D pixel image of the physical environment (300) acquired at a different time T from the same viewpoint as the second image (302); - Compare (232') the second image (302) with the third image (303) to identify one or several regions where the second image (302) and the third image (303) are different; - Discard (233') regions smaller than a predefined region threshold, and for each region larger than the predefined region threshold, identify (234') and / or store the region in the second image (302) where the region corresponding to a missing or newly present object to be detected and located has been identified.
13. A non-transitory computer-readable medium encoded with executable instructions that, when executed, cause one or more data processing systems to perform the following operations: - Obtain or receive (210) a first image (301) representing the physical environment (300), wherein, The first image is a 3D point cloud image, and the 3D point cloud image includes location data for the points in the point cloud image; - Obtain or receive (220) a second image (302) representing the physical environment (300), wherein the second image (302) is a 2D pixel image of the physical environment (300); - Detect (230) the object (310) in one or several regions of the second image (302); - For each region in the second image (302) in which the object (310) has been detected, find (240) the corresponding region in the first image (301); and - Provide (250) the corresponding region in the first image (301) via an interface as the location where the object (310) has been detected.
14. The non-transitory computer-readable medium according to claim 13, wherein, To detect (230) the object (310) in one or more regions in the second image (302), when the executable instructions are executed, cause the one or more data processing systems to perform the following operations: - In the second image, find (231) pixels whose color values fall within a range of color values defined by a clustering of color values representing the main color of the object (310); - Identify (233) one or more shapes formed by one or more groups or clusters of the found pixels; - For each identified shape, compare (235) the color histogram of the pixels in the second image (302) belonging to the identified shape with the color histogram of the pixels of the object (310) for which the range of color values has been determined, and if the comparison results in a color histogram difference higher than a predefined threshold, discard the identified shape, otherwise identify (236) and / or store the region in the second image (302) in which the shape corresponding to the object (310) to be detected and located has been identified.
15. The non-transitory computer-readable medium according to claim 13, wherein, To detect (230) the object (310) in one or more regions in the second image (302), when the executable instructions are executed, cause the one or more data processing systems to perform the following operations: - Receive (231') a third image (303), which is a 2D pixel image of the physical environment (300) acquired at a different time T from the same viewpoint as the second image (302); - Compare (232') the second image (302) with the third image (303) to identify one or more regions in which the second image (302) and the third image (303) are different; - Discard (233') regions smaller than a predefined region threshold, and for each region larger than the predefined region threshold, identify (234') and / or store the region in the second image (302) in which the region corresponding to a missing or newly present object to be detected and located has been identified, for example by enclosing the relevant region with a bounding box.