Photometric stereoscopic fly-shooting with image alignment
By performing a calibration phase on the calibrated object, calculating the object's direction, velocity, and camera distance, generating a 3D offset vector, and aligning the image, the problem of image alignment in aerial photography scenarios is solved, achieving high-precision 3D calculation and defect detection.
Patent Information
- Application Number
- CN202510870490.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-31
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-14
AI Technical Summary
In aerial photography scenarios, it is difficult to align object images taken from different angles and under different lighting conditions to achieve high-precision 3D calculations and defect detection.
By using a calibration object in the calibration phase, the object's direction, velocity, and camera distance are calculated to generate a 3D offset vector. This vector is then applied to align the image during the photometric stereo vision process. Combined with image calibration functions and perspective distortion processing, image alignment is achieved.
It enables high-precision three-dimensional calculation and defect detection of object surfaces under aerial photography conditions, improving the accuracy and consistency of detection.
Smart Images

Figure CN120948480A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the detection of camera components. More specifically, this application relates to an illumination controller with noise and signal isolation functions for detecting camera components. Background Technology
[0002] Inspection cameras are used in industrial products to help detect defects in finished products. For example, if a manufacturer is producing metal castings, one or more inspection cameras may be placed on the manufacturing and / or assembly line to inspect the produced metal castings or portions thereof to detect any quality control issues. Inspection camera assemblies may include cameras mounted on or near multiple independently controlled light sources. These light sources may be activated in a coordinated sequence controlled by a lighting controller to illuminate the finished product from different angles and at different times. Summary of the Invention
[0003] On the one hand, this application provides a system comprising:
[0004] Lighting fixture, including multiple lamps;
[0005] A calibration object with shape and size;
[0006] A camera aimed at the conveyor belt;
[0007] A computer system includes at least one hardware processor and a non-transitory computer-readable medium, the medium storing instructions that, when executed by the at least one hardware processor, perform the following operations:
[0008] Adjust the illumination from the lighting device and acquire multiple images of the calibration object from the camera under different lighting conditions;
[0009] The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images;
[0010] The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector;
[0011] The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object;
[0012] The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object;
[0013] The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
[0014] On the other hand, this application provides a method comprising, at the controller:
[0015] Make the calibration object pass under the camera;
[0016] As the calibration object passes under the camera, the illumination from the lighting device is adjusted, and multiple images of the calibration object are acquired from the camera under different lighting conditions;
[0017] The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images;
[0018] The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector;
[0019] The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object;
[0020] The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object;
[0021] The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
[0022] Thirdly, this application provides a non-transitory machine-readable storage medium having instructions embodied thereon that are executable by one or more machines to perform operations on a controller, the operations including:
[0023] Make the calibration object pass under the camera;
[0024] As the calibration object passes under the camera, the illumination from the lighting device is adjusted, and multiple images of the calibration object are acquired from the camera under different lighting conditions;
[0025] The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images;
[0026] The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector;
[0027] The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object;
[0028] The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object;
[0029] The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
[0030] As can be seen from the technical solution provided in this application, the orientation and velocity of the calibration object, as well as the distance between the calibration object and the camera, can be calculated. Based on these three pieces of information, a three-dimensional offset vector is created. Then, this three-dimensional offset vector can be applied during photometric stereo processing to align objects in different images, so as to accurately calculate the orientation and scale of the surface. Attached Figure Description
[0031] Figure 1 A block diagram of a detection system based on some examples is shown.
[0032] Figure 2 This is a diagram illustrating an example of a calibration object according to an example embodiment.
[0033] Figures 3A-3H These are example images taken when calibrating the detection system using calibration object 200, according to an example embodiment.
[0034] Figure 4 This is a flowchart illustrating a method for operating the controller according to an example embodiment.
[0035] Figure 5 This is a block diagram illustrating a mobile device according to an example embodiment.
[0036] Figure 6 It is a block diagram of a machine in the form of a computer system example, within which instructions can be executed to cause the machine to perform any or more of the methods discussed herein. Detailed Implementation
[0037] In inspection environments, particularly industrial or quality control environments, aerial photography refers to a technique used to detect and measure defects, flaws, or problems on the surface of an object without it being directly in front of the camera. In other words, the object can be moving during the inspection, such as moving on a conveyor belt.
[0038] This technology is useful in environments requiring high-speed detection. It can quickly capture and analyze data from fast-moving objects or processes, making it suitable for automated detection systems.
[0039] Photogrammetry typically uses advanced imaging or scanning techniques to provide precise measurements and detailed images of surfaces. This high level of detail helps identify small or subtle defects that may affect product quality or performance.
[0040] To generate an accurate 3D computer representation of an object's surface, it's typically necessary to capture multiple distinct images of the object from different angles and under varying lighting conditions. However, achieving high accuracy is challenging when attempting this in aerial photography scenarios because aligning multiple images in a way that the system knows the position of each part of the object in each image is difficult. For example, if there's a small bump on the object, it might appear in different positions in each image because the object is moving, and also because the relative position and orientation of the conveyor belt and camera can cause the viewing angle to differ between images. In other words, it's not simply a matter of finding a specific part of the object in one image and aligning it with the same specific part in another image; because the two images may have different viewing angles, even if a specific part of the object is aligned, other parts of the same object may end up misaligned.
[0041] In one example embodiment, a calibration phase is performed where the calibration object is transmitted through the system and multiple images of the calibration object are captured. From these images, the orientation and velocity of the calibration object, as well as the distance between the calibration object and the camera, can be calculated. Based on these three pieces of information, a three-dimensional offset vector is created. This three-dimensional offset vector can then be applied during the photometric stereo vision process to align objects in different images, enabling accurate calculation of surface orientation and scale.
[0042] Figure 1 A block diagram of an inspection system 100 according to some examples is shown. The inspection system 100 includes a photomask 102, a camera 108, a controller 106, an industrial computer 112, and a factory computer 116. The factory computer 116 communicates with the computer 112 via a wired or wireless factory network 124.
[0043] The photomask 102 is used to illuminate a target object 104, such as a metal casting or other product. The photomask 102 includes a housing containing multiple light sources, as will be described in more detail below. In some examples, the light sources include multiple light-emitting diodes (LEDs) or displays arranged to provide flexibility in illuminating the target object 104. The light sources are selectively activated by a controller 106 via a power line 110. A light source is a unit of illumination that can be individually addressed by the controller 106 to illuminate the target object 104. Therefore, a single light source may include a single LED or multiple LEDs that can be addressed as a group. A light source may also include a subset of light-emitting units, such as a group or block of pixels in a flexible display. Preferably, the photomask 102 includes at least ten individually addressable light sources arranged within the photomask 102 to provide illumination flexibility.
[0044] Camera 108 can be mounted on photomask 102 via bracket 114. It captures images of the illuminated target object 104 through a hole in the top of photomask 102. Camera 108 is triggered by controller 106 via trigger line 118, synchronized with the activation of the light source in photomask 102.
[0045] Controller 106 controls the operation of camera 108 and the illumination of target object 104 by photomask 102. Controller 106 receives instructions from computer 112 via control line 122. Controller 106 may be implemented by a hardware processor disposed in camera 108. Controller 106 may also include hardware components, which may include a central processing unit (“CPU”), bus, volatile and non-volatile storage devices, storage cells, non-transitory computer-readable media, data processor, processing device, control device, transmitter, receiver, antenna, transceiver, input device, output device, network interface device, and other types of components well known to those skilled in the art. These hardware components within the user equipment can be used independently of other devices disclosed herein to perform the various applications, methods, or algorithms disclosed herein.
[0046] The controller 106 illuminates the target object according to one or more optimal lighting configurations. The lighting configuration can be defined as a matrix, where each value represents the operating state of an individual controllable light source, such as one or more light-emitting diodes (LEDs) and / or pixel groups on a flexible display screen. The matrix may also include brightness or color values for a specific configuration. The lighting configurations can also be arranged into a sequence specifying the order in which lighting configurations are to be executed for a particular target object 104, so that the camera 108 can capture multiple images under different lighting conditions.
[0047] Computer 112 runs software that provides a user interface for specifying lighting configurations and sequences, and can be loaded into controller 106. Computer 112 also instructs controller 106 to operate via control line 122 and receives images captured by camera 108 via data line 120.
[0048] Factory computer 116 provides overall control of the factory and can receive operational data and captured images from computer 112 via factory network 124. Factory computer 116 can also provide instructions to control or initiate the operation of detection system 100 based on, for example, other factory operations (such as the movement of target object 104 over photomask 102).
[0049] An object can be placed on conveyor belt 126, and conveyor belt 126 can be moved, thereby moving the object so that when one or more light sources on photomask 102 are illuminated, the object is at least partially below camera 108. As previously described, this can be performed under fly-shooting conditions, where conveyor belt 126 is not moving, so the object does not stop below camera 108. Instead, multiple images of the object are captured from different angles and under different lighting conditions, but instead of moving around the object to capture these different angles, camera 108 moves while camera 108 remains stationary.
[0050] As previously stated, in order to achieve image alignment when performing defect detection on multiple images of an actual component, a calibration operation must first be performed. During calibration, the calibration object passes through the detection system via conveyor belt 126. In one example embodiment, the calibration object is a milky white glass chessboard with a central pattern different from the chessboard. It should be noted that a chessboard-based calibration object is merely one possible type of calibration object that can be used, and nothing in this disclosure should be construed as limiting the scope of protection solely to chessboard-based calibration objects, unless otherwise expressly stated. Figure 2 This is a diagram illustrating an example of calibration object 200 according to an exemplary embodiment. As previously described, calibration object 200 is a milky white glass checkerboard. Because it is made of glass, it has a semi-reflective surface for reflecting light from one or more light sources.
[0051] As can be seen, the calibration object 200 is essentially a checkerboard pattern 202, although a unique marker pattern 204 is placed at one point on the calibration object 200. In some example embodiments, this unique marker pattern 204 is placed at the center of the calibration object 200, but this is not strictly necessary.
[0052] The unique marking pattern 204 can be any pattern that is distinguishable from the checkerboard pattern 202. For example, here the unique marking pattern 204 is a series of horizontal and vertical lines that contrast with the diagonals of the checkerboard pattern 202.
[0053] During calibration, the calibration object 200 passes through the detection system, and as the calibration object 200 moves beneath the camera (i.e., during a fly-shot), the camera captures images of the calibration object 200 at different times. During these captures, the light source also alternates so that images of the calibration object 200 are captured not only from different angles but also under different lighting conditions.
[0054] Figures 3A-3H These are example images taken when calibrating the detection system using calibration object 200 according to an example embodiment. Here, each image represents a different image taken at different times and under different lighting conditions, but all are images of the same calibration object 200 moving below the camera.
[0055] Once these images are captured, the second derivative of the image intensity is calculated in each image, and possible corner locations are identified as peaks on the surface based on the second derivative. More specifically, in an image, the first derivative measures the rate of change of pixel intensity. The second derivative measures the rate of change of the first derivative and is useful for edge detection because zero-crossing points (points where the sign of the second derivative changes) typically correspond to edges. In one example embodiment, the second derivative can be approximated using the Laplacian operator.
[0056] Find corner locations grouped with consistent spacing (e.g., 3x3 groups), and then iteratively expand the chessboard in all directions by searching for identified corners within a small window, extending the existing identified corners. This iterative process is repeated until the image edge is reached or no peak satisfying a threshold is found in the search window.
[0057] Once the corner points of the calibration object are determined, its distance and orientation can be determined. This can be done using a camera calibration function, which takes the detected corner point locations from multiple images and returns parameters describing lens distortion and a camera matrix describing the mapping between 3D and 2D points. It also returns a rotation vector describing the calibration object relative to the camera orientation. The camera calibration function operates by solving a system of equations relating the 3D coordinates of points in the calibration object to their 2D image coordinates. The function can also employ optimization techniques to refine the estimation of camera parameters and attempt to minimize reprojection error, the difference between observed image points and projected 3D points.
[0058] Then trigonometry can be used, taking advantage of the angle information about the lens provided by the camera calibration function and the known actual size of the calibration object, to solve for the distance from the calibration object to the camera in each image.
[0059] At this point, a plane can be fitted to the position / orientation of all observed images of the calibration object. Since the calibration object moves within the same plane, the orientation vector of the calibration object can be averaged over all images. Similarly, the shortest distance from the camera to the calibration object plane can be calculated in each image, and then the distances over all images can be averaged.
[0060] Furthermore, the motion observed in the images can be used to calculate motion vectors within the transport plane. More specifically, now that the orientation of the transport plane is determined, perspective distortion of the images can be used to simulate a top view of the transporter. For each pair of consecutive images, this top view can be generated, edges in each image can be detected, and then correlations within the offset range can be used to find motion vectors that best align pixels with the top view.
[0061] The perspective distortion can then be reversed, and the endpoints of the vector can be projected onto the transport plane to provide a three-dimensional motion vector. This motion vector can then be averaged over all consecutive image pairs to obtain the final estimated direction of the motion vector. The timestamp difference between each consecutive image pair and the motion vector found between these images can be used to estimate the transport speed, and the speed estimates for all image pairs are averaged to obtain the average transport speed.
[0062] Therefore, all images are sorted from earliest to latest by timestamp. For each pair of consecutive images, the pixels in the earlier image are projected onto the three-dimensional coordinates of the plane on the conveyor belt using lens and conveyor calibration. These three-dimensional coordinates are then moved a distance determined by the calibration speed and image timestamps along the calibrated motion vector, and then the three-dimensional coordinates are reprojected back into the two-dimensional image.
[0063] In some example implementations, after determining the time-based motion, the alignment is further improved by calculating edge surfaces for each image using Gaussian blur and image derivatives. For each pair of consecutive images, the correlation of the edge surfaces is calculated over all offsets within a small pixel radius. The location of the highest correlation represents the two-dimensional offset vector used for fine-tuning the alignment. This two-dimensional offset vector is then projected onto a calibrated conveyor belt plane, and then onto a calibrated motion vector. The length of this projected vector and the calibrated conveyor speed then provide a fine-tuned time offset for the first image of the image pair. This process is performed for each pair of images, adjusting the timestamps to the fine-tuned values. The time-based motion process can then be repeated using the fine-tuned timestamps to obtain the finally aligned images.
[0064] In some example implementations, the methods described above run on a graphics processing unit (GPU) instead of a central processing unit (CPU). Running on a GPU instead of a CPU requires making several algorithmic decisions. Specifically, as mentioned earlier, Gaussian blur is used. This is done instead of median filtering because, although median filtering is more robust to lighting changes, it cannot be implemented on a GPU. Furthermore, correlation is used instead of normalized cross-correlation. Again, this sacrifices some quality on challenging images for the sake of speed.
[0065] Figure 4 This is a flowchart illustrating a method for aligning component images captured by a drone mechanism, according to an example embodiment.
[0066] In operation 410, the conveyor belt is moved so that the calibration object passes under the camera. In operation 420, as the calibration object passes under the camera, the illumination from the lighting equipment is adjusted, and multiple images of the calibration object are captured from the camera under different lighting conditions.
[0067] In operation 430, the multiple images are organized chronologically based on the timestamps associated with each image. In operation 440, the multiple images are input into a camera calibration function to obtain a camera matrix, distortion coefficients, and a rotation vector. The camera matrix defines the mapping from 3D points to 2D points. The distortion coefficients describe the degree of distortion in each of the multiple images. The rotation vector describes the amount of rotation of the calibration object in each of the multiple images. The distortion coefficients describe the lens distortion in each of the multiple images, such that knowing the distortion coefficients allows mapping the coordinates of pixels in the images to vectors in 3D space.
[0068] In operation 450, the distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape within the calibration object. In operation 460, the velocity and orientation of the calibration object are calculated using the camera matrix, distortion coefficients, rotation vector, and multiple images. In operation 470, the calculated distance, velocity, and orientation are used to modify the multiple images so that the calibration object is aligned in a single plane across all images.
[0069] In operation 475, the part with the defect to be analyzed is placed on a conveyor belt. In operation 480, the conveyor belt is moved so that the part passes under the camera. In operation 485, while the part passes under the camera, the illumination from the lighting equipment is adjusted, and a second set of multiple images is captured from the camera under different lighting conditions. In operation 490, the images in the second set of multiple images are aligned using a three-dimensional offset vector.
[0070] Figure 5 This is a block diagram 500 illustrating a software architecture 502 that can be installed on any one or more of the aforementioned devices. Figure 5 This is merely a non-limiting example of a software architecture, and it will be understood that many other architectures can be implemented to facilitate the functionality described herein. In various embodiments, software architecture 502 is derived from, for example, references to... Figure 6 The hardware implementation of machine 600 includes a processor 610, memory 630, and input / output (I / O) components 650. In this example architecture, software architecture 502 can be conceptualized as a layered stack, where each layer provides specific functionality. For example, software architecture 502 includes layers such as operating system 504, libraries 506, frameworks 508, and applications 510. Operationally, application 510 invokes application programming interface (API) calls 512 through the software stack and receives messages 514 in response to API calls 512, consistent with some embodiments.
[0071] In various implementations, the operating system 504 manages hardware resources and provides general services. For example, the operating system 504 includes a kernel 520, services 522, and drivers 524. According to some embodiments, the kernel 520 acts as an abstraction layer between the hardware and other software layers. For example, the kernel 520 provides functions such as memory management, processor management (e.g., scheduling), component management, network and security settings. Services 522 can provide other general services to other software layers. Drivers 524 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 524 may include display drivers, camera drivers, Bluetooth or Bluetooth Low Energy drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi drivers, audio drivers, power management drivers, and so on.
[0072] In some embodiments, library 506 provides the low-level general-purpose infrastructure used by application 510. Library 506 may include system libraries 530 (e.g., the C standard library) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Furthermore, library 506 may include API libraries 532, such as media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as MPEG-4, H.264 or AVC, MP3, AAC, AMR audio codecs, JPEG or JPG, PNG), graphics libraries (e.g., the OpenGL framework for two-dimensional (2D) and three-dimensional (3D) rendering in a graphics context on a display), database libraries (e.g., SQLite to provide various relational database functions), network libraries (e.g., WebKit to provide web browsing functionality), and so on. Library 506 may also include a variety of other libraries 534 to provide many other APIs to application 510.
[0073] Framework 508 provides advanced general-purpose infrastructure that application 510 can use. For example, framework 508 provides various graphical user interface functions, advanced resource management, advanced location services, and so on. Framework 508 can provide a wide range of other APIs that application 510 can use, some of which may be specific to a particular operating system 504 or platform.
[0074] In one example embodiment, application 510 includes a homepage application 550, a contacts application 552, a browser application 554, an e-book reader application 556, a location application 558, a media application 560, a messaging application 562, a game application 564, and a variety of other applications, such as a third-party application 566. Application 510 is a program that performs the functions defined in the program. One or more applications 510 can be created using various programming languages and in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, third-party application 566 (e.g., used by an entity other than a specific platform vendor using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can run on platforms such as iOS. TM ANDROID TM , Mobile software on a phone or other mobile operating system.
[0075] Figure 6 A schematic representation of a machine 600 in the form of a computer system is shown, within which a set of instructions can be executed to cause the machine 600 to perform any or more methods discussed herein. Specifically, Figure 6 A schematic representation of machine 600 is shown as an example of a computer system, wherein instructions 616 (e.g., software, program, application, applet, application program, or other executable code) cause machine 600 to perform any or more of the methods discussed herein. For example, instruction 616 may cause machine 600 to execute... Figure 4 Method 400. Alternatively, or alternatively, instruction 616 can implement... Figure 1-4And so on. Instruction 616 transforms a general, unprogrammed machine 600 into a specific machine 600 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, machine 600 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, machine 600 may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 600 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablets, laptops, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, network devices, network routers, network switches, bridges, or any machine capable of executing instruction 616, which sequentially or otherwise specifies the actions to be taken by machine 600. Furthermore, although only a single machine 600 is shown, the term "machine" should also be understood to include a collection of machines 600 that individually or collectively execute instructions 616 to perform any or more methods discussed herein.
[0076] Machine 600 may include processor 610, memory 630, and input / output components 650, which may be configured to communicate with each other, for example, via bus 602. In example embodiments, processor 610 (e.g., CPU, Reduced Instruction Set Computing (RISC) processor, Complex Instruction Set Computing (CISC) processor, Graphics Processing Unit (GPU), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Radio Frequency Integrated Circuit (RFIC), another processor, or any suitable combination) may include, for example, processors 612 and 614 capable of executing instructions 616. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that can execute instructions 616 simultaneously. Although Figure 6 Multiple processors 610 are shown, but machine 600 may include a single processor 612 with a single core, a single processor 612 with multiple cores (e.g., a multi-core processor 612), multiple processors 612, 614 with single cores, multiple processors 612, 614 with multiple cores, or any combination thereof.
[0077] Memory 630 may include main memory 632, static memory 634, and memory cell 636, which are accessible by processor 610 via bus 602. Main memory 632, static memory 634, and memory cell 636 store instructions 616 embodying any one or more methods or functions described herein. During execution of instructions 616 by machine 600, instructions 616 may also reside wholly or partially in main memory 632, static memory 634, memory cell 636, at least one processor 610 (e.g., in the processor's cache memory), or any suitable combination thereof.
[0078] Input / output component 650 may include a wide variety of components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement data, and so on. The specific input / output component 650 included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, while displayless server machines may not include such touch input devices. It should be understood that input / output component 650 may include many components not included in the above description. Figure 6 Other components are shown in the diagram. The input / output components 650 are grouped by function only for the purpose of simplifying the following discussion, and this grouping is in no way limiting. In various example embodiments, the input / output components 650 may include output components 652 and input components 654. Output component 652 may include visual components (such as displays like plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRTs), acoustic components (such as speakers), haptic components (such as vibration motors, resistive mechanisms), other signal generators, and so on. Input component 654 may include alphanumeric input components (such as keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (such as mice, touchpads, trackballs, joysticks, motion sensors, or other indicating instruments), haptic input components (such as physical buttons, touchscreens or other haptic input components that provide the position and / or force of a touch or touch gesture), audio input components (such as microphones), and so on.
[0079] In a further example embodiment, input / output component 650 may include biometric component 656, motion component 658, environmental component 660 or location component 662, and numerous other components. For example, biometric component 656 may include expression (e.g., hand expressions, facial expressions, voice expressions, body posture, or eye tracking) detection components, biosignal (e.g., blood pressure, heart rate, body temperature, sweating, or brainwave) measurement components, identity recognition (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition) components, and so on. Motion component 658 may include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), and so on. Environmental component 660 may include, for example, a light sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for safety detection of hazardous gas concentrations or measurement of pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Position component 662 may include a position sensor component (e.g., a Global Positioning System (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer from which altitude can be derived), a direction sensor component (e.g., a magnetometer), and so on.
[0080] Communication can be implemented using a wide variety of technologies. Input / output component 650 may include communication component 664, operable to couple machine 600 to network 680 or device 670 via coupling 682 and coupling 672, respectively. For example, communication component 664 may include a network interface component or other suitable device interfaced with network 680. In further examples, communication component 664 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth components (e.g., Bluetooth Low Energy), Wi-Fi components, and other communication components for providing communication in other ways. Device 670 may be another machine or any of a wide variety of peripheral devices (e.g., coupled via USB).
[0081] Furthermore, the communication component 664 can detect identifiers or include components operable to detect identifiers. For example, the communication component 664 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes (such as Universal Product Code (UPC) barcodes), multi-dimensional barcodes (such as QR codes, Aztec codes, data matrix codes, data map codes, maximum codes, PDF417, super codes, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various types of information can be obtained through the communication component 664, such as location geolocation via Internet Protocol (IP), location triangulation via Wi-Fi signals, location by detecting NFC beacon signals that may indicate a specific location, and so on.
[0082] Various memories (i.e., the memories of 630, 632, 634 and / or the memory of processor 610) and / or storage units 636 may store one or more sets of instructions 616 and data structures (e.g., software) embodying or utilized by any one or more methods or functions described herein. When executed by processor 610, these instructions (e.g., instructions 616) cause various operations to be performed to implement the disclosed embodiments.
[0083] As used herein, the terms “machine-readable medium” 638, “machine storage medium,” “device storage medium,” and “computer storage medium” have the same meaning and are used interchangeably. These terms refer to one or more storage devices and / or media (e.g., centralized or distributed databases and associated caches and servers) that store executable instructions and / or data. Therefore, these terms should be understood to include, but are not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and / or device storage media include non-volatile memory, such as semiconductor storage devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), field-programmable gate arrays (FPGAs), and flash memory devices; disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine storage medium,” “computer storage medium,” and “device storage medium” specifically exclude carrier waves, modulated data signals, and other such media, at least some of which fall within the scope of the term “signal media” discussed below.
[0084] In various example embodiments, one or more portions of network 680 may be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless local area network (WLAN), wide area network (WAN), wireless wide area network (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a Common Old-Style Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi network, another type of network, or a combination of two or more such networks. For example, network 680 or a portion of network 680 may include a wireless or cellular network, and coupling 682 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, Coupler 682 can implement any of a variety of data transmission technologies, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimization (EVDO), General Packet Radio Service (GPRS), GSM Evolution Enhanced Data Rate (EDGE), 3rd Generation Partnership Project (3GPP) including 8G, 4th Generation Wireless Network (4G), Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standard, other standards defined by various standards-setting organizations, other telematics protocols, or other data transmission technologies.
[0085] Instruction 616 can be transmitted or received on network 680 using a transmission medium via a network interface device (e.g., a network interface component included in communication component 664) and utilizing many well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instruction 616 can be transmitted to or received from device 670 using a transmission medium via coupling 672 (e.g., peer-to-peer coupling). The terms “transmission medium” and “signal medium” have the same meaning and are used interchangeably in this disclosure. The terms “transmission medium” and “signal medium” should be understood to include any intangible medium capable of storing, encoding, or carrying instructions 616 for execution by machine 600, and include digital or analog communication signals or other intangible media to facilitate communication of such software. Therefore, the terms “transmission medium” and “signal medium” should be understood to include any form of modulated data signal, carrier wave, etc. The term “modulated data signal” refers to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal.
[0086] The terms “machine-readable medium,” “computer-readable medium,” and “device-readable medium” have the same meaning and are used interchangeably in this disclosure. These terms are defined to include both machine storage media and transmission media. Therefore, these terms include storage devices / media as well as carrier / modulated data signals.
Claims
1. A system comprising: Lighting fixture, including multiple lamps; A calibration object with shape and size; A camera aimed at the conveyor belt; A computer system includes at least one hardware processor and a non-transitory computer-readable medium, the medium storing instructions that, when executed by the at least one hardware processor, perform the following operations: Adjust the illumination from the lighting device and acquire multiple images of the calibration object from the camera under different lighting conditions; The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images; The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector; The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object; The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object; The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
2. The system according to claim 1, characterized in that, The operation further includes: for each of the multiple images: Calculate the second derivative of the image intensity of the corresponding image; The corner points of the calibrated object in the corresponding image are identified based on the second derivative.
3. The system according to claim 2, characterized in that, The shape is a square with a checkerboard pattern.
4. The system according to claim 3, characterized in that, Corner identification involves finding groups of corner locations with consistent spacing, and then iteratively expanding the calibrated object in all directions by searching for corner locations within a small window that continues searching for corner locations until the peak in the search window no longer meets a threshold or an edge of the corresponding image is detected.
5. The system according to claim 1, characterized in that, The operation further includes: for each pair of consecutive images in the plurality of images: Perspective distortion is performed on the first image in each pair of consecutive images based on the rotation vector to simulate a top view of the calibrated object.
6. The system according to claim 1, characterized in that, The operation also includes creating a three-dimensional offset vector using the aligned image.
7. The system according to claim 6, characterized in that, The operation also includes: The component to be analyzed is passed beneath the camera; As the component passes beneath the camera, it adjusts the illumination from the lighting device and acquires a second set of multiple images from the camera under different lighting conditions. The images in the second group of multiple images are aligned using the three-dimensional offset vector.
8. The system according to claim 1, characterized in that, The camera matrix defines the mapping from three-dimensional points to two-dimensional points.
9. The system according to claim 1, characterized in that, The distortion coefficients describe the degree of distortion in each of the plurality of images.
10. The system according to claim 1, characterized in that, The rotation vector describes the amount of rotation of the calibrated object in each of the multiple images.
11. A method comprising, at a controller: Make the calibration object pass under the camera; As the calibration object passes under the camera, the illumination from the lighting device is adjusted, and multiple images of the calibration object are acquired from the camera under different lighting conditions; The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images; The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector; The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object; The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object; The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
12. The method of claim 11, further comprising: For each of the plurality of images: Calculate the second derivative of the image intensity of the corresponding image; The corner points of calibrated objects in the corresponding image are identified based on the second derivative.
13. The method of claim 12, characterized in that, The shape is a square with a checkerboard pattern.
14. The method of claim 13, characterized in that, Corner identification involves finding groups of corner locations with consistent spacing, and then iteratively expanding the calibrated object in all directions by searching for corner locations within a small window that continues searching for corner locations until the peak in the search window no longer meets a threshold or an edge of the corresponding image is detected.
15. The method of claim 11, further comprising: For each pair of consecutive images in a plurality of images: Perspective distortion is performed on the first image in each pair of consecutive images based on the rotation vector to simulate a top view of the calibrated object.
16. The method of claim 11, further comprising creating a three-dimensional offset vector using the aligned image.
17. The method of claim 16, further comprising: The component to be analyzed is passed beneath the camera; As the component passes beneath the camera, it adjusts the illumination from the lighting device and acquires a second set of multiple images from the camera under different lighting conditions. The images in the second group of multiple images are aligned using a three-dimensional offset vector.
18. A non-transitory machine-readable storage medium having instructions embodied thereon that are executable by one or more machines to perform operations on a controller, the operations including: Make the calibration object pass under the camera; As the calibration object passes under the camera, the illumination from the lighting device is adjusted, and multiple images of the calibration object are acquired from the camera under different lighting conditions; The multiple images are organized in chronological order based on the timestamp associated with each of the multiple images; The multiple images are input into a camera calibration function to obtain the camera matrix, distortion coefficients, and rotation vector; The distance between the calibration object and the camera is calculated using the rotation vector and the dimensions of the shape in the calibration object; The camera matrix, distortion coefficients, rotation vectors, and multiple images are used to calculate the velocity and orientation of the calibrated object; The calculated directions and velocities are used to modify the multiple images so that the calibrated objects in all images are aligned on a single plane.
19. The non-transitory machine-readable storage medium of claim 18, characterized in that, The operation also includes: For each of the plurality of images: Calculate the second derivative of the image intensity of the corresponding image; The corner points of calibrated objects in the corresponding image are identified based on the second derivative.
20. The non-transitory machine-readable storage medium of claim 19, characterized in that, The shape is a square with a checkerboard pattern.