Image acquisition method and system based on space-time coding
By adopting spatiotemporal encoding technology in image acquisition, using time encoding rules and spatial encoding rules, the problem of low image quality caused by ambient light and indirect light interference is solved, and higher quality image acquisition and output are achieved.
Patent Information
- Application Number
- CN202510424345.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art cannot effectively distinguish and eliminate ambient light and indirect light, resulting in low image output quality and lack of universality for the imaging environment and the surface state of the object.
An image acquisition method based on space-time encoding is adopted, and an image with encoding information is collected through preset time encoding rules and spatial encoding rules, and the target image is output through the decoding process. The time encoding rules eliminate ambient light by controlling the time series; the spatial encoding rules separate direct light and indirect light by controlling the light source lighting area.
It improves image output quality, enhances the ability to distinguish ambient and indirect light, achieves higher imaging flexibility and accuracy, and is suitable for a variety of imaging environments.
Smart Images

Figure CN120151670A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine vision technology, and particularly to an image acquisition method and system based on spatio-temporal coding. Background Art
[0002] An image sensor converts an optical signal into an electrical signal based on the photoelectric effect and outputs a digital image through a processing circuit. However, the image sensor cannot distinguish the source of photons. For example, light directly reflected from the object surface, indirectly reflected light, or ambient light, resulting in the output image quality being significantly affected by the imaging environment, the state and shape of the object surface.
[0003] To solve the interference of ambient light, the signal-to-noise ratio can be enhanced by increasing the power of the active light source or using a filter to suppress the influence of ambient light; the imaging conditions can also be optimized by adjusting the layout angle between the light source and the camera.
[0004] However, increasing the light source power or using a filter can only partially suppress ambient light, and cannot fundamentally distinguish and eliminate the ambient light component. Moreover, it relies on the layout optimization of the light source and the camera, requires repeated adjustment for different objects, lacks universality, and cannot solve the interference of indirect light (such as multiple reflections or diffracted light on the object surface) on imaging, resulting in low output quality of the image. Summary of the Invention
[0005] This application provides an image acquisition method and system based on spatio-temporal coding to solve the problem of low output quality of images.
[0006] In a first aspect, this application provides an image acquisition method based on spatio-temporal coding, including:
[0007] Presetting coding rules, where the coding rules include a time coding rule and a space coding rule. The time coding rule is a coding rule generated by controlling a time sequence, and the space coding rule is a coding rule generated by controlling the illuminated area of the light source;
[0008] Collecting a first set of images and / or a second set of images based on the coding rules. The first set of images are images collected based on the time coding rule, and the second set of images are images collected based on the space coding rule;
[0009] Decoding the first set of images and / or the second set of images to output a target image.
[0010] In some feasible embodiments, the preset coding rules include:
[0011] Presetting the number of lighting times and a first reference time, and generating a first time sequence based on the number of lighting times;
[0012] Light up the light source at the first reference time and collect the first test image;
[0013] Calculate the average brightness and overexposed pixel ratio of the first test image, and adjust the first reference time based on the average brightness and overexposed pixel ratio to generate a second time series.
[0014] In some feasible embodiments, the adjusting the first reference time based on the average brightness and overexposed pixel ratio to generate a second time series includes:
[0015] Calculate the number of overexposed pixels, where the number of overexposed pixels is the number of pixels in the first test image whose pixel values are greater than or equal to the overexposure threshold;
[0016] Based on the number of overexposed pixels, calculate the overexposure ratio, where the overexposure ratio is the ratio of the number of overexposed pixels to the number of pixels in the first test image; if the average brightness is less than the first brightness threshold and if the overexposure ratio is less than or equal to the first ratio threshold, adjust the first reference time to a second reference time, and the duration of the second reference time is greater than the duration of the first reference time;
[0017] If the average brightness is greater than the second brightness threshold and if the overexposure ratio is greater than the first ratio threshold, adjust the first reference time to a third reference time, where the first brightness threshold is less than the second brightness threshold, and the duration of the third reference time is less than the duration of the first reference time;
[0018] Generate a second time series based on the second reference time or the third reference time, where the second time series is generated as a multiple of the first time series.
[0019] In some feasible embodiments, the spatial coding rule includes a first spatial coding, a second spatial coding, and a third spatial coding;
[0020] The preset coding rule includes:
[0021] Preset a transmitted light spot, where the light spot is obtained based on the object to be measured and has a coding rule;
[0022] Generate a first spatial coding based on the light spot and collect a second test image based on the first spatial coding;
[0023] Calculate the image complexity, where the image complexity is the complexity of the second test image;
[0024] If the image complexity is greater than the complexity threshold, reduce the size of the light spot and generate a second spatial coding based on the adjusted light spot;
[0025] If the image complexity is less than or equal to the complexity threshold, increase the spot size, and generate a third spatial encoding based on the adjusted spot.
[0026] In some feasible embodiments, calculating the image complexity includes:
[0027] Performing edge detection on the second test image to obtain an edge detection result, where the edge detection result includes one or more of an edge pixel ratio, an edge length, and an edge direction entropy;
[0028] Calculating the image complexity based on the edge detection result.
[0029] In some feasible embodiments, collecting the first set of images and / or the second set of images based on the encoding rule includes:
[0030] Collecting the first set of images based on the encoding rule, where the first set of images is the images collected based on the time encoding rule;
[0031] Collecting the second set of images based on the first set of images through the spatial encoding rule.
[0032] In some feasible embodiments, collecting the first set of images and / or the second set of images based on the encoding rule includes:
[0033] Collecting a third test image;
[0034] Calculating the ambient light intensity of the third test image;
[0035] If the ambient light intensity is greater than the intensity threshold, set a first acquisition order to collect images based on the first acquisition order, where the first acquisition order is to collect the first set of images first and then the second set of images;
[0036] If the ambient light intensity is less than or equal to the intensity threshold, set a second acquisition order to collect images based on the second acquisition order, where the second acquisition order is to collect the second set of images first and then the first set of images.
[0037] In some feasible embodiments, decoding the first set of images and / or the second set of images to output a target image includes:
[0038] If the first set of images is collected based on the encoding rule, obtain the luminance value and the light source lighting time based on the time encoding rule;
[0039] Establishing a time response model, where the time response model is a model established based on the luminance value and the light source lighting time;
[0040] And based on the time response model, calculate a first response coefficient and a second response coefficient, where the first response coefficient is the ambient light coefficient and the second response coefficient is the system light response coefficient;
[0041] Ignore the first response coefficient to output a first set of images, where the first set of images is the images corresponding to the second response coefficient.
[0042] In some feasible embodiments, decoding the first set of images and / or the second set of images to output a target image includes:
[0043] If a second set of images is acquired based on the encoding rule, obtain a first luminance value and a second luminance value based on the spatial encoding rule, where the first luminance value is the maximum luminance when the area is lit, and the second luminance value is the minimum luminance when the area is not lit;
[0044] Calculate a direct light component and an indirect light component, where the direct light component is the difference between the first luminance value and the second luminance value, and the indirect light component is twice the second luminance value;
[0045] Remove the indirect light component to output a second set of images, where the second set of images is the images corresponding to the direct light component.
[0046] In a second aspect, the present application provides an image acquisition system based on spatio-temporal encoding, including:
[0047] An encoding unit for presetting an encoding rule, where the encoding rule includes a time encoding rule and a spatial encoding rule, the time encoding rule is an encoding rule generated by controlling a time sequence, and the spatial encoding rule is an encoding rule generated by controlling the area where the light source is lit;
[0048] An acquisition unit for acquiring a first set of images and / or a second set of images based on the encoding rule, where the first set of images is the images acquired based on the time encoding rule, and the second set of images is the images acquired based on the spatial encoding rule;
[0049] A decoding unit for decoding the first set of images and / or the second set of images to output a target image.
[0050] As can be seen from the above technical solutions, the present application provides an image acquisition method and system based on spatio-temporal coding. The method includes: presetting coding rules, where the coding rules include a time coding rule and a space coding rule. The time coding rule is a coding rule generated by controlling a time sequence, and the space coding rule is a coding rule generated by controlling the illuminated area of a light source; then acquiring a first set of images and / or a second set of images based on the coding rules. The first set of images are images acquired based on the time coding rule, and the second set of images are images acquired based on the space coding rule; outputting a target image by decoding the first set of images and / or the second set of images. The method uses a controllable light source that can perform time coding and space coding, acquires images with coding information, and then decodes them, thereby improving the output quality of the images. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a schematic flowchart of the image acquisition method based on spatio-temporal coding provided by the embodiment of the present application;
[0053] Figure 2 It is a schematic diagram of the composition of the image brightness under ordinary imaging provided by the embodiment of the present application;
[0054] Figure 3 It is a schematic flowchart of generating a second time sequence provided by the embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of the time coding of the light source illumination provided by the embodiment of the present application;
[0056] Figure 5 It is a schematic diagram of a bar-shaped light spot provided by the embodiment of the present application;
[0057] Figure 6 It is a schematic diagram of a checkerboard light spot provided by the embodiment of the present application;
[0058] Figure 7 It is a schematic diagram only distinguished by space coding provided by the embodiment of the present application;
[0059] Figure 8 It is a schematic diagram of acquisition under different illumination times provided by the embodiment of the present application;
[0060] Figure 9 It is a schematic diagram of the decoding method provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0062] In machine vision, an image sensor is a component that captures external light. Its working process can be divided into four stages: photon reception, charge conversion and storage, signal processing, and digital output. In the photon reception and charge conversion stage, each pixel of the image sensor contains a tiny photosensitive unit, such as a photodiode. When light irradiates the photosensitive unit, the energy of the photons is absorbed, exciting electrons in the semiconductor material to form charges. The stronger the light, the more charges are excited.
[0063] The generated charges will be temporarily stored in the charge storage area within the pixel. The CCD sensor transfers charges row by row through an electric field, while the CMOS sensor allows each pixel to store charges independently and then reads them one by one through a circuit, so that the charges are retained in an orderly manner, reducing signal confusion.
[0064] In the signal processing and noise reduction stage, after the charges are converted into voltage signals, the voltage signals generated by the tiny charges are weak and need to be amplified to a processable range. The sensor will perform two samplings, where the two samplings include background noise and noise sampling containing signals, to deduct the noise generated by the circuit itself. The circuit filters out high-frequency interference, such as electromagnetic wave interference, and retains the effective signals.
[0065] In the digitization and output stage, the amplified analog voltage is converted into a digital signal by an analog-to-digital converter. For example, a 12-bit converter will divide the voltage into 4096 levels of precision, and finally generate values from 0 to 4095. These values are arranged and combined by pixels to form the original image data, which is transmitted to a computer through a standard interface (such as USB, HDMI) for subsequent processing.
[0066] Determines the ability of the sensor to simultaneously capture the brightest and darkest details, similar to the effect of the HDR mode of a camera, affecting the imaging quality in low-light environments. A high-sensitivity sensor can clearly image in moonlight. Even in complete darkness, the sensor will generate slight noise due to heat.
[0067] The processing speed and storage capacity can be improved by laminating the photosensitive layer and the processing circuit. It only records the instant of light change and is suitable for capturing high-speed moving objects, such as splashing water droplets. Special materials such as quantum dots are used to expand the photosensitive range to achieve the capture of infrared and ultraviolet light. Through this series of steps, the image sensor accurately converts external light into digital signals that can be recognized by a computer.
[0068] The imaging quality of an image sensor is limited by its inherent characteristic of being unable to distinguish the source of photons. For directly reflected light, photons are emitted from the light source, reflected once by the object surface, and then directly enter the sensor, carrying the true topographical information of the target object. For indirectly reflected light, photons enter the sensor after multiple reflections, refractions on the object surface, or scattering by environmental objects, forming interference noise, such as specular highlights and environmental halos. In ambient stray light, non-system light sources, such as sunlight and lamp light, directly irradiate the sensor, superimposing on the active light source signal.
[0069] Directly reflected light, i.e., light directly reflected from the object surface, indirectly reflected light, i.e., light indirectly coming after multiple reflections from the object surface, and ambient light, i.e., scattered light from the surrounding environment, will all affect the quality of the output image.
[0070] Exemplarily, materials such as polished metal and mirror plastics cause strong specular reflections, forming bright spots in the image that obscure tiny defects such as scratches and pits. The curvature of spherical objects causes the reflection angles of light to disperse, and edge detection algorithms misidentify the reflection gradient as the actual contour. In an industrial workshop scenario, when the detection device is close to a window, sunlight directly irradiates the surface of metal workpieces, and the sensor simultaneously receives the active light source and ambient light, resulting in color distortion.
[0071] Specifically, characteristics of the imaging environment such as brightness, contrast, and color distribution, properties of the object surface such as roughness, glossiness, and color, as well as geometric features of the object such as shape, size, and position, will all to a certain extent determine the image quality captured by the image sensor. Therefore, in many practical application scenarios, due to the complexity and uncontrollability of these factors, the image quality output by the image sensor is poor.
[0072] This is not only manifested as a decrease in indicators such as image clarity and resolution but may also cause problems such as image noise, distortion, and blurring. These problems will limit the application and development of machine vision technology in certain fields, especially in high-precision measurement, recognition, analysis, etc. fields with strict requirements for image quality.
[0073] For the suppression of ambient light, the ambient light can be covered by a high-power LED array or a laser light source. For example, in pulse synchronous illumination, the light source is strictly synchronized with the camera exposure time, and high-intensity light pulses are output within an extremely short time window. Another example is a multi-band light source, which uses a specific wavelength light source in combination with a narrow-band filter to suppress visible light interference.
[0074] However, high-power light sources result in a large system volume and require a forced heat dissipation device, making them inapplicable to portable devices. The narrow-band filter solution cannot adapt to multi-spectral application scenarios. In scenarios where the light intensity changes rapidly, fixed-parameter light sources cannot respond in real time.
[0075] In some embodiments, multi-view imaging fusion can also be used, that is, by arranging multiple cameras and light sources at different angles, for example, an annular light source and coaxial light, and fusing images under different lighting conditions through an algorithm. However, multi-view imaging is costly and requires the configuration of multiple cameras and complex mechanical structures, such as a six-axis robot, resulting in an exponential increase in integration complexity.
[0076] In summary, that is to say, it is impossible to fundamentally remove ambient light, resulting in low output quality of the image.
[0077] To solve the problem of low output quality of the image, some embodiments of the present application provide an image acquisition method based on spatio-temporal coding. By using a controllable light source that can perform time coding and space coding, images with coding information are acquired and then decoded to improve the output quality of the image.
[0078] As Figure 1 shown, the method includes the following steps:
[0079] S100: Preset a coding rule.
[0080] In this embodiment, the coding rule includes a time coding rule and a space coding rule. The time coding rule is a coding rule generated by controlling the time sequence and can encode information according to the passage of time or changes in brightness; the space coding rule is a coding rule generated by controlling the illuminated area of the light source and can encode information according to the specific area illuminated by the light source.
[0081] Time coding is achieved by finding the following functional relationship between the light source lighting time and the image brightness:
[0082] y = f(t);
[0083] where t = 0 is the brightness y of the image when the light source is not lit 0 = f(0), and when the light source is lit for 1 time unit, t = 1, the brightness y of the image 1 = f(1). Then, y 1 - y 0 is to eliminate the influence of the ambient light source.
[0084] It is equivalent to the brightness of the image when the light source is lit for 1 time unit t = 1 in a pure black environment; thus, the image brightness at any time when the light source is lit can be obtained through the functional relationship, and then the image data of any area of the image at any time when the light source is lit can be output.
[0085] In this state, the intelligent camera can output an image or simply output the functional relationship, and when an image is needed, the required image can be generated according to the functional relationship.
[0086] Spatial encoding is achieved by lighting the light sources in different regions to find the following functional relationship between the illuminated region of the light source and the image brightness:
[0087] y = f(s);
[0088] where s = 1 represents that the light source in the region is lit. At this time, the image brightness y 1 = f(1). When s = 0, it represents that the light source in the region is not lit and the surrounding light sources are lit. At this time, the image brightness y 0 = f(0). Through this relationship, the intelligent camera can output image data that is only illuminated by the light source in this region, thereby further eliminating the influence of other regions of its own light source on this region.
[0089] In principle, time encoding can eliminate ambient light interference, and spatial encoding can eliminate the interference of light reflected and diffracted from other regions of its own illumination into this region. For example, Figure 2 as shown, the entire region is composed of the image brightness under ordinary imaging. Using time encoding is similar to dividing by the line pointed by A, into ambient light and system light. Using spatial encoding is similar to dividing by the line pointed by B, into direct light and indirect light.
[0090] For the time encoding sequence, including the first time sequence and the second time sequence, in some embodiments, a preset number of lighting times and a first reference time are set, and the first time sequence is generated based on the number of lighting times; the light source is lit at the first reference time, and the first test image is collected; the average brightness and the overexposed pixel ratio of the first test image are calculated to adjust the first reference time based on the average brightness and the overexposed pixel ratio to generate the second time sequence.
[0091] The number of lighting times is the number of pulsed emissions of the light source in a single trigger. For example, 5 times. The initial value of the number of lighting times is set according to the object movement speed, which is used to balance the time resolution and processing efficiency in a dynamic scene. The first reference time is the initial time of time encoding and is the minimum duration of a single lighting.
[0092] The first time sequence is a linearly increasing sequence (1t, 2t, 3t,...) or a geometric sequence (t, 2t, 4t,...) generated based on the first reference time. The first time sequence is dynamically adjusted, and the dynamic adjustment types include stopping subsequent acquisitions after reaching the brightness index and generating a new sequence based on overexposed pixels.
[0093] In some embodiments, the number of overexposed pixels is calculated; then, based on the number of overexposed pixels, the overexposed ratio is calculated, and the overexposed ratio is the ratio of the number of overexposed pixels to the number of pixels in the first test image.
[0094] The overexposure quantity is the number of pixels in the first test image whose pixel values are greater than or equal to the overexposure threshold, that is, the total number of pixels in a single-frame image whose pixel values exceed the linear response range of the sensor. Among them, the overexposure threshold can be determined according to the characteristics of the sensor. It can be understood that for different objects, the overexposure threshold is different.
[0095] Preset a first brightness threshold, a second brightness threshold, and a first ratio threshold. The first brightness threshold is less than the second brightness threshold. For example, the first brightness threshold is 50, the second brightness threshold is 200, and the first ratio threshold is 10%.
[0096] Such as Figure 3 As shown, if the average brightness is less than the first brightness threshold and if the overexposure ratio is less than or equal to the first ratio threshold, adjust the first reference time to the second reference time, and the duration of the second reference time is greater than the duration of the first reference time.
[0097] If the average brightness is greater than the second brightness threshold and if the overexposure ratio is greater than the first ratio threshold, adjust the first reference time to the third reference time, and the duration of the third reference time is less than the duration of the first reference time.
[0098] Based on the second reference time or the third reference time, generate a second time series, and the second time series is generated as a multiple of the first time series.
[0099] Exemplarily, the number of lighting times is five, and the lighting time for each time is 1, 2, 3, 4, 5 respectively; when the intelligent camera processes internally, it will detect the brightness index of the original image collected each time. For example, when the brightness index is reached at the third time, the shooting of 4 and 5 will not be carried out; when an external signal is input each time, an additional dynamic detection process is added to find the appropriate lighting time ratio, and according to the lighting time ratio, the value of the time encoding is adjusted. For example, if the time ratio is 2, then the five lighting times are 2, 4, 6, 8, 10 respectively, and the brightness index can be the average brightness of the image or the proportion of overexposed pixels in the image.
[0100] Such as Figure 4 As shown, the camera timing represents the exposure time window of the camera, marked as exposure time m, that is, the time period during which the camera sensor receives the optical signal. During this time period, the camera completes the acquisition of a single-frame image. The acquisition period k is the total period for the camera and the light source to work together, including a complete image acquisition process, such as multiple light source illuminations and camera exposures. The light source encoding 1 and the light source encoding 2 are the lighting timings of two independent light sources or different regions of the same light source, representing different encoding rules or time parameters.
[0101] The light source lighting times n1 and n2 represent the specific lighting times of the light source within the coding period (such as 1 ms, 2 ms), which are used to generate different optical signal characteristics. For the lighting state of the light source, a high level represents lighting, and a low level represents extinction.
[0102] Camera exposure synchronization. The camera turns on the sensor during the exposure time m, and the light source lights up according to the coding rule during this period. For example, in the time periods n1 and n2. It can be understood that the light source lighting time needs to completely cover the camera exposure window; otherwise, signal loss or noise increase will occur.
[0103] Light source coding 1: The lighting time is n1 in the first sub-period, generating a specific optical signal (such as high brightness in a short time). Light source coding 2: The lighting time is n2 in the second sub-period, generating a differentiated optical signal (such as low brightness in a long time). The optical signals encoded at different times are used to separate the ambient light (constant component) and the system light source contribution (time-varying component). Within each period k, light source coding 1 and coding 2 are executed in sequence, and the camera synchronously acquires multiple groups of images (for example, the acquisition period k = 3 means repeating 3 times).
[0104] The light source lights up according to the preset time sequence (n1, n2,...), and the camera acquires the corresponding image groups. By fitting the function of the pixel brightness and the lighting time, the ambient light and the light source contribution are separated.
[0105] Light source coding 1 and light source coding 2 can represent different regions (such as the left half region and the right half region) or different wavelengths (such as red light and blue light), which are used for signal separation in the spatial or spectral dimension.
[0106] Exemplarily, the first reference time is 1 ms, the number of lighting times is 5, the first brightness threshold is 150, the second brightness threshold is 200, and the first ratio threshold is 10%. For the first test, an image with t = 1 ms is acquired, and the calculated average brightness is 120, and the exposure ratio is 5%. The first reference time is adjusted to the second reference time 1×1.5 = 1.5 ms, and the second sequence is generated: [1.5 ms, 3 ms, 4.5 ms, 6 ms, 7.5 ms] (ratio method, k1 = 1.5).
[0107] For the second test, an image with t = 1.5 ms is acquired, and the calculated average brightness is 170, and the exposure ratio is 8%. The average brightness is within the target range, and the exposure ratio < 10%. Then the current sequence, that is, the first time sequence, is continued for acquisition.
[0108] For spatial coding, the spatial coding rule includes first spatial coding, second spatial coding, and third spatial coding.
[0109] In some embodiments, a preset emission light spot is provided, a first spatial encoding is generated based on the light spot, and a second test image is acquired based on the first spatial encoding; then the image complexity is calculated, where the image complexity is the complexity of the second test image; if the image complexity is greater than the complexity threshold, the size of the light spot is reduced, and a second spatial encoding is generated based on the adjusted light spot; if the image complexity is less than or equal to the complexity threshold, the size of the light spot is increased, and a third spatial encoding is generated based on the adjusted light spot.
[0110] The light spot is an optical pattern with a specific encoding rule projected by a light source. For example, it can be stripes, checkerboards, concentric circles, etc. It can be generated through a pre-stored mode or dynamically. The pre-stored mode is that an FPGA stores standard templates such as Gray codes and sine stripes. For example, an 8-bit Gray code covers a 1024×768 area. Dynamic generation is to calculate the optimal light spot in real time according to the object contour. For example, radial stripes are generated for a curved object.
[0111] As Figure 5 and Figure 6 shown, among them, the stripe encoding is equally spaced black and white stripes, such as Gray codes and sine waves, which can reduce the error between the plane and a simple curved surface; the checkerboard encoding is a regular black and white checkerboard partition, such as 8×8, which can be applied to complex curved surfaces to increase the spatial resolution by 3 times; the composite encoding is stripes and random dot matrices, such as those dynamically generated by a DMD, which are applied to multi-directional reflection surfaces and can increase the anti-interference ability by 80%.
[0112] For the selection of the light spot, it can be selected according to different conditions. For example, the surface characteristics of the object to be measured, the intensity of ambient light interference, and the encoding efficiency requirements.
[0113] If the surface of the object is reflective, such as metal or glass, a checkerboard light spot is preferably selected to separate the ambient light reflection noise using its periodic characteristics. If the surface of the object has diffuse reflection and complex texture, such as fabric or rough material, a strip light spot is selected to enhance the feature contrast.
[0114] Under strong ambient light, a light spot with a high modulation depth (such as a binary-encoded strip light spot) needs to be selected. In a weak ambient light scenario, a Gray code checkerboard light spot can be selected to increase the encoding density. For a dynamic scenario (such as a moving object) that requires fast projection decoding, a binary bar code (with high encoding efficiency) is preferably selected. For static high-precision measurement, a multi-frequency phase encoding + checkerboard light spot (with strong anti-noise ability) can be selected.
[0115] In this embodiment, the light spot is a light spot with an encoding rule obtained based on the object to be measured. In some embodiments, the surface characteristics of the object to be measured are analyzed, and then the encoding rule is matched to perform light spot transmission and verification.
[0116] Specifically, the input is a pre-acquired object contour image, with the light source in full-bright mode, analyzing the reflectance distribution (specular / diffuse reflection ratio) and geometric complexity (curvature change rate, edge density). Exemplarily, for an electroplated part, the specular reflection ratio on the surface is >90%, triggering high-frequency stripe coding; for a frosted plastic, diffuse reflection is dominant, using low-frequency wide stripes.
[0117] The pre-stored pattern library can be set, and the coding rules can be matched by calling the pre-stored pattern library. For example, when the object type is a flat metal plate, the matching coding rule is horizontal equidistant stripes with a spacing of 5 mm; when the object type is a gear tooth surface, the matching coding rule is radial stripes with a spacing of 1 mm.
[0118] For the dynamic generation algorithm, it is triggered when the image complexity is greater than the threshold. When the image complexity exceeds the threshold, the spot size is reduced, which enhances the analytical ability of complex surface features by improving the spatial resolution. When the surface complexity of the object is greater than the threshold, large-sized spots will cause the light intensity superposition of adjacent reflection regions, forming false edge signals.
[0119] Small-sized spots can align the checkerboard coding units with the microscopic structure of the object, preventing the coding from spanning multiple feature regions. An optimal observation operator is constructed to improve the feature distinguishability of high-complexity scenes.
[0120] Finally, the light spots are projected according to the coding rules by a DLP projector or an LED array, the light spot images are collected, the modulation transfer function (MTF) is calculated to verify the effectiveness. If MTF > 0.6 (threshold), the coding is qualified; if MTF ≤ 0.6, denser light spots are regenerated.
[0121] Exemplarily, for the detection of highly reflective metal parts, due to specular reflection causing spot distortion and decoding failure, dynamic adjustment can be performed. If the initial coding is horizontal stripes (5 mm spacing) and spot breakage occurs, it can be switched to orthogonal dual-frequency stripes, i.e., horizontal 5 mm + vertical 5 mm, with a phase difference of 90°, which can improve the decoding success rate.
[0122] After generating the first spatial coding through the selected light spots, images are collected through the first spatial coding. During the spatial coding process, the complexity of the measured target affects the direct and indirect light separation effect. The more complex the object, the more refined the coding light is required. Therefore, image complexity calculation is performed. Image complexity can be calculated in various ways. For example, image edge detection can be used to calculate the length of the image edge and the complexity of the edge.
[0123] In some embodiments, edge detection is performed on the second test image to obtain an edge detection result, which includes one or more of an edge pixel ratio, an edge length, and an edge direction entropy, to calculate the image complexity based on the edge detection result.
[0124] The image edge is the area where the gray value changes in the image, and these areas reflect the boundary information of the objects in the image. The complexity of the edge can be measured from multiple perspectives, including the number of edges, the total length of the edges, the change of the edge direction, and the tortuosity of the edges, etc. The richness and complexity of the edges are usually proportional to the complexity of the image. That is to say, the richer and more complex the edges are, the higher the overall complexity of the image is.
[0125] The more complex the image is, the more complex the surface morphology of the measured object is. When light hits the object surface, multiple reflections will occur, thus affecting the imaging effect. Edge detection operators include Sobel operator, Canny operator, etc. The Sobel operator is an edge detection operator based on the first derivative, which detects edges by calculating the gradients of the image in the horizontal and vertical directions. The Canny operator is a multi-stage edge detection algorithm, which has the advantages of low false detection rate, high positioning accuracy and single edge response.
[0126] The Sobel operator calculates the horizontal / vertical gradient through a 3×3 convolution kernel, and the Canny operator performs Gaussian filtering, gradient calculation, non-maximum suppression, and double-threshold connection.
[0127] The indicators for evaluating the edge complexity include the edge pixel ratio, the edge length, and the edge direction entropy. The ratio of the number of edge pixels to the total number of pixels in the image, the larger the ratio, the richer the edges are, and the higher the complexity may be. The longer the edge length is, the higher the complexity is. The entropy of the edge direction can measure the diversity of the edge direction. The larger the entropy value is, the more complex the edge direction is. According to the complexity indicators, the size of the barcode or checkerboard is set to achieve the best effect.
[0128] Among them, the edge pixel ratio is obtained by counting the number of white pixels in the binary image output by Sobel / Canny. The total edge length is the sum of the pixel lengths of the continuous edge chains, which is calculated by the chain code method or approximate calculation. The chain code method tracks the edge direction and accumulates the step size; the approximate calculation is calculated by the number of edge pixels and the hypotenuse correction coefficient. The edge direction entropy is the degree of disorder of the edge direction distribution, which reflects the geometric structure complexity and can be obtained by statistically analyzing the Sobel gradient direction histogram.
[0129] Exemplarily, the input image is converted into a grayscale image; the Sobel operator calculates the gradient magnitude and direction, then performs binarization processing, outputs an edge binary image, counts the number of white pixels in the binary image, calculates the total edge length and the edge direction entropy, and then determines the complexity. If the complexity is greater than the complexity threshold, the spot size is reduced or switched to composite coding, the stripe width is reduced from 5 mm to 2 mm or switched to a combination of radial stripes and random dot matrices.
[0130] In some embodiments, the comprehensive complexity can also be calculated by multi-detector fusion, that is, Sobel, Canny, and Laplacian operators are used jointly, and weighted comprehensive complexity is calculated. Weights are set for Sobel, Canny, and Laplacian operators to comprehensively calculate the complexity.
[0131] In this embodiment, whether it is through preset time coding or space coding, the coding and decoding are controlled by the controller, which is used as the main control unit to generate trigger signals, execute preset coding rules, and coordinate the synchronization of light sources and cameras. It can include an ARM processor and programmable logic, which is solidified in the Flash memory and includes time / space coding sequences, trigger timing parameters, etc.
[0132] The encoding rules are written into the controller through the interface, and the decoding algorithm is burned into the BRAM of the FPGA simultaneously. When an external trigger is input, the controller controls the camera to collect image data according to the burned encoding rules, and the FPGA decoder decodes and outputs specific image data or corresponding functional relationships based on the input raw data. When the controller and decoder programs are burned, the rules are fixed and will not change. If you want to change them, you need to dynamically set the encoding and decoding rules supported by the current program.
[0133] S200: Acquire a first group of images and / or a second group of images based on a coding rule.
[0134] According to the above coding rules, the first group of images are images collected based on the time coding rules, which can reflect the changes in the time series or brightness series. The second group of images are images collected based on the space coding rules, which can reflect the changes in the light source lighting area. In this way, the coding rules can be effectively used to collect image data corresponding to the coding rules.
[0135] In some embodiments, a first group of images is acquired based on the coding rule, and the first group of images is images acquired based on the time coding rule; then based on the first group of images, a second group of images is acquired through the spatial coding rule. In other words, the second group of images is images acquired based on the first group of images, that is, the images processed by spatial coding need to remove ambient light first, and then separate direct light and indirect light. If spatial coding is used alone, ambient light will be classified as indirect light.
[0136] like Figure 7 As shown, after time encoding, the output image is the separated ambient light and system light, and after space encoding, the output image is the ambient light, system direct light, and system indirect light.
[0137] It is understandable that temporal encoding has no necessary connection with spatial encoding and can be used alone or in series. However, in order to remove ambient light, in this embodiment, before using spatial encoding, temporal encoding is used to process the image.
[0138] If used in series, in some cases, when temporal encoding is not used to process the image before using spatial encoding, the acquisition order is determined by the ambient light intensity.
[0139] In some embodiments, a third test image is acquired, and the ambient light intensity of the third test image is calculated; if the ambient light intensity is greater than the intensity threshold, a first acquisition order is set to acquire images based on the first acquisition order, where the first acquisition order is to acquire the second group of images after acquiring the first group of images; if the ambient light intensity is less than or equal to the intensity threshold, a second acquisition order is set to acquire images based on the second acquisition order, where the second acquisition order is to acquire the first group of images after acquiring the second group of images.
[0140] By measuring the ambient light intensity, for example, the average gray value, the optical power density, and dynamically selecting the acquisition order of temporal encoding, spatial encoding or spatial encoding, temporal encoding, the problem of the active light source signal being submerged due to strong ambient light interference, insufficient signal-to-noise ratio in weak light scenarios (such as night detection, enclosed cavities), and imaging quality fluctuations caused by dynamic light mutations (such as welding arcs, cloud occlusion) can be solved.
[0141] The ambient light intensity is a quantization value of the external light intensity of the non-system active light source, which can be detected by a photosensitive sensor or image analysis method. The photosensitive sensor is independent of the photodiode of the camera and outputs the optical power density; the image analysis method acquires a single-frame image after turning off the active light source and calculates the average gray value of the whole image. The intensity threshold is the critical light intensity that triggers the switching of the acquisition order and can be set according to the actual situation.
[0142] When the ambient light intensity is greater than the intensity threshold, that is, in the strong light state, temporal encoding is first performed to eliminate ambient light interference, and then spatial encoding is performed to eliminate multiple reflections. For example, the light source is lit in a time series (t1 = 1ms, t2 = 3ms, t3 = 5ms), the FPGA decodes and outputs the image after ambient light compensation, projects the spatial encoding spot, and acquires and decodes the direct light component.
[0143] When the ambient light intensity is less than or equal to the intensity threshold, that is, in the weak light state, spatial encoding is first performed to extract the geometric features of the object), and then the signal is enhanced through temporal encoding. For example, a high-density stripe spot is projected, the geometric contour is acquired and decoded, the light source is lit according to the brightness gradient (P1 = 30%, P2 = 60%, P3 = 100%), and the FPGA fuses the spatio-temporal data to generate an image with a high signal-to-noise ratio, which can be applied in an enclosed cavity.
[0144] In some embodiments, multi-threshold hierarchical control can also be set. The intensity thresholds include a first intensity threshold, a second intensity threshold, and a third intensity threshold. When the ambient light intensity is less than the first intensity threshold, it is in the ultra-weak light state, and the acquisition sequence is spatial encoding, time encoding, and brightness enhancement. When the ambient light intensity is greater than or equal to the first intensity threshold and less than the second intensity threshold, it is in the mixed state, and the acquisition sequence is spatial encoding first and then time encoding. When the ambient light intensity is greater than or equal to the second intensity threshold and less than the third intensity threshold, it is in the standard strong light state, and the acquisition sequence is time encoding first and then spatial encoding. When the ambient light intensity is greater than or equal to the third intensity threshold, it is in the enhanced anti-interference state, and the acquisition sequence is time encoding first and then spatial encoding, and secondary time encoding can also be performed.
[0145] The first set of acquired images is the image with the ambient light separated out. The second set of images can be the image with the direct light and indirect light separated out, but the indirect light contains the ambient light, or it can be the image with the direct light and indirect light separated out, but the indirect light does not contain the ambient light, that is, the image obtained by first performing time encoding and then spatial encoding.
[0146] S300: Decode the first set of images and / or the second set of images to output the target image.
[0147] After the image data is acquired, then decode the first set of images and / or the second set of images to output the target image. Through the decoding process, the information encoded in the images can be extracted, and then the target image can be generated.
[0148] For the first set of images, the intelligent camera decodes the time encoding (processing each pixel). For the same pixel of a set of encoded images, the corresponding relationship between the pixel value and the light source lighting time, and the relationship between the brightness of the image and the light source lighting time are obtained.
[0149] In some embodiments, obtain the brightness value and the light source lighting time based on the time encoding rule; establish a time response model, where the time response model is a model established based on the brightness value and the light source lighting time; and based on the time response model, calculate a first response coefficient and a second response coefficient, where the first response coefficient is the ambient light coefficient and the second response coefficient is the system light response coefficient; ignore the first response coefficient to output the first set of images, and the first set of images is the image corresponding to the second response coefficient.
[0150] As Figure 8 shown, the brightness value is the gray value of a single pixel or a pixel region at a specific lighting time, and the light source lighting time is the effective light emission duration of the controllable light source within the camera exposure period.
[0151] The time response model is as follows:
[0152] z = k 0 + k 1 n + k 2 n 2 = φ(n);
[0153] Where z is the pixel value (image brightness value), n is the exposure time of the light source (the time when the light source is on), and k 0 , k 1 , k 2 is the response coefficient. Among them, k 0 is the first response coefficient, k 1 , k 2 is the second response coefficient. It can be seen from the model that k 0 is the influence of ambient light on the imaging system.
[0154] To obtain the response coefficient of the camera, for each n i , the sum of the squares of the differences between the pixel value calculated through n i and z i should be minimized, that is:
[0155]
[0156] Taking k 0 , k 1 , k 2 as variables, taking the partial derivative of the above formula, the following 3 groups of equations are obtained:
[0157]
[0158] Solving the values of each k by the elimination method. After obtaining the values of k, for any light source on time n, its corresponding ideal camera response value is the following formula:
[0159] z n = k 0 + k 1 x n + k 2 x n 2 ;
[0160] Where x n is the light source on time.
[0161] It can be foreseen that the image output is the camera response value after removing the ambient interference term k 0 The ambient interference is removed. The image output is the camera response value corresponding to any light source on time. The image output is the camera response value obtained by fusing the camera response values corresponding to several light source on times according to a fixed rule. The camera response value is the pixel value of the image, which is the number of electrons collected by each pixel during one acquisition inside the camera.
[0162] For the second set of images, the intelligent camera decodes the spatial encoding (processing each pixel), and for the same pixel of a set of encoded images, obtains the correspondence between the pixel value and the illuminated area of the light source.
[0163] In some embodiments, if the second set of images is acquired based on the encoding rule, the first luminance value and the second luminance value are obtained based on the spatial encoding rule, where the first luminance value is the maximum luminance when the area is illuminated, and the second luminance value is the minimum luminance when the area is not illuminated; calculate the direct light component and the indirect light component, where the direct light component is the difference between the first luminance value and the second luminance value, and the indirect light component is twice the second luminance value; eliminate the indirect light component to output the second set of images, and the second set of images is the image corresponding to the direct light component.
[0164] The first luminance value is the pixel peak gray value of the imaging area corresponding to when the target area is fully powered on and illuminated, and is obtained by directly irradiating the target area with the light source. The second luminance value is the pixel valley gray value of the imaging area corresponding to when the target area is turned off but the adjacent area is illuminated, and is obtained by turning off or blocking the light source.
[0165] When there is direct light irradiation, the brightness is higher, and when there is no direct light irradiation, the brightness is lower. Therefore, the direct light component and the indirect light component are separated by the following formula:
[0166]
[0167] where L d is the direct light component, L g is the indirect light component, L max is the first luminance value, L min is the second luminance value.
[0168] In addition to decoding through the direct light component and the indirect light component, as Figure 9 shown, each corresponding pixel value of spatial encoding 1 and encoding 2 has a logical inverse relationship. For example, encoding 1 is a checkerboard substrate, such as 3x3 black blocks spaced with white blocks, and encoding 2 is an inverted checkerboard, where the white blocks replace the positions of the black blocks. When a certain area of encoding 1 is highlighted (white block), the corresponding area of encoding 2 is dark (black block). The complementary characteristic can achieve the cancellation of ambient light and only retain the true reflection information of the object surface.
[0169] In some embodiments, for direct light extraction, the common ambient light is eliminated by the difference and then the ratio with a preset value to restore the true reflection intensity, where the preset value is 2; for indirect light extraction, the ambient light and the multiple reflection components are retained by summing and averaging. Through the double-image difference solution of complementary spatial encoding, the accurate stripping of ambient light and the extraction of the true reflection characteristics of the object are realized.
[0170] The method provided by this application realizes efficient control of the image acquisition process by combining the time encoding rule and the space encoding rule. Among them, the time encoding rule realizes precise regulation of the image acquisition time point by controlling the time sequence, and the space encoding rule realizes precise control of the image acquisition spatial position by controlling the light source lighting area. This encoding method that combines time and space not only improves the flexibility and accuracy of image acquisition, but also effectively improves the efficiency of image acquisition.
[0171] Based on the above image acquisition method based on spatio-temporal encoding, some embodiments of this application also provide an image acquisition system based on spatio-temporal encoding, including:
[0172] An encoding unit for presetting an encoding rule, the encoding rule including a time encoding rule and a space encoding rule, the time encoding rule being an encoding rule generated by controlling the time sequence, and the space encoding rule being an encoding rule generated by controlling the light source lighting area;
[0173] An acquisition unit for acquiring a first set of images and / or a second set of images based on the encoding rule, the first set of images being images acquired based on the time encoding rule, and the second set of images being images acquired based on the space encoding rule;
[0174] A decoding unit for decoding the first set of images and / or the second set of images to output a target image.
[0175] It can be understood that it also includes a controller for generating a trigger signal, executing a preset encoding rule, and coordinating the synchronous actions of the light source and the camera.
[0176] For the effects of the above system embodiments during operation, reference can be made to the effects of the above method embodiments, which will not be elaborated here.
[0177] For the similar parts between the embodiments provided by this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.
Claims
1. An image acquisition method based on spatiotemporal coding, characterized in that: include: Preset coding rules, the coding rules include time coding rules and space coding rules, the time coding rules are coding rules generated by controlling the time sequence, and the space coding rules are coding rules generated by controlling the light source lighting area; Acquire a first group of images and / or a second group of images based on the coding rule, wherein the first group of images are images acquired based on the time coding rule, and the second group of images are images acquired based on the space coding rule; The first group of images and / or the second group of images are decoded to output a target image.
2. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The preset encoding rules include: Preset the number of lighting times and the first reference time, and generate a first time series based on the number of lighting times; Lighting up the light source at the first reference time to capture a first test image; The average brightness and the overexposed pixel ratio of the first test image are calculated, and the first reference time is adjusted based on the average brightness and the overexposed pixel ratio to generate a second time series.
3. The image acquisition method based on spatiotemporal coding according to claim 2, characterized in that: The step of adjusting the first reference time based on the average brightness and the overexposed pixel ratio to generate a second time series includes: Calculating the overexposure number, where the overexposure number is the number of pixels of the first test image whose pixel values are greater than or equal to an overexposure threshold; Based on the overexposure quantity, calculating an overexposure ratio, where the overexposure ratio is a ratio of the overexposure quantity to the number of pixels of the first test image; if the average brightness is less than a first brightness threshold, and if the overexposure ratio is less than or equal to a first ratio threshold, adjusting the first reference time to a second reference time, wherein the duration of the second reference time is greater than the duration of the first reference time; If the average brightness is greater than a second brightness threshold, and if the overexposure ratio is greater than a first ratio threshold, adjusting the first reference time to a third reference time, the first brightness threshold is less than the second brightness threshold, and the duration of the third reference time is less than the duration of the first reference time; A second time series is generated based on the second reference time or the third reference time, wherein the second time series is generated based on a multiple of the first time series.
4. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The space coding rule includes a first space coding, a second space coding and a third space coding; The preset encoding rules include: A preset emission light spot, wherein the light spot is a light spot obtained based on the object to be measured and has a coding rule; generating a first spatial code based on the light spot, and acquiring a second test image based on the first spatial code; Calculating image complexity, where the image complexity is the complexity of the second test image; If the image complexity is greater than a complexity threshold, reducing the spot size, and generating a second spatial code based on the adjusted spot; If the image complexity is less than or equal to the complexity threshold, the spot size is increased, and a third spatial code is generated based on the adjusted spot.
5. The image acquisition method based on spatiotemporal coding according to claim 4, characterized in that: The calculating image complexity comprises: Performing edge detection on the second test image to obtain an edge detection result, wherein the edge detection result includes one or more of edge pixel ratio, edge length, and edge direction entropy; The image complexity is calculated based on the edge detection result.
6. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The collecting of the first group of images and / or the second group of images based on the encoding rule comprises: Acquire a first group of images based on the coding rule, the first group of images being images acquired based on the time coding rule; Based on the first group of images, a second group of images is acquired using the spatial encoding rule.
7. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The collecting of the first group of images and / or the second group of images based on the encoding rule comprises: Acquiring a third test image; Calculating the ambient light intensity of the third test image; If the ambient light intensity is greater than the intensity threshold, setting a first acquisition sequence to acquire images based on the first acquisition sequence, wherein the first acquisition sequence is to acquire a second group of images after acquiring a first group of images; If the ambient light intensity is less than or equal to the intensity threshold, a second acquisition sequence is set to acquire images based on the second acquisition sequence, wherein the second acquisition sequence is to acquire the first group of images after acquiring the second group of images.
8. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The decoding of the first group of images and / or the second group of images to output a target image comprises: If the first group of images is acquired based on the coding rule, the brightness value and the light source lighting time are acquired based on the time coding rule; Establishing a time response model, wherein the time response model is a model established based on the brightness value and the lighting time of the light source; and based on the time response model, calculating a first response coefficient and a second response coefficient, wherein the first response coefficient is an ambient light coefficient, and the second response coefficient is a system light response coefficient; The first response coefficient is ignored to output a first group of images, where the first group of images is images corresponding to the second response coefficient.
9. The image acquisition method based on spatiotemporal coding according to claim 1, characterized in that: The decoding of the first group of images and / or the second group of images to output a target image comprises: If a second group of images is acquired based on the coding rule, a first brightness value and a second brightness value are acquired based on the spatial coding rule, wherein the first brightness value is the maximum brightness when the area is lit, and the second brightness value is the minimum brightness when the area is not lit; Calculate a direct light component and an indirect light component, wherein the direct light component is a difference between a first brightness value and a second brightness value, and the indirect light component is twice the second brightness value; The indirect light component is eliminated to output a second group of images, where the second group of images are images corresponding to the direct light component.
10. An image acquisition system based on spatiotemporal coding, characterized in that: The method for image acquisition based on spatiotemporal coding according to any one of claims 1 to 9 comprises: A coding unit, used to preset coding rules, wherein the coding rules include time coding rules and space coding rules, wherein the time coding rules are coding rules generated by controlling the time sequence, and the space coding rules are coding rules generated by controlling the light source lighting area; an acquisition unit, configured to acquire a first group of images and / or a second group of images based on the coding rule, wherein the first group of images are images acquired based on the time coding rule, and the second group of images are images acquired based on the space coding rule; A decoding unit is used to decode the first group of images and / or the second group of images to output a target image.