An image processing method, apparatus and device

CN122760656APending Publication Date: 2026-09-15HANGZHOU HIKROBOT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610943016.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0003]若第一原始图像的采集时间与第二原始图像的采集时间不同,那么,待测物体是高速运动物体时,即使采集时间的时间间隔较小,待测物体在该时间间隔的运动位移量也较大,从而导致待测物体在第一原始图像和第二原始图像中对应的状态不一致,破坏多目几何一致性

Benefits of technology

[0009] As can be seen from the above technical solutions, in this embodiment, the spatial offset of the object under test in the first time interval can be determined based on prior external motion information. Motion error compensation is then performed on the first original image based on this spatial offset to obtain a first candidate image. The depth information of the object under test is then determined based on the first candidate image and the second original image. Thus, even if the acquisition time of the first original image differs from that of the second original image, and the object under test is a high-speed moving object, motion error compensation can ensure that the corresponding state of the object under test is consistent in the first candidate image and the second original image, resulting in accurate depth information, correct 3D reconstruction, and good 3D reconstruction performance. High-precision, high-stability, high-speed 3D imaging is achieved in continuous motion scenes without stopping the movement of the object under test. This effectively reduces the impact of synchronization errors and motion artifacts on depth calculation, improves detection speed and stability in industrial automation and online inspection, and enhances imaging accuracy without significantly increasing hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122760656A_ABST
    Figure CN122760656A_ABST
Patent Text Reader

Abstract

The application provides an image processing method, device and equipment, the method comprising: acquiring a first original image collected by a first camera and a second original image collected by a second camera; wherein the first original image and the second original image both comprise a to-be-measured object; determining a spatial offset of the to-be-measured object in a first time interval based on external motion prior information, the first time interval being a time interval between a second collection time and a first collection time; performing motion error compensation on the first original image based on the spatial offset to obtain a first candidate image, the first candidate image and the second original image being used to determine depth information of the to-be-measured object. Through the technical scheme of the application, high-precision, high-stability and high-speed 3D imaging is realized in a continuous motion scene without stopping the motion of the to-be-measured object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular to an image processing method, apparatus and device. Background Technology

[0002] In the fields of industrial automation and online inspection, 3D imaging technology is widely used in scenarios such as dimensional measurement, appearance inspection, and positioning guidance. Taking multi-view cameras as an example, a multi-view camera can include multiple cameras, such as a first camera and a second camera. The first camera can acquire a first original image of the object under test, and the second camera can acquire a second original image of the object under test. Based on this, the depth information of the object under test can be determined based on the first and second original images, and then a 3D reconstruction of the object can be performed.

[0003] If the acquisition time of the first original image differs from that of the second original image, then when the object under test is moving at high speed, even if the time interval between acquisitions is small, the displacement of the object under test within that time interval will be large. This will lead to inconsistencies in the corresponding states of the object under test in the first and second original images, disrupting multi-view geometric consistency. Consequently, accurate depth information of the object under test cannot be obtained, resulting in abnormal 3D reconstruction results and poor 3D reconstruction quality. Summary of the Invention

[0004] This application provides an image processing method, the method comprising: Acquire a first raw image captured by a first camera and a second raw image captured by a second camera; wherein both the first raw image and the second raw image include the object to be measured; The spatial offset of the object under test is determined based on prior information about external motion, where the first time interval is the time interval between the second acquisition time and the first acquisition time. Based on the spatial offset, motion error compensation is performed on the first original image to obtain a first candidate image. The first candidate image and the second original image are used to determine the depth information of the object under test.

[0005] This application provides an image processing apparatus, the apparatus comprising: The acquisition module is used to acquire a first raw image captured by a first camera and a second raw image captured by a second camera; wherein both the first raw image and the second raw image include the object to be measured; A determination module is used to determine the spatial offset of the object under test in a first time interval based on external motion prior information; wherein the first time interval is the time interval between the second acquisition time and the first acquisition time; a processing module is used to perform motion error compensation on the first original image based on the spatial offset to obtain a first candidate image; wherein the first candidate image and the second original image are used to determine the depth information of the object under test.

[0006] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the image processing method of the example above in this application.

[0007] This application provides a computer program product, which includes a computer program that, when executed by a processor, implements the image processing method executorified above in this application.

[0008] This application provides a machine-readable storage medium storing machine-executable instructions that can be executed by a processor; wherein the processor is configured to execute the machine-executable instructions to implement the image processing method executable in the above example of this application.

[0009] As can be seen from the above technical solutions, in this embodiment, the spatial offset of the object under test in the first time interval can be determined based on prior external motion information. Motion error compensation is then performed on the first original image based on this spatial offset to obtain a first candidate image. The depth information of the object under test is then determined based on the first candidate image and the second original image. Thus, even if the acquisition time of the first original image differs from that of the second original image, and the object under test is a high-speed moving object, motion error compensation can ensure that the corresponding state of the object under test is consistent in the first candidate image and the second original image, resulting in accurate depth information, correct 3D reconstruction, and good 3D reconstruction performance. High-precision, high-stability, high-speed 3D imaging is achieved in continuous motion scenes without stopping the movement of the object under test. This effectively reduces the impact of synchronization errors and motion artifacts on depth calculation, improves detection speed and stability in industrial automation and online inspection, and enhances imaging accuracy without significantly increasing hardware costs. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating an image processing method according to one embodiment of this application; Figure 2 This is a flowchart illustrating an image processing method according to one embodiment of this application; Figure 3AThis is a schematic diagram of a conveyor belt with periodically specific markings in one embodiment of this application; Figure 3B This is a schematic diagram illustrating the triggering of image acquisition in one embodiment of this application; Figure 3C This is a schematic diagram of motion error compensation in one embodiment of this application; Figure 3D This is a schematic diagram of motion artifact compensation in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of an image processing apparatus according to one embodiment of this application; Figure 5 This is a hardware structure diagram of an electronic device according to one embodiment of this application. Detailed Implementation

[0011] This application proposes an image processing method that can be applied to electronic devices such as personal computers, laptops, smart terminals, servers, IoT devices, and cloud devices, without limitation. See also Figure 1 The diagram shown is a flowchart of the image processing method, which may include: Step 101: Acquire a first original image captured by a first camera and a second original image captured by a second camera. The first acquisition time of the first original image is different from the second acquisition time of the second original image. Both the first original image and the second original image include the object to be measured.

[0012] Step 102: Determine the spatial offset of the object to be measured in the first time interval based on the prior information of external motion. The first time interval can be the time interval between the second acquisition time and the first acquisition time.

[0013] Step 103: Perform motion error compensation on the first original image based on the spatial offset to obtain a first candidate image; wherein, the first candidate image and the second original image are used to determine the depth information of the object to be measured.

[0014] For example, if there is one second camera and at least one first camera, for each first camera, motion error compensation can be performed on the first original image captured by the first camera based on spatial offset to obtain a first candidate image. Alternatively, if there is one first camera and at least one second camera, motion error compensation can be performed on the second original image captured by the second camera based on spatial offset to obtain a second candidate image. For ease of description, we will take obtaining a first candidate image by performing motion error compensation on the first original image based on spatial offset as an example.

[0015] For example, the external motion prior information may include the motion velocity of the object under test. Determining the spatial offset of the object under test in a first time interval based on the external motion prior information may include, but is not limited to: determining the first time interval based on the motion velocity; or, when the first camera acquires a first original image, determining a first acquisition time for the first original image, and when the second camera acquires a second original image, determining a second acquisition time for the second original image, and determining the first time interval based on the first and second acquisition times. Then, the spatial offset can be determined based on the motion velocity and the first time interval.

[0016] For example, both the first original image and the second original image include specific markers. Determining the first time interval based on the motion speed may include, but is not limited to: determining the first pixel position corresponding to the specific marker in the first original image and the second pixel position corresponding to the specific marker in the second original image; determining the actual position difference between the second pixel position and the first pixel position; determining the position change based on the actual position difference and the calibrated theoretical position difference; and determining the first time interval based on the motion speed and the position change.

[0017] For example, the external motion prior information may include the distance between two adjacent objects to be measured and a second time interval between the first signal and the second signal; the motion speed may be determined based on the distance and the second time interval. Based on this, when the preceding object to be measured moves into the field of view of the first and second cameras, the first signal is used to trigger the first and second cameras to acquire images; when the object to be measured moves into the field of view of the first and second cameras, the second signal is used to trigger the first camera to acquire a first raw image and the second signal is used to trigger the second camera to acquire a second raw image.

[0018] For example, after performing motion error compensation on the first original image based on the spatial offset to obtain a first candidate image, a first offset of the object under test can be determined based on external motion prior information during the first exposure duration. Motion artifact compensation is then performed on the first candidate image based on the first offset to obtain a first target image; wherein, the first exposure duration is the exposure duration of the first camera. A second offset of the object under test is determined based on external motion prior information during the second exposure duration. Motion artifact compensation is then performed on the second original image based on the second offset to obtain a second target image; wherein, the second exposure duration is the exposure duration of the second camera. Depth information of the object under test is determined based on the first target image and the second target image.

[0019] For example, the first target image is obtained by performing motion artifact compensation on the first candidate image based on the first offset, which may include, but is not limited to: determining the artifact region based on the gradient value of each pixel position in the first candidate image, or determining the artifact region based on the frequency value of each pixel position in the first candidate image; wherein, the region composed of pixel positions with gradient values ​​less than a preset first threshold is regarded as the artifact region; wherein, the region composed of pixel positions with frequency values ​​less than a preset second threshold is regarded as the artifact region; and the first target image is obtained by performing motion artifact compensation on the artifact region based on the first offset.

[0020] For example, the second target image is obtained by performing motion artifact compensation on the second original image based on the second offset, which may include, but is not limited to: determining the artifact region based on the gradient value of each pixel position in the second original image, or determining the artifact region based on the frequency value of each pixel position in the second original image; wherein, the region composed of pixel positions with gradient values ​​less than a preset first threshold is regarded as the artifact region; wherein, the region composed of pixel positions with frequency values ​​less than a preset second threshold is regarded as the artifact region; and the second target image is obtained by performing motion artifact compensation on the artifact region based on the second offset.

[0021] For example, after acquiring the first original image captured by the first camera and the second original image captured by the second camera, the external motion prior information, the first original image and the second original image can be input into the trained depth estimation model. Then, the depth information of the object to be measured can be determined by the depth estimation model based on the external motion prior information, the first original image and the second original image.

[0022] For example, the first camera and the second camera can be combined to form a multi-view camera; or, the first camera and the second camera can be combined to form a structured light 3D camera; or, the first camera and the second camera can be combined to form a speckle 3D camera.

[0023] As can be seen from the above technical solutions, in this embodiment, the spatial offset of the object under test in the first time interval can be determined based on prior external motion information. Motion error compensation is then performed on the first original image based on this spatial offset to obtain a first candidate image. The depth information of the object under test is then determined based on the first candidate image and the second original image. Thus, even if the acquisition time of the first original image differs from that of the second original image, and the object under test is a high-speed moving object, motion error compensation can ensure that the corresponding state of the object under test is consistent in the first candidate image and the second original image, resulting in accurate depth information, correct 3D reconstruction, and good 3D reconstruction performance. High-precision, high-stability, high-speed 3D imaging is achieved in continuous motion scenes without stopping the movement of the object under test. This effectively reduces the impact of synchronization errors and motion artifacts on depth calculation, improves detection speed and stability in industrial automation and online inspection, and enhances imaging accuracy without significantly increasing hardware costs.

[0024] The technical solutions described above in the embodiments of this application will be explained below in conjunction with specific application scenarios.

[0025] In the fields of industrial automation and online inspection, 3D imaging technology is widely used in scenarios such as dimensional measurement, appearance inspection, and positioning guidance. A first original image of the object under test can be acquired using a first camera, and a second original image can be acquired using a second camera. Based on these two original images, the depth information of the object under test is determined, and then a 3D reconstruction of the object is performed.

[0026] For example, the object to be tested can be a stationary object, a low-speed moving object, or a high-speed moving object. The following example will be a high-speed moving object.

[0027] If the acquisition time of the first original image differs from that of the second original image (due to non-negligible differences in exposure time or trigger time between different cameras), then even if the time interval between acquisitions is small, the displacement of the object under test within that time interval will be large. This will result in inconsistencies in the corresponding states of the object under test in the first and second original images, disrupting multi-view geometric consistency. Consequently, accurate depth information of the object under test cannot be obtained, leading to poor 3D reconstruction results.

[0028] Furthermore, if the object under test is controlled to stop moving, i.e., a high-speed motion-stop-shoot method is used, an additional stopping mechanism or timing control device is required, which not only reduces detection efficiency but also increases system complexity and cost. In addition, if the motion problem is mitigated by increasing the frame rate, shortening the exposure time, or enhancing the brightness of the light source, it leads to a decrease in the signal-to-noise ratio, an increase in power consumption, and a significant increase in system cost.

[0029] In response to the above findings, this application proposes a high-speed 3D imaging scheme based on external motion prior information triggering and combined with spatiotemporal compensation. This scheme can achieve high-precision and high-stability high-speed 3D imaging without stopping the motion of the object under test, and can be applied to industrial continuous motion scenarios.

[0030] For example, in the fields of industrial automation and online inspection, such as dimensional measurement, appearance inspection, and positioning guidance, the depth information of the object under test can be determined based on prior information about external motion. This achieves high-precision, high-stability, high-speed 3D imaging in continuous motion scenarios without relying on ultra-high frame rate cameras, effectively solving the problems of low efficiency and high cost caused by reliance on stationary shooting.

[0031] This application proposes an image processing method that can be applied to electronic devices, such as personal computers, laptops, smart terminals, servers, IoT devices, and cloud devices. In this embodiment, images are acquired using at least two cameras, referred to as the first camera and the second camera. The number of first cameras is at least one, and the number of second cameras is at least one.

[0032] For example, a first camera and a second camera can form a multi-view camera system. This means that a multi-view camera system is deployed in the scene of an image processing method. A multi-view camera system includes at least two cameras, referred to as the first camera and the second camera. Multi-view cameras can provide richer perspective and depth information.

[0033] A multi-view camera is an imaging system consisting of two or more cameras that have been spatially calibrated. Multi-view cameras acquire depth information by using the geometric relationships between images from multiple perspectives.

[0034] For example, a first camera and a second camera can form a structured light 3D camera. That is, a structured light 3D camera is deployed in the scene of the image processing method. The structured light 3D camera includes at least two cameras, referred to as the first camera and the second camera. The structured light 3D camera can be a 3D measurement device, consisting of a laser and at least two cameras (cameras). The laser projects a laser beam onto the surface of the object to be measured, and the cameras take pictures of the object to obtain the original image reflected by the object.

[0035] After obtaining the original image, a 3D reconstruction of the object under test can be performed based on the original image to obtain a 3D reconstructed image of the object under test. For structured light 3D cameras, the laser beam projected by the laser (structured light laser) can be a structured light beam, and the original image can include a structured light laser pattern.

[0036] For example, a first camera and a second camera can form a speckle 3D camera. That is, a speckle 3D camera is deployed in the scene of the image processing method. The speckle 3D camera includes at least two cameras, referred to as the first camera and the second camera. The speckle 3D camera can be a 3D measurement device, consisting of a laser and at least two cameras (cameras). The laser projects a laser beam onto the surface of the object to be measured, and the cameras take pictures of the object to obtain the original image reflected by the object.

[0037] After obtaining the original image, a 3D reconstruction of the object under test can be performed based on the original image to obtain a 3D reconstructed image of the object under test. For speckle 3D cameras, the laser beam projected by the laser (speckle laser) can be a speckle beam, and the original image can include a speckle laser pattern.

[0038] Of course, the above are just a few examples of 3D measurement devices (which can also be called 3D equipment). 3D measurement devices can include structured light 3D cameras, speckle 3D cameras, etc., and are not limited thereto. 3D measurement devices can be built based on the principle of laser triangulation (such as systems based on the principle of laser active illumination triangulation). 3D measurement devices can include binocular passive imaging systems (which utilize some triangular relationships for binocular distance determination). 3D measurement devices can include AI-generated multi-view cameras trained using models. That is, the 3D measurement device does not specify the principle of laser triangulation; it only needs to be able to achieve 3D measurement functions based on lasers. 3D measurement devices include collaborative measurement systems of lasers and cameras (at least two cameras). The 3D measurement devices in the embodiments of this application cover all measurement devices based on laser triangulation.

[0039] In this embodiment, if the image processing method is applied to a 3D measuring device (i.e., the electronic device can be the 3D measuring device itself), the 3D measuring device can acquire original images (such as a first original image and a second original image) and implement the image processing method based on the original images. Alternatively, if the image processing method is applied to a management device of the 3D measuring device (such as a personal computer, laptop, smart terminal, server, IoT device, cloud device, etc.), i.e., the electronic device can be the management device of the 3D measuring device, the 3D measuring device can acquire original images (such as a first original image and a second original image; the original images can be laser images or other types of images, not limited to laser images) and send the original images to the management device, which then implements the image processing method based on the original images.

[0040] In summary, for scenarios involving 3D measurement equipment, this image processing method can be applied to electronic devices, which can be either the 3D measurement equipment itself or a management device for it. For scenarios involving multi-view cameras (i.e., where lasers are not required), this image processing method can also be applied to electronic devices, which can be management devices for the multi-view cameras. For ease of description, this embodiment uses the example of an electronic device being a management device, meaning the image processing method can be applied to either 3D measurement equipment or a management device for multi-view cameras.

[0041] In scenarios involving 3D measurement equipment or multi-view cameras, at least two cameras are involved. These at least two cameras are referred to as the first camera and the second camera. There is at least one first camera and at least one second camera. For ease of description, we will take one first camera and one second camera as an example.

[0042] See Figure 2 The diagram shown is a flowchart of an image processing method, which may include: Step 201: The management device acquires prior information about external motion.

[0043] For example, external motion prior information refers to prior information provided by external devices or the environment, without relying on the multi-view camera's own image calculations, used to characterize the motion characteristics of the object under test or the supporting mechanism (i.e., the supporting mechanism of the object under test). External motion prior information may include, but is not limited to, the motion velocity and motion direction of the object under test (stable or predictable motion direction information). External motion prior information may also include displacement, periodic characteristics, time stamps, or any combination thereof. External motion prior information is predictable or repeatable in time and in space.

[0044] During the continuous movement of the object under test (the measured object) along a conveyor belt, rotary table, or other motion mechanism, prior information that characterizes the macroscopic motion properties of the measured object can be obtained, which can be denoted as external motion prior information. For example, since the motion direction of the conveyor belt, rotary table, or other motion mechanism is known, this motion direction can be taken as the motion direction of the measured object.

[0045] If the object to be measured is moving at a constant speed, and the speed of the conveyor belt, turntable, or other moving mechanism is known, then that speed can be taken as the speed of the object to be measured.

[0046] In one possible implementation, assuming the object under test is moving at a uniform speed or a non-uniform speed, the speed of the motion mechanism is not taken as the speed of the object under test. In order to obtain a more accurate speed, the external motion prior information may also include the distance between two adjacent objects under test and the second time interval between the first signal and the second signal. The speed of the object under test is determined based on the distance and the second time interval. For example, the quotient of the distance and the second time interval can be taken as the speed of the object under test.

[0047] For example, when the preceding object to be measured (i.e., the current object to be measured, such as object a1) moves into the field of view of the first and second cameras, the management device can send a first signal to the first and second cameras. The first signal is used to trigger the first and second cameras to acquire an image (i.e., an image of object a2). When object a1 moves into the field of view of the first and second cameras, the management device can send a second signal to the first and second cameras. The second signal is used to trigger the first and second cameras to acquire an image (i.e., an image of object a1).

[0048] For example, when a test object is placed on a motion mechanism (such as a conveyor belt, a rotary table, or other motion mechanism), the distance between two adjacent test objects (such as test object a2 and test object a1) can be known. Therefore, the management device can know the distance between test object a2 and test object a1.

[0049] For example, when the object under test a2 moves into the field of view of the first camera and the second camera, the management device sends a first signal to the first camera and the second camera, and the management device can know the time of transmission of the first signal. When the object under test a1 moves into the field of view of the first camera and the second camera, the management device sends a second signal to the first camera and the second camera, and the management device can know the time of transmission of the second signal.

[0050] Obviously, the time interval between the transmission time of the second signal and the transmission time of the first signal, that is, the time interval between the first signal and the second signal, can be denoted as the second time interval.

[0051] In summary, the management device can obtain the distance between two adjacent objects to be measured and the second time interval between the first signal and the second signal, and then determine the speed of the objects to be measured based on the distance and the second time interval. Thus, the external motion prior information can include the distance between two adjacent objects to be measured, the second time interval between the first signal and the second signal, and the speed of the objects to be measured.

[0052] For example, an encoder can be pre-deployed and connected to the mechanical structure of a motion mechanism (such as a conveyor belt, rotary table, or other motion mechanism). For instance, the encoder can be connected to the mechanical structure of a conveyor belt. During the movement of the motion mechanism, due to its connection to the mechanical structure, the encoder can determine the position of the motion mechanism. For instance, when the motion mechanism moves to a certain position 1 (which can be pre-configured), the encoder detects that a measured object (such as measured object a2) has moved into the field of view of the first and second cameras, and sends an encoder signal to the management device. Upon receiving the encoder signal, the management device sends a first signal to the first and second cameras. When the motion mechanism moves to a certain position 2, the encoder detects that a new measured object (such as measured object a1) has moved into the field of view of the first and second cameras, and sends an encoder signal to the management device. Upon receiving the encoder signal, the management device sends a second signal to the first and second cameras, and so on.

[0053] For example, a photoelectric switch (also called a photoelectric sensor) or a proximity sensor can be pre-deployed; taking a photoelectric sensor as an example, when the object to be measured moves into the field of view of the first and second cameras, the photoelectric sensor can detect the object; when the object to be measured does not move into the field of view of the first and second cameras, the photoelectric sensor cannot detect the object.

[0054] For example, when the object to be measured, a2, moves into the field of view of the first and second cameras, the photoelectric sensor detects a2 and sends a trigger signal (also called a photoelectric sensor signal) to the management device. Upon receiving this trigger signal, the management device sends a first signal to the first and second cameras. Similarly, when the object to be measured, a1, moves into the field of view of the first and second cameras, the photoelectric sensor detects a1 and sends a trigger signal to the management device. Upon receiving this trigger signal, the management device sends a second signal to the first and second cameras, and so on.

[0055] In addition to encoders or photoelectric sensors, other sensors that can reflect the motion state of the object under test can be deployed. The sensor detects that the object under test has moved into the field of view of the first and second cameras, sends a sensor signal to the management device, triggers the management device to send a first signal and a second signal to the first and second cameras, and then obtains the second time interval between the first signal and the second signal.

[0056] In one possible implementation, the external motion prior information may also include specific markers (specific markers may also be called visual feature markers or conveyor belt visual markers). The specific markers may be any markers deployed on the surface of the motion mechanism (such as a conveyor belt, a rotary table, or other motion mechanism, etc.) (such as specific markers set on the surface of the conveyor belt), and the location of the specific markers can be found from the image.

[0057] For example, see Figure 3A The diagram shows a conveyor belt with periodic specific markings. The surface of the conveyor belt is provided with periodic specific markings. The specific markings appear as a repeating structure along a fixed direction in the camera's field of view. The direction of motion of the object under test can be determined by these specific markings. By combining the spacing and time information of the specific markings, the speed or displacement of the object under test can also be estimated.

[0058] Step 202: The management device acquires a first original image captured by the first camera and a second original image captured by the second camera. The first acquisition time of the first original image is different from the second acquisition time of the second original image. Both the first original image and the second original image include the object to be measured.

[0059] For example, in a multi-camera scenario, the first camera can acquire a first raw image of the object under test and send it to the management device. The second camera can acquire a second raw image of the object under test and send it to the management device.

[0060] For scenarios involving structured light 3D cameras or speckle 3D cameras, a laser beam is projected onto the surface of the object under test using a laser. A first camera then captures an image of the object, obtaining a first raw image (such as a laser image or other type of image) reflected by the object. This first raw image is then sent to a management device. A second camera then captures an image of the object, obtaining a second raw image (such as a laser image) reflected by the object. This second camera is also sent to the management device.

[0061] For example, when the object under test moves into the field of view of the first and second cameras, the management device sends a second signal to both cameras (the first signal is triggered by the preceding object under test). Upon receiving the second signal, the first camera can acquire a first raw image of the object under test and send it to the management device. Upon receiving the second signal, the second camera can acquire a second raw image of the object under test and send it to the management device.

[0062] In summary, the second signal is used to trigger the first camera to acquire the first original image of the object under test, and the second signal is used to trigger the second camera to acquire the second original image of the object under test.

[0063] For example, the first camera can periodically acquire first raw images, rather than triggering the acquisition of first raw images upon receiving a signal from the management device. If the object under test moves into the field of view of the first and second cameras, the first raw image includes the object under test, and the first camera sends the first raw image to the management device. If the object under test does not move into the field of view of the first and second cameras, the first raw image does not include the object under test, and the first camera may or may not send the first raw image to the management device.

[0064] The second camera can periodically acquire second raw images, rather than triggering acquisition upon receiving a signal from the management device. If the object under test moves into the field of view of both the first and second cameras, the second raw image includes the object, and the second camera sends the second raw image to the management device. If the object under test does not move into the field of view of either the first or second camera, the second raw image does not include the object, and the second camera may or may not send the second raw image to the management device.

[0065] For example, during the acquisition process, differences in hardware response, exposure control, or communication delays can lead to time discrepancies in the effective imaging times of images from different cameras (i.e., the first camera and the second camera), resulting in synchronization time errors. Consequently, the first acquisition time of the first raw image differs from the second acquisition time of the second raw image. For instance, the first exposure start time of the first raw image may differ from the second exposure start time of the second raw image, or the first exposure end time of the first raw image may differ from the second exposure end time of the second raw image.

[0066] In summary, the management device can obtain a first original image and a second original image. Both the first and second original images include the object to be measured, and the first acquisition time of the first original image is different from the second acquisition time of the second original image. When the first camera acquires the first original image, the management device determines and records the first acquisition time of the first original image. When the second camera acquires the second original image, the management device determines and records the second acquisition time of the second original image.

[0067] See Figure 3B The diagram shown illustrates the triggering of image acquisition. The object to be measured can also be referred to as the workpiece to be measured. Figure 3B The following explanation uses the workpiece under test as an example. Specific markings can also be called visual feature markings (or simply visual features) or conveyor belt visual markings. Figure 3B The following explanation uses visual features as an example.

[0068] When the object under test moves into the preset detection area (meaning the object under test has moved into the field of view of the first and second cameras), the management device sends a second signal to the first and second cameras, triggering the first camera to acquire a first raw image of the object under test and triggering the second camera to acquire a second raw image of the object under test. The management device can record the first acquisition time of the first raw image and the second acquisition time of the second raw image. Alternatively, when a specific marker (such as a specific marker deployed in front of the object under test) moves into the preset detection area (meaning the object under test has moved into the field of view of the first and second cameras), the management device sends a second signal to the first and second cameras, triggering the first camera to acquire a first raw image of the object under test and triggering the second camera to acquire a second raw image of the object under test.

[0069] Triggered image acquisition is a mechanism used to control, coordinate, or mark the timing of image acquisition by a multi-view camera. Triggered image acquisition includes, but is not limited to, controlling the start time of exposure, exposure window, image readout, frame selection, or timestamp marking. The triggering method can be hardware triggering, software triggering, or a combination of hardware and software triggering.

[0070] Based on prior information about external motion, trigger decision logic can be executed to control or mark the image acquisition timing of multi-camera systems. Trigger conditions may include, but are not limited to, at least one of the following: the object under test moves to a preset detection area; a specific marker (such as a specific marker deployed in front of the object under test) moves to a preset detection area; or the motion speed is within a preset stable range (e.g., in a scenario where the object under test moves at a constant speed, acquisition can be triggered periodically, i.e., the first and second cameras can periodically acquire raw images).

[0071] When triggering image acquisition, it is not limited to all cameras starting exposure completely simultaneously. Instead, it is used to establish a temporal correlation between multi-view images (such as the first original image and the second original image) and external motion priors. Triggering image acquisition can be used to control exposure, frame selection, timestamp marking, or acquisition window alignment.

[0072] Step 203: The management device determines the first time interval between the second acquisition time and the first acquisition time.

[0073] For example, when the first camera acquires a first raw image, the management device can record the first acquisition time of the first raw image; when the second camera acquires a second raw image, the management device can record the second acquisition time of the second raw image. Based on this, the management device can determine a first time interval, i.e., the difference between the second acquisition time and the first acquisition time, based on the first acquisition time and the second acquisition time.

[0074] For example, the first acquisition time could be the first exposure start time of the first original image, and the second acquisition time could be the second exposure start time of the second original image, with the difference between the first and second exposure start times used as the first time interval. Alternatively, the first acquisition time could be the first exposure end time of the first original image, and the second acquisition time could be the second exposure end time of the second original image, with the difference between the first and second exposure end times used as the first time interval.

[0075] For example, the prior information of external motion includes the motion speed of the object under test, and the management device can determine the first time interval between the second acquisition time and the first acquisition time based on the motion speed of the object under test.

[0076] The first original image may include a specific marker, and the management device determines the first pixel position corresponding to the specific marker in the first original image. For example, the management device extracts the visual features of each pixel position in the first original image. If the visual features of a certain pixel position match the stored reference visual features (i.e., the visual features of the specific marker, which can be pre-stored) (e.g., the similarity between the two is greater than a threshold), then that pixel position is taken as the first pixel position corresponding to the specific marker in the first original image.

[0077] The second original image may include a specific marker, and the management device determines the second pixel position corresponding to that specific marker in the second original image. For example, the management device extracts the visual features of each pixel position in the second original image, and if the visual features of a certain pixel position match the stored reference visual features, then that pixel position is taken as the second pixel position corresponding to that specific marker in the second original image.

[0078] The management device can determine the actual position difference between the second pixel position and the first pixel position, denoted as (x1, y1), where x1 is the difference between the horizontal coordinate of the second pixel position and the horizontal coordinate of the first pixel position, and y1 is the difference between the vertical coordinate of the second pixel position and the vertical coordinate of the first pixel position.

[0079] The management equipment can also acquire the calibrated theoretical position difference. For example, since there is a certain distance between the first camera and the second camera, during the imaging process, for a certain point in physical space, the pixel position of that point in the image of the first camera is different from the pixel position of that point in the image of the second camera. The difference between these two pixel positions is the theoretical position difference, which can be pre-calibrated.

[0080] For example, a calibration object is deployed in physical space. Image 1 of the calibration object is captured by a first camera, and image 2 of the calibration object is captured by a second camera. The pixel position 1 of the calibration object in image 1 is determined, and the pixel position 2 of the calibration object in image 2 is determined. The position difference between pixel position 2 and pixel position 1 can be a theoretical position difference, which can be pre-calibrated and denoted as (x2, y2). x2 can be the difference between the x-coordinate of pixel position 2 and the x-coordinate of pixel position 1, and y2 can be the difference between the y-coordinate of pixel position 2 and the y-coordinate of pixel position 1.

[0081] The management equipment determines the change in position based on the difference between the actual position and the theoretical position. For example, the distance between the difference between the actual position and the theoretical position (such as Euclidean distance) can be determined as the change in position.

[0082] Then, the management device determines a first time interval between the second acquisition time and the first acquisition time based on the motion speed of the object under test and the change in position. For example, if the change in position is a change in the image coordinate system, it can be converted into a spatial change in the world coordinate system, such as based on the transformation relationship between the image coordinate system and the world coordinate system. Then, the quotient between the spatial change and the motion speed can be used as the first time interval.

[0083] Step 204: The management device determines the spatial offset of the object under test in the first time interval based on the external motion prior information. For example, the external motion prior information includes the motion speed of the object under test. The spatial offset can be determined based on the motion speed of the object under test and the first time interval. For example, the product of the motion speed and the first time interval can be used as the spatial offset.

[0084] For example, the motion speed can be decomposed into motion speed 1 in the X-axis direction and motion speed 2 in the Y-axis direction. The spatial offset can include spatial offset 1 in the X-axis direction and spatial offset 2 in the Y-axis direction. The product of motion speed 1 and the first time interval can be used as spatial offset 1, and the product of motion speed 2 and the first time interval can be used as spatial offset 2.

[0085] Obviously, if the object to be measured moves along the X-axis, then the velocity 2 is 0 and the spatial offset 2 is 0. If the object to be measured moves along the Y-axis, then the velocity 1 is 0 and the spatial offset 1 is 0.

[0086] Step 205: The management device performs motion error compensation on the first original image based on the spatial offset to obtain a first candidate image; or, performs motion error compensation on the second original image based on the spatial offset to obtain a second candidate image; or, performs motion error compensation on the first original image based on the spatial offset to obtain a first candidate image, and performs motion error compensation on the second original image to obtain a second candidate image.

[0087] For example, spatial offset is an offset in the world coordinate system. Converting spatial offset to pixel offset in the image coordinate system, such as based on the transformation relationship between image and world coordinate systems, allows us to convert spatial offset to pixel offset, denoted as pixel offset 1 and pixel offset 2. Pixel offset 1 corresponds to spatial offset 1, and pixel offset 2 corresponds to spatial offset 2. If the object being measured moves along the X-axis, pixel offset 2 is 0; if the object moves along the Y-axis, pixel offset 1 is 0.

[0088] Then, the external motion prior information includes the motion direction of the object under test, and motion error compensation can be performed on the first original image and / or the second original image based on the pixel offset and the motion direction.

[0089] For example, for any pixel position P1 in the first original image, the corresponding pixel position P2 in the first candidate image is the pixel offset 1 (which can be determined as positive or negative based on the motion direction of the object being measured), and the pixel offset 2 (which can also be determined as positive or negative based on the motion direction of the object being measured). Therefore, the pixel value of pixel position P2 is the same as the pixel value of pixel position P1. Clearly, after performing the above processing on each pixel position in the first original image, the motion error compensated first candidate image can be obtained.

[0090] For example, for any pixel position P3 in the second original image, the corresponding pixel position P4 in the second candidate image is equal to the pixel offset 1 in the X-axis direction and the pixel offset 2 in the Y-axis direction. Therefore, the pixel value of pixel position P4 is the same as the pixel value of pixel position P3. After performing the above processing on each pixel position in the second original image, the motion error compensated second candidate image can be obtained.

[0091] For example, the process by which a management device performs motion error compensation on a first or second original image based on spatial offset can be called spatiotemporal compensation or synchronization error compensation. Motion error compensation refers to the process of mapping synchronization time errors or imaging differences caused by continuous motion to spatial or temporal offsets in the image domain, feature domain, or depth calculation process, based on camera calibration relationships and prior information about external motion, and then correcting, aligning, or compensating for these offsets. For instance, based on prior information such as motion direction and speed, the time synchronization error caused by triggering, exposure, or readout between multiple cameras can be modeled as a spatial offset along a known motion direction, and the image can be processed for error compensation based on this spatial offset.

[0092] In the above process, the management device can determine the spatial offset of the object under test in the first time interval based on the motion speed and the first time interval, and perform motion error compensation on the first original image or the second original image based on the spatial offset. In addition to the above motion error compensation method, the management device can also remap the first original image or the second original image based on a preset motion model (such as a rigid translation model, an affine model, or a combination thereof). For example, the first original image can be rigidly translated or affinely transformed to obtain a first candidate image, or the second original image can be rigidly translated or affinely transformed to obtain a second candidate image.

[0093] Alternatively, the management device can align images from different viewpoints to a unified reference time or coordinate system based on triggering features. For example, the management device can ensure that the first acquisition time of the first original image is the same as the second acquisition time of the second original image, such as controlling the first camera and the second camera to receive the second signal at the same time, and then controlling the first camera and the second camera to acquire images at the same time.

[0094] Alternatively, error compensation can be performed in the image domain, feature domain, or depth calculation process. Error compensation in the image domain means that the management device performs motion error compensation on the first or second original image based on spatial offset. Error compensation in the feature domain means that the management device extracts features from the first original image to obtain first image features, and then performs motion error compensation on the first image features based on spatial offset; the compensation process is similar to that for the first original image. Alternatively, the management device extracts features from the second original image to obtain second image features, and then performs motion error compensation on the second image features based on spatial offset. Error compensation during depth calculation means that the management device inputs the first original image (or second original image) and external motion prior information into a deep learning model, and then performs motion error compensation on the first original image (or second original image) through the deep learning model; there are no restrictions on this compensation process.

[0095] See Figure 3CThe diagram illustrates motion error compensation. The image in the upper left corner represents the first original image captured by the first camera, and the image in the upper right corner represents the second original image captured by the second camera. A time synchronization error exists between the second and first original images, caused by differences in image acquisition time. The image in the lower left corner represents the first original image captured by the first camera, and the image in the lower right corner represents the second candidate image after motion error compensation (obtained by performing motion error compensation on the second original image). The time synchronization error between the second candidate image and the first original image is eliminated.

[0096] In summary, without motion error compensation, images acquired by different cameras correspond to the spatial state of the object under test at different time sections. By introducing external motion prior information and performing motion error compensation processing, multi-view images can correspond to the same reference time or the same spatial state.

[0097] Step 206: The management device acquires the first target image and the second target image.

[0098] In one possible implementation, if motion error compensation is performed on the first original image to obtain a first candidate image, then the first candidate image is used as the first target image, and the second original image is used as the second target image. Alternatively, if motion error compensation is performed on the second original image to obtain a second candidate image, then the second candidate image is used as the second target image, and the first original image is used as the first target image.

[0099] In one possible implementation, if motion error compensation is performed on the first original image to obtain a first candidate image, then motion artifact compensation is performed on the first candidate image to obtain a first target image, and motion artifact compensation is performed on the second original image to obtain a second target image. Alternatively, if motion error compensation is performed on the second original image to obtain a second candidate image, then motion artifact compensation is performed on the second candidate image to obtain a second target image, and motion artifact compensation is performed on the first original image to obtain the first target image.

[0100] When performing motion artifact compensation on the first candidate image to obtain the first target image, a first offset of the object under test during the first exposure time can be determined based on external motion prior information. The first target image is then obtained by performing motion artifact compensation on the first candidate image based on this first offset. Furthermore, when performing motion artifact compensation on the first original image to obtain the first target image, the first target image can be obtained by performing motion artifact compensation on the first original image based on this first offset. The first exposure time is the exposure time of the first camera.

[0101] When performing motion artifact compensation on the second original image to obtain the second target image, a second offset of the object under test can be determined based on external motion prior information during the second exposure duration. The second target image is then obtained by performing motion artifact compensation on the second original image based on this second offset. Furthermore, when performing motion artifact compensation on the second candidate image to obtain the second target image, the second target image can be obtained by performing motion artifact compensation on the second candidate image based on the second offset. The second exposure duration is the exposure duration of the second camera. The second exposure duration of the second camera can be the same as or different from the first exposure duration of the first camera.

[0102] For example, in the process of determining the first offset of the object under test based on the prior information of external motion, the prior information of external motion may include the motion speed of the object under test. The first offset (i.e., the first spatial offset) can be determined based on the motion speed of the object under test and the first exposure time. For example, the product of the motion speed and the first exposure time can be used as the first offset.

[0103] Then, the first offset in the world coordinate system is converted into the first pixel offset in the image coordinate system. Based on the first pixel offset and the motion direction of the object to be tested, motion artifact compensation is performed on the first candidate image (or the first original image) to obtain the first target image. The compensation process is similar to step 205.

[0104] For example, in the process of determining the second offset of the object under test during the second exposure time based on external motion prior information, the external motion prior information may include the motion velocity of the object under test. The second offset (i.e., the second spatial offset) can be determined based on the motion velocity of the object under test and the second exposure time. For instance, the product of the motion velocity and the second exposure time can be used as the second offset. Obviously, if the second exposure time is the same as the first exposure time, then the second offset is the same as the first offset.

[0105] Then, the second offset in the world coordinate system is converted into the second pixel offset in the image coordinate system. Based on the second pixel offset and the motion direction of the object to be tested, motion artifact compensation is performed on the second original image (or the second candidate image) to obtain the second target image. The compensation process is similar to step 205.

[0106] For example, based on synchronization error compensation, and combining external motion prior information (such as motion direction and speed) with image features (such as feature matching), the management device can also perform motion artifact compensation on multi-view images (such as the first candidate image and the second original image) to further reduce image blurring caused by rapid movement of the object under test or the camera. For instance, motion artifacts refer to image blurring, distortion, or information mismatch caused by exposure time or imaging sequence when the object under test is in continuous motion.

[0107] Motion artifact compensation can include one or more of the following methods, which can be performed sequentially or in parallel: Image remapping is based on motion models (such as rigid translation models, affine models, or combinations thereof). For example, a first target image is obtained by rigidly translating or affine the first candidate image, and a second target image is obtained by rigidly translating or affine the second original image. There are no restrictions on this remapping process.

[0108] Adaptive motion model processing: For non-rigid motion scenes, more complex motion models (such as thin-plate spline models or elastic models) are used for image remapping. For example, the first candidate image is remapped using a thin-plate spline model or elastic model to obtain the first target image, and the second original image is remapped using a thin-plate spline model or elastic model to obtain the second target image. There are no restrictions on this process.

[0109] Deblurring based on motion direction and displacement: Based on the motion direction and displacement (i.e., the first offset and the second offset), the image is blurred (e.g., PSF, deconvolution, Wiener filtering, etc.) to restore a clear image. For example, motion artifact compensation is performed on the first candidate image based on the first offset to obtain the first target image. When compensating for motion artifacts on the first candidate image, methods such as PSF deconvolution and Wiener filtering can also be used. Motion artifact compensation is then performed on the second original image based on the second offset to obtain the second target image.

[0110] Multi-frame or multi-view information fusion: Aligned multi-view images are fused to average motion noise. For example, a first candidate image and a second original image can be fused (e.g., weighted fusion), where the weighting coefficient of the first candidate image is greater than that of the second original image, to obtain a first target image. Alternatively, a first candidate image and a second original image can be fused (e.g., weighted fusion), where the weighting coefficient of the first candidate image is less than that of the second original image, to obtain a second target image.

[0111] Artifact Detection and Local Compensation: Residual artifact regions are identified using image quality metrics (such as gradient analysis or frequency domain detection), and targeted compensation is applied to these regions. For example, artifact regions can be determined based on the gradient values ​​of each pixel location within the first candidate image (calculated based on the first candidate image). Regions with gradient values ​​less than a preset first threshold (configured according to actual needs) are considered artifact regions. Alternatively, artifact regions can be determined based on the frequency values ​​of each pixel location within the first candidate image (calculated based on the first candidate image). Regions with frequency values ​​less than a preset second threshold (configured according to actual needs) are considered artifact regions. Based on this, motion artifact compensation is performed on the artifact regions within the first candidate image using a first offset to obtain the first target image, instead of performing motion artifact compensation on the entire first candidate image.

[0112] Artifact regions are determined based on the gradient values ​​of each pixel location within the second original image (calculated based on the second original image). Regions consisting of pixel locations with gradient values ​​less than a preset first threshold are considered artifact regions. Alternatively, artifact regions are determined based on the frequency values ​​of each pixel location within the second original image (calculated based on the second original image). Regions consisting of pixel locations with frequency values ​​less than a preset second threshold are considered artifact regions. Based on these methods, motion artifact compensation is performed on the artifact regions within the second original image using a second offset to obtain the second target image.

[0113] See Figure 3D The diagram illustrates motion artifact compensation. The left image represents the first candidate image, which is an image captured by the first camera under high-speed motion (including motion blur). Alternatively, the left image represents the second original image, which is an image captured by the second camera under high-speed motion. The right image represents either the first or second target image, which is the image after motion artifact compensation.

[0114] Based on prior information about external motion (such as motion speed v, motion direction, etc.), and combined with the exposure time Δt (such as the first exposure time or the second exposure time), the displacement Δt (such as the first displacement or the second displacement) can be calculated, and then motion artifact compensation can be performed on the first candidate image (the second original image).

[0115] Step 207: The management device determines the depth information of the object to be measured based on the first target image and the second target image. The depth information of the object to be measured can be a depth map or three-dimensional point cloud data of the object to be measured.

[0116] For example, after obtaining the first target image and the second target image, a depth imaging process can be completed. Depth imaging refers to generating a depth map or 3D point cloud data reflecting the spatial structure of the object under test through image processing or model calculation. For instance, based on the first and second target images, a stereo matching algorithm (such as block matching, semi-global matching, or graph cut algorithm) can be used to determine the depth information of the object under test. Alternatively, the first and second target images can be input into a depth estimation model (such as a convolutional neural network, stereo matching network, or multi-view stereo network) to determine the depth information of the object under test. Or, external motion prior information, the first and second target images can be input into a depth estimation model (a depth estimation model capable of fusing multimodal conditions), where the external motion prior information (as an additional constraint) can assist the depth estimation model in determining the depth information of the object under test. Of course, the above are just a few examples; the goal is simply to determine the depth information of the object under test.

[0117] In summary, based on multi-view images (such as the first target image and the second target image) after motion compensation and spatial alignment, multi-view depth calculation can be performed to output the depth information of the object under test.

[0118] In one possible implementation, after acquiring the first raw image captured by the first camera and the second raw image captured by the second camera, the external motion prior information, the first raw image, and the second raw image can be input into a trained depth estimation model. The depth information of the object under test can be determined by the depth estimation model based on the external motion prior information, the first raw image, and the second raw image.

[0119] For example, a depth estimation model can be pre-trained, with no restrictions on its network structure, training method, or parameter configuration. The input data for the depth estimation model can be external motion prior information, raw images captured by the first camera, and raw images captured by the second camera.

[0120] Based on this, the management device can input external motion prior information (such as motion direction, motion speed, etc.), the first original image (i.e. the original image before compensation) and the second original image (i.e. the original image before compensation) into the depth estimation model. The depth estimation model is an AI model that can integrate multimodal conditions.

[0121] After obtaining the first and second original images, the depth estimation model can determine the depth information of the object under test based on these images. When determining the depth information of the object, the depth estimation model considers external motion prior information and performs depth estimation based on this prior information. That is, external motion prior information participates in the model calculation as input features, intermediate feature modulation conditions, or depth estimation constraints. This allows the depth estimation model to explicitly or implicitly consider the influence of motion factors on the imaging results during the depth estimation process, resulting in output results that incorporate motion information and achieve more robust depth estimation.

[0122] As can be seen from the above technical solutions, in this embodiment, external motion prior information is introduced as the basis for acquisition triggering. Based on the external motion prior information, the time synchronization error and motion artifacts between cameras are compensated, achieving high-precision and high-stability high-speed 3D imaging in continuous motion scenes without stopping the movement of the object under test. This effectively reduces the impact of synchronization error and motion artifacts on depth calculation, improves the detection speed and stability in industrial automation and online inspection, and improves imaging accuracy without significantly increasing hardware costs. A multi-view depth imaging method based on external motion prior information is proposed. Under the premise that the multi-view cameras have completed spatial calibration, image acquisition is controlled by an external trigger signal. Using the motion direction prior, the time synchronization error between multi-view cameras caused by triggering, exposure, or readout is modeled and mapped as a calculable image spatial offset along the known motion direction. This offset is then compensated and aligned in the spatial domain, overcoming the problem of inconsistent viewpoints caused by continuous object movement. The synchronization error problem of multi-view cameras in continuous motion scenes is solved by external motion prior information, improving the dynamic scene measurement accuracy of multi-view cameras and achieving high-precision, high-speed, and low-cost depth estimation.

[0123] Based on the same concept as the above method, this application proposes an image processing apparatus, see [link to previous application]. Figure 4 The diagram shown is a structural schematic of the image processing device, which may include: The acquisition module 41 is used to acquire a first original image captured by a first camera and a second original image captured by a second camera; wherein the first original image and the second original image include the object to be measured; The determining module 42 is used to determine the spatial offset of the object under test in a first time interval based on external motion prior information; wherein, the first time interval is the time interval between the second acquisition time and the first acquisition time; the processing module 43 is used to perform motion error compensation on the first original image based on the spatial offset to obtain a first candidate image; wherein, the first candidate image and the second original image are used to determine the depth information of the object under test.

[0124] For example, the external motion prior information includes the motion speed of the object under test. When the determining module 42 determines the spatial offset of the object under test in the first time interval based on the external motion prior information, it is specifically used to: determine the first time interval based on the motion speed; or, when the first camera acquires the first original image, determine the first acquisition time of the first original image, when the second camera acquires the second original image, determine the second acquisition time of the second original image, and determine the first time interval based on the first acquisition time and the second acquisition time. The spatial offset is determined based on the motion speed and the first time interval.

[0125] For example, the first original image and the second original image include specific markers. When the determining module 42 determines the first time interval based on the motion speed, it is specifically used to: determine the first pixel position corresponding to the specific marker in the first original image and the second pixel position corresponding to the specific marker in the second original image; determine the actual position difference between the second pixel position and the first pixel position; determine the position change amount based on the actual position difference and the calibrated theoretical position difference; and determine the first time interval based on the motion speed and the position change amount.

[0126] For example, the external motion prior information includes the distance between two adjacent test objects and a second time interval between the first signal and the second signal; the motion speed is determined based on the distance and the second time interval; wherein, when the preceding test object moves into the field of view of the first camera and the second camera, the first signal is used to trigger the first camera and the second camera to acquire images; when the test object moves into the field of view of the first camera and the second camera, the second signal is used to trigger the first camera to acquire the first original image and the second signal is used to trigger the second camera to acquire the second original image.

[0127] For example, the determining module 42 is further configured to: determine a first offset of the object under test based on the external motion prior information during a first exposure duration; perform motion artifact compensation on the first candidate image based on the first offset to obtain a first target image; wherein the first exposure duration is the exposure duration of the first camera; determine a second offset of the object under test during a second exposure duration based on the external motion prior information; perform motion artifact compensation on the second original image based on the second offset to obtain a second target image; wherein the second exposure duration is the exposure duration of the second camera; and determine the depth information of the object under test based on the first target image and the second target image.

[0128] For example, when the determining module 42 performs motion artifact compensation on the first candidate image based on the first offset to obtain the first target image, it is specifically used to: determine the artifact region based on the gradient value of each pixel position in the first candidate image, or, determine the artifact region based on the frequency value of each pixel position in the first candidate image; the region composed of pixel positions with gradient values ​​less than a preset first threshold is taken as the artifact region; the region composed of pixel positions with frequency values ​​less than a preset second threshold is taken as the artifact region; and the first target image is obtained by performing motion artifact compensation on the artifact region based on the first offset.

[0129] For example, the processing module 43 is further configured to, after obtaining the first original image and the second original image, input the external motion prior information, the first original image and the second original image into a trained depth estimation model, and determine the depth information of the object to be measured based on the external motion prior information, the first original image and the second original image through the depth estimation model.

[0130] For example, the first camera and the second camera form a multi-view camera; or, the first camera and the second camera form a structured light 3D camera; or, the first camera and the second camera form a speckle 3D camera.

[0131] Based on the same concept as the above method, this application proposes an electronic device, see [link to previous application]. Figure 5 As shown, the electronic device includes a processor 51 and a machine-readable storage medium 52, the machine-readable storage medium 52 storing machine-executable instructions that can be executed by the processor 51; the processor 51 is used to execute the machine-executable instructions to implement the image processing method disclosed in the above example of this application.

[0132] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the image processing method disclosed in the above examples of this application.

[0133] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0134] Based on the same concept as the methods described above, this application also provides a computer program product, which may include a computer program. When executed by a processor, the computer program implements the image processing method disclosed in the examples above.

[0135] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An image processing method, characterized in that, The method includes: Acquire a first raw image captured by a first camera and a second raw image captured by a second camera; wherein both the first raw image and the second raw image include the object to be measured; The spatial offset of the object under test is determined based on prior information about external motion in the first time interval, where the first time interval is the time interval between the second acquisition time and the first acquisition time. Based on the spatial offset, motion error compensation is performed on the first original image to obtain a first candidate image. The first candidate image and the second original image are used to determine the depth information of the object under test.

2. The method according to claim 1, characterized in that, The external motion prior information includes the motion velocity of the object under test, and determining the spatial offset of the object under test in the first time interval based on the external motion prior information includes: The first time interval is determined based on the motion speed; or, when the first camera acquires the first original image, a first acquisition time of the first original image is determined, and when the second camera acquires the second original image, a second acquisition time of the second original image is determined, and the first time interval is determined based on the first acquisition time and the second acquisition time. The spatial offset is determined based on the motion speed and the first time interval.

3. The method according to claim 2, characterized in that, Both the first original image and the second original image include specific markers. Determining the first time interval based on the motion speed includes: Determine the first pixel position of the specific marker in the first original image and the second pixel position of the specific marker in the second original image; Determine the actual position difference between the second pixel position and the first pixel position, and determine the position change based on the actual position difference and the calibrated theoretical position difference; The first time interval is determined based on the motion speed and the change in position.

4. The method according to claim 2, characterized in that, The external motion prior information includes the distance between two adjacent objects to be measured and the second time interval between the first signal and the second signal; the motion speed is determined based on the distance and the second time interval. When the preceding object to be tested moves into the field of view of the first camera and the second camera, the first signal is used to trigger the first camera and the second camera to acquire images. When the object under test moves into the field of view of the first camera and the second camera, the second signal is used to trigger the first camera to acquire the first original image and the second signal is used to trigger the second camera to acquire the second original image.

5. The method according to any one of claims 1-4, characterized in that, After obtaining the first candidate image by performing motion error compensation on the first original image based on the spatial offset, the method further includes: Based on the external motion prior information, the first offset of the object under test is determined at the first exposure time, and motion artifact compensation is performed on the first candidate image based on the first offset to obtain the first target image; wherein, the first exposure time is the exposure time of the first camera; Based on the external motion prior information, the second offset of the object under test is determined at the second exposure time, and motion artifact compensation is performed on the second original image based on the second offset to obtain the second target image; wherein, the second exposure time is the exposure time of the second camera; The depth information of the object under test is determined based on the first target image and the second target image.

6. The method according to claim 5, characterized in that, The step of performing motion artifact compensation on the first candidate image based on the first offset to obtain the first target image includes: The artifact region is determined based on the gradient value of each pixel position in the first candidate image, or based on the frequency value of each pixel position in the first candidate image; wherein, the region composed of pixel positions with gradient values ​​less than a preset first threshold is the artifact region; wherein, the region composed of pixel positions with frequency values ​​less than a preset second threshold is the artifact region. Based on the first offset, motion artifact compensation is performed on the artifact region to obtain the first target image.

7. The method according to any one of claims 1-4, characterized in that, After acquiring the first raw image captured by the first camera and the second raw image captured by the second camera, the method further includes: The external motion prior information, the first original image, and the second original image are input into a trained depth estimation model, and the depth estimation model determines the depth information of the object under test based on the external motion prior information, the first original image, and the second original image.

8. The method according to any one of claims 1-4, characterized in that, The first camera and the second camera form a multi-view camera; or, The first camera and the second camera together form a structured light 3D camera; or, The first camera and the second camera together form a speckle 3D camera.

9. An image processing apparatus, characterized in that, The device includes: An acquisition module is used to acquire a first raw image captured by a first camera and a second raw image captured by a second camera; wherein the first raw image and the second raw image include the object to be measured; The determination module is used to determine the spatial offset of the object under test in a first time interval based on external motion prior information; wherein, the first time interval is the time interval between the second acquisition time and the first acquisition time; The processing module is used to perform motion error compensation on the first original image based on the spatial offset to obtain a first candidate image; wherein the first candidate image and the second original image are used to determine the depth information of the object to be measured.

10. The apparatus according to claim 9, characterized in that, The external motion prior information includes the motion velocity of the object under test. When determining the spatial offset of the object under test based on the external motion prior information, the determining module is specifically used for: The first time interval is determined based on the motion speed; or, when the first camera acquires the first original image, a first acquisition time of the first original image is determined, and when the second camera acquires the second original image, a second acquisition time of the second original image is determined, and the first time interval is determined based on the first acquisition time and the second acquisition time. The spatial offset is determined based on the motion speed and the first time interval; Alternatively, the first original image and the second original image include specific markers, and when the determining module determines the first time interval based on the motion speed, it is specifically used to: determine the first pixel position corresponding to the specific marker in the first original image and the second pixel position corresponding to the specific marker in the second original image; Determine the actual position difference between the second pixel position and the first pixel position, and determine the position change based on the actual position difference and the calibrated theoretical position difference; The first time interval is determined based on the motion speed and the position change. Alternatively, the external motion prior information includes the distance between two adjacent objects to be tested and a second time interval between the first signal and the second signal; the motion speed is determined based on the distance and the second time interval; wherein, when the preceding object to be tested moves into the field of view of the first camera and the second camera, the first signal is used to trigger the first camera and the second camera to acquire images; when the object to be tested moves into the field of view of the first camera and the second camera, the second signal is used to trigger the first camera to acquire the first original image and the second signal is used to trigger the second camera to acquire the second original image; Alternatively, the determining module is further configured to: determine a first offset of the object under test based on the external motion prior information during a first exposure duration; perform motion artifact compensation on the first candidate image based on the first offset to obtain a first target image; wherein the first exposure duration is the exposure duration of the first camera; determine a second offset of the object under test during a second exposure duration based on the external motion prior information; perform motion artifact compensation on the second original image based on the second offset to obtain a second target image; wherein the second exposure duration is the exposure duration of the second camera; and determine the depth information of the object under test based on the first target image and the second target image. Alternatively, when the determining module performs motion artifact compensation on the first candidate image based on the first offset to obtain the first target image, it is specifically used to: determine the artifact region based on the gradient value of each pixel position in the first candidate image, or determine the artifact region based on the frequency value of each pixel position in the first candidate image; wherein, the region composed of pixel positions with gradient values ​​less than a preset first threshold is regarded as the artifact region; the region composed of pixel positions with frequency values ​​less than a preset second threshold is regarded as the artifact region; and the first target image is obtained by performing motion artifact compensation on the artifact region based on the first offset. Alternatively, the processing module is further configured to, after obtaining the first original image and the second original image, input the external motion prior information, the first original image and the second original image into a trained depth estimation model, and determine the depth information of the object to be measured based on the external motion prior information, the first original image and the second original image through the depth estimation model; Alternatively, the first camera and the second camera can form a multi-view camera; or, the first camera and the second camera can form a structured light 3D camera; or, the first camera and the second camera can form a speckle 3D camera.

11. An electronic device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of any one of claims 1-8.