Image acquisition system, marking method, processing method, electronic equipment and vehicle

Through the design of three camera systems combined with the spectrometer and attenuation filter, the problem of image labeling in dark environments is solved, and the synchronous acquisition and multiplexing of multimodal images in sufficient light environments is achieved, improving the efficiency and accuracy of image labeling in dark environments.

CN120455839APending Publication Date: 2025-08-08BYD CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510660691.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In dark environments, deterioration in the image quality in the prior art makes it very difficult to label images. The existing cameras have low response speeds under low light conditions and are easily affected by the environment, making it difficult to effectively acquire and label images.

Method used

Three camera systems are adopted, including the first camera, the second camera and the third camera. Through the combination of a spectrometer and attenuation filter, RGB images with sufficient light, RGB images in dark environments and event frame images are collected respectively. The optical path design and attenuation filter are used to synchronize the multi-modal images in a sufficient light environment, and then label them on the normal light image and multiplex them on other images.

Benefits of technology

After labeling images in an environment with sufficient lighting, it can be easily synchronously and multiplexed into images in dark environments, reducing the difficulty and cost of image labeling in dark environments, and improving the efficiency and accuracy of image acquisition and labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455839A_ABST
    Figure CN120455839A_ABST
Patent Text Reader

Abstract

The invention discloses an image acquisition system, a processing method, electronic equipment and a vehicle. The image acquisition system comprises a first camera, a second camera and a third camera, the first spectroscope is used for splitting the first light beam into a second light beam and a third light beam, and the second light beam enters the first camera; the second spectroscope is used for splitting the third light beam into a fourth light beam and a fifth light beam, the fourth light beam enters the second camera, and the fifth light beam enters the third camera; and the attenuation optical filter is used for performing intensity attenuation on the third light beam or the fifth light beam, so that the light beam subjected to intensity attenuation enters the third camera. According to the invention, in an environment with sufficient illumination, an illumination normal image, a dark light image and an event camera image are synchronously acquired through the light path design and the use of the attenuation optical filter. And the marking on the image with normal illumination is very easy, and the image is synchronously multiplexed to the other two images, so that the difficulty of image marking in a dark environment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and more specifically, to an image acquisition system, a labeling method, a processing method, an electronic device, and a vehicle. Background Art

[0002] Cameras are widely used in vehicle vision sensors and can capture rich visual semantic information. However, image quality deteriorates easily in dark environments, making image annotation in dark environments very difficult. Summary of the Invention

[0003] The embodiments of the present application provide an image acquisition system, an annotation method, a processing method, an electronic device, and a vehicle, which reduce the difficulty of image annotation in dark environments.

[0004] In order to achieve the above-mentioned object, according to a first aspect of the present application, an image acquisition system is provided, comprising:

[0005] a first camera, a second camera, and a third camera;

[0006] a first beam splitter, the first beam splitter being configured to split the first light beam into a second light beam and a third light beam, wherein the second light beam enters the first camera;

[0007] a second beam splitter, the second beam splitter being configured to split the third light beam into a fourth light beam and a fifth light beam, the fourth light beam entering the second camera, and the fifth light beam entering the third camera;

[0008] An attenuation filter is used to attenuate the intensity of the third light beam or the fifth light beam, so that the light beam with attenuated intensity enters the third camera.

[0009] Optionally, the attenuation filter is located between the first beam splitter and the second beam splitter, and the attenuation filter is used to attenuate the intensity of the third light beam; or, the attenuation filter is located between the second beam splitter and the third camera, and the attenuation filter is used to attenuate the intensity of the fifth light beam.

[0010] Optionally, the first camera and the third camera have the same shooting frequency.

[0011] Optionally, the first camera, the second camera and the third camera have the same resolution.

[0012] Optionally, the first camera and the third camera are RGB cameras, and the second camera is an event camera.

[0013] Optionally, the light splitting ratio of the first beam splitter and the second beam splitter is 50:50, and the light attenuation of the attenuation filter is 60% to 80%.

[0014] According to a second aspect of the present application, there is provided an image annotation method, comprising:

[0015] Acquire a first image captured by a first camera, event frame image data captured by a second camera, and a third image captured by a third camera in the image acquisition system described above, wherein the brightness of the third image is lower than that of the first image, and the event frame image data is used to record changes in scene brightness;

[0016] annotating the first image, and multiplexing the annotated result into the event frame image data and the third image;

[0017] A labeled data set is obtained according to the labeled first image, the event frame image data, and the third image.

[0018] Optionally, obtaining a second image captured by the second camera at each moment and recording brightness changes of the scene;

[0019] determining a length of a time period according to a timestamp of the first image and / or the third image;

[0020] All the second images within the time period are fitted to obtain the event frame image data, wherein the first image, the third image and the event frame image data have consistent time stamps.

[0021] Optionally, fitting all the second images within the time period to obtain the event frame image data includes:

[0022] Obtaining a change polarity value of each pixel point in the second image at each moment in response to a brightness change of the scene;

[0023] The change polarity value of each pixel point in all the second images within the time period is fitted to obtain the event frame image data.

[0024] According to a third aspect of the present application, there is provided an image processing method, comprising:

[0025] Acquire low-brightness images;

[0026] Inputting the low-brightness image into a first annotation model to obtain a first annotation result, wherein the first annotation model is used to annotate the low-brightness image;

[0027] Comparing the first annotation result with the annotation dataset obtained by the above-mentioned image annotation method to obtain a first annotation error;

[0028] The first annotation model is optimized according to the first annotation error, so as to process the low-brightness image using the optimized first annotation model to obtain a first annotation result.

[0029] Optionally, inputting the low-brightness image into a first annotation model to obtain a first annotation result includes:

[0030] After performing the feature extraction and sampling on the low-brightness image, data annotation is performed to obtain the first annotation result.

[0031] According to a fourth aspect of the present application, there is provided an image processing method, comprising:

[0032] Acquire low-light images and event frame pictures;

[0033] Inputting the low-brightness image and the event frame picture into a second annotation model to obtain a second annotation result, wherein the second annotation model is used to fuse and annotate the low-brightness image and the event frame picture;

[0034] Comparing the second annotation result with the annotation dataset obtained by the above-mentioned image annotation method to obtain a second annotation error;

[0035] The second annotation model is optimized according to the second annotation error, so as to process the low-brightness image and the event frame picture with the optimized second annotation model to obtain a second annotation result.

[0036] Optionally, inputting the low-brightness image and the event frame picture into a second annotation model to obtain a second annotation result includes:

[0037] Fusing the low-brightness image and the event frame image to obtain a fusion weight coefficient;

[0038] Extract features from the low-brightness image and the event frame image respectively to obtain a first feature and a second feature;

[0039] The second labeling result is obtained according to the fusion weight coefficient, the first feature and the second feature.

[0040] According to a fifth aspect of the present application, an electronic device is provided, including:

[0041] Memory, on which computer programs / instructions are stored;

[0042] A processor is configured to execute the computer program / instructions in the memory to implement the steps of the above-mentioned image annotation method or the above-mentioned image processing method.

[0043] According to a sixth aspect of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the image annotation method or the steps of the image processing method as described above are implemented.

[0044] According to a seventh aspect of the present application, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image annotation method or the above-mentioned image processing method.

[0045] According to an eighth aspect of the present application, a vehicle is provided, which includes the electronic device as described above, or includes a computer-readable storage medium as described above, or includes the image acquisition system as described above.

[0046] This application uses optical path design and the use of attenuation filters to simultaneously capture normal-light images, dark-light images, and event camera images. It is very easy to annotate the normal-light image and simultaneously multiplex it onto the other two images, reducing the difficulty of image annotation in dark environments.

[0047] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 A schematic diagram of the structure of an image acquisition system provided in some embodiments of the present application;

[0050] Figure 2 A schematic diagram of another image acquisition system structure provided by certain embodiments of the present application;

[0051] Figure 3 A schematic diagram of a flow chart of an image annotation method provided in certain embodiments of the present application;

[0052] Figure 4 A flowchart of another image annotation method provided in certain embodiments of the present application;

[0053] Figure 5 A flowchart of an image processing method provided in some embodiments of the present application;

[0054] Figure 6A flowchart of an image acquisition and annotation processing method provided by certain embodiments of the present application;

[0055] Figure 7 A schematic diagram of a first picture, a second picture, and a third picture provided in some embodiments of the present application;

[0056] Figure 8 Certain embodiments of the present application provide a second image fitting schematic diagram. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application. In addition, it should be understood that the specific embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application.

[0058] Conventional RGB cameras are widely used in vehicle vision sensors, capable of capturing rich visual semantic information. However, they have a slow response speed and are highly susceptible to environmental influences. For example, low light and other dark environments can significantly degrade the quality of RGB images. This deterioration in image quality in dark environments can lead to performance degradation in downstream perception algorithms, and related technologies are particularly difficult to label images in dark environments.

[0059] In order to solve the above problems, the present invention provides an image acquisition system. Figure 1 and Figure 2 Shown, including:

[0060] First camera 1, second camera 2 and third camera 3;

[0061] a first beam splitter 4 , the first beam splitter 4 is used to split the first light beam 6 into a second light beam 7 and a third light beam 8 , the second light beam 7 enters the first camera 1 ;

[0062] a second beam splitter 6 , the second beam splitter 6 is used to split the third light beam 8 into a fourth light beam 9 and a fifth light beam 10 , the fourth light beam 9 enters the second camera 2 , and the fifth light beam 10 enters the third camera 3 ;

[0063] The attenuation filter 5 is used to attenuate the intensity of the third light beam 8 or the fifth light beam 10 , so that the light beam with attenuated intensity enters the third camera 3 .

[0064] It can be understood that the first camera 1, the second camera 2, and the third camera 3 can include but are not limited to functions for driving assistance, safety monitoring, environmental perception, recording and evidence collection, etc., such as wide-angle cameras, infrared / night vision cameras, multispectral / thermal imaging cameras, etc.; the beam splitter can include but is not limited to an optical element mainly used to split a beam of light into two (or more) beams for optical path design with different paths. The beam splitting ratio of the beam splitter can include but is not limited to a 50:50 beam splitter, a 20:80 beam splitter, a 30:70 beam splitter, etc.; the attenuation filter can include but is not limited to an optical filter used to control light intensity to reduce the brightness or energy of the incident light while maintaining the spectral characteristics unchanged. The optical density of the attenuation filter can include but is not limited to a transmitted light intensity of 30% of the incident light (light attenuation 70%), 20% (light attenuation 80%), 10% (light attenuation 90%), etc.

[0065] Specifically, the image acquisition system includes a first camera 1, a second camera 2, a third camera 3, a first beam splitter 4, a second beam splitter 6, and an attenuation filter 5. The system's optical path is designed to first capture a first light beam 6 from the scene. After passing through the first beam splitter 4, the first light beam 6 is split into two, yielding a second light beam 7 and a third light beam 8. The second light beam 7 enters the first camera 1. The third light beam 8 passes through the second beam splitter 6, splitting it into two, yielding a fourth light beam 9 and a fifth light beam 10. The fourth light beam 9 enters the second camera 2, and the fifth light beam 10 enters the third camera 3. In this way, the first camera 1, the second camera 2, and the third camera 3 record different scene data for the same scene at the same time using the second light beam 7, the fourth light beam 9, and the fifth light beam 10, respectively. The attenuation filter 5 can be used to attenuate the intensity of the third light beam 8 and the fifth light beam 10, thereby making the image brightness of the third camera 3 lower than that of the first camera 1.

[0066] In a specific embodiment, if Figure 1 and Figure 2 As shown, after the first light beam of the scene with sufficient external light enters the device, it is first split into two beams by the No. 1 spectroscope. The second light beam 7 is photographed by the No. 1 RGB camera to obtain a well-lit RGB image; the third light beam 8 passes through an attenuation filter 5 to attenuate and reduce the brightness of the light, and then passes through the No. 2 spectroscope to be split into two beams again. The fifth light beam 10 is photographed by the No. 3 RGB camera to obtain an RGB image in a dark environment, and the fourth light beam 9 is photographed by the No. 2 event camera to obtain an event image frame. The resolution of these three cameras needs to be consistent, and the degree of attenuation of the light by the attenuation filter 5 can be adjusted automatically according to actual needs. The three images obtained through the data acquisition device are as follows Figure 7As shown in the figure below. The left image is an RGB image in a well-lit environment, the middle image is an RGB image in a dark environment, and the right image is an event frame image. In this way, through the optical path design and the use of attenuation filter 5, the normal-light image, the dark-light image, and the event camera data are simultaneously collected in a well-lit environment, and there is no pixel difference between the camera images. Annotation is very easy on the normal-light image and can be synchronously multiplexed on the other two images, reducing the cost of data collection and annotation in dark environments.

[0067] In certain embodiments, as Figure 1 As shown, the attenuation filter 5 is located between the first beam splitter 4 and the second beam splitter 6, and the attenuation filter 5 is used to attenuate the intensity of the third light beam 8; or, as shown Figure 2 As shown, the attenuation filter 5 is located between the second beam splitter 6 and the third camera 3 , and the attenuation filter 5 is used to attenuate the intensity of the fifth light beam 10 .

[0068] Specifically, the attenuation filter 5 can be set between the first beam splitter 4 and the second beam splitter 6. In this case, the attenuation filter 5 will attenuate the intensity of the third light beam 8, which is equivalent to attenuating the intensity of the fourth light beam 9 and the fifth light beam 10. In this way, the brightness of the image obtained by the second camera 2 and the third camera 3 can be lower than that of the image obtained by the first camera 1. The attenuation filter 5 can also be set between the second beam splitter 6 and the third camera 3. In this way, the brightness of the image obtained by the third camera 3 can be lower than that of the image obtained by the first camera 1. By using the attenuation filter 5, a simulated dark environment can be captured in an environment with sufficient light. It is very easy to mark on the normal light image and can be synchronously multiplexed on the other two images, reducing the cost of data collection and marking in dark environments.

[0069] In some embodiments, the first camera 1 and the third camera 3 have the same shooting frequency.

[0070] Specifically, the first camera 1 and the third camera 3 have the same shooting frequency, for example, 10 Hz. This consistent shooting frequency ensures strict temporal synchronization between the two sets of images, reducing matching errors caused by time skew. This synchronized shooting frequency simplifies the multi-camera image alignment and fusion process, reduces computational complexity, and thus improves the system's real-time performance and reliability.

[0071] In some embodiments, the first camera 1 , the second camera 2 , and the third camera 3 have the same resolution.

[0072] Specifically, the resolutions of the first camera 1, the second camera 2, and the third camera 3 are consistent. Consistent resolution ensures pixel-level alignment of multiple images from multiple cameras, thus avoiding blurred or distorted stitching boundaries caused by resolution differences.

[0073] In some embodiments, the first camera 1 and the third camera 3 are RGB cameras, and the second camera 2 is an event camera.

[0074] Among them, it can be understood that an RGB camera can be a camera that uses three color channels of red, green, and blue to capture color information within the visible light spectrum; an event camera can be a visual sensor that asynchronously detects brightness changes at the pixel level, with a high dynamic range and microsecond timing resolution without motion blur, and is not affected by the absolute ambient brightness.

[0075] Specifically, the first camera 1 and the third camera 3 are RGB cameras, and the second camera 2 is an event camera. By introducing event cameras, multimodal fusion can be achieved in dark environments by combining the visual semantic information of RGB cameras with the high dynamic range of event data from event cameras. This improves the performance of perception algorithms in dark environments and enhances the performance and environmental adaptability of downstream perception algorithms.

[0076] In some embodiments, the light splitting ratio of the first beam splitter 4 and the second beam splitter 6 is 50:50, and the light attenuation of the attenuation filter 5 is 60% to 80%.

[0077] Specifically, the 50:50 split ratio simplifies the system calibration process and avoids complex adjustment steps, thereby improving the device's ease of use and reliability. The filter, with an attenuation range of 60% to 80%, ensures effective attenuation of light intensity while retaining sufficient optical signal, thus meeting the light intensity requirements of subsequent systems while avoiding overload or damage.

[0078] The present application provides an image annotation method, including:

[0079] Acquire a first image captured by a first camera, event frame image data captured by a second camera, and a third image captured by a third camera in the image acquisition system described above, wherein the brightness of the third image is lower than that of the first image, and the event frame image data is used to record changes in scene brightness;

[0080] Annotate the first image, and reuse the annotation results on the event frame image data and the third image;

[0081] A labeled data set is obtained according to the labeled first image, the event frame image data, and the third image.

[0082] It can be understood that the event frame image data may include but is not limited to data information of the scene brightness changes recorded in real time by the second camera 2. The event frame image data may be data obtained by processing an event frame image, or data obtained by fusing multiple event frame images within a time period; annotation may include but is not limited to manual, semi-automatic or automatic marking of key targets, areas or features in the image to clearly describe the content, location, attributes and other information in the image; reuse may include but is not limited to directly or indirectly applying the existing annotation data (such as target boxes, segmentation masks, attribute labels, etc.) of an image to another image in some way, such as by directly copying annotations, reuse based on geometric transformations, etc.; the annotation dataset may include but is not limited to a dataset including the first image, event frame image data and the third image at different times, and their corresponding image annotation labels.

[0083] Specifically, combined Figure 1-8 As shown, the first camera 1, the second camera 2, and the third camera 3 in the image acquisition system respectively acquire a first image, event frame image data, and a third image. Due to the effect of the attenuation filter 5, the brightness of the third image is lower than that of the first image. The event frame image data is used to record real-time scene brightness changes. Next, key targets, areas, or features in the first image are annotated to obtain annotation results. These annotation results are then directly reused in the event frame image data and the third image. In a specific embodiment, because the resolution of the first camera 1, the second camera 2, and the third camera 3 is consistent, the annotation results of the first image can be directly copied and annotated into the event frame image data and the third image. Finally, the first image, event frame image data, and third image at different times, as well as their corresponding image annotation labels, are stored in a dataset to obtain an annotation dataset for use by downstream algorithms, models, etc. In this way, in an environment with sufficient lighting, through the optical path design and the use of the attenuation filter 5, the normal-light image, the dark-light image and the event camera data are collected synchronously, and there is no pixel difference between the camera images; it is very easy to mark on the normal-light image, and it can be synchronously multiplexed on the other two images, reducing the cost of data collection and marking in a dark environment.

[0084] Further, obtaining a second image captured by the second camera at each moment to record the brightness change of the scene;

[0085] determining a length of a time period based on a time stamp of the first image and / or the third image;

[0086] All second images within the time period are fitted to obtain event frame image data, wherein the first image, the third image and the event frame image data have consistent time stamps.

[0087] Among them, it can be understood that the timestamp may include but is not limited to the time information recorded in the image file, indicating that the image was taken or generated; the length of the time period may include but is not limited to the time period near a certain timestamp (the length of the time period can be specified by yourself), such as the time when the timestamp is T1, and the preset time length is t, then the time period length can be [T1-t, T1+t]; fitting may include but is not limited to fusing multiple images through methods such as simple averaging, weighted averaging, median fusion, etc. to generate a new image.

[0088] Specifically, combined Figure 6-8 As shown, the second camera 2 captures a second image in real time, recording the scene brightness changes at each moment. Because the first and third cameras 1 and 3 have the same acquisition frequency, the time period length can be determined based on the timestamp T1 of the first and / or third images, such as [T1-t, T1+t], where t is a preset time period. All second images within this time period [T1-t, T1+t] are then fitted, such as using a simple averaging method, and ultimately fused into the event frame image data. This also ensures that the timestamps of the first, third, and event frame image data are consistent. Fitting multiple images can merge clear areas across multiple frames, improving the overall image resolution and detail, while reducing storage requirements and real-time processing latency. The consistent timestamps of the first, third, and event frame image data simplify the multi-camera image alignment and fusion process, reducing computational complexity and thereby improving the real-time performance and reliability of the system.

[0089] Furthermore, all second images within the time period are fitted to obtain event frame image data, including:

[0090] Obtaining a change polarity value of each pixel in the second image at each moment in response to a change in scene brightness;

[0091] The change polarity value of each pixel point in all the second images within the time period is fitted to obtain event frame image data.

[0092] It can be understood that the change polarity value may include but is not limited to the change in light intensity of each pixel in the image. If the light intensity exceeds a certain threshold, a polarity value is assigned to the pixel (usually +1 indicates an increase in brightness, and -1 indicates a decrease in brightness).

[0093] Specifically, combined Figure 6-8As shown, after obtaining the change polarity value of each pixel in the second image at each moment in response to the scene brightness change, a fitting (e.g., simple averaging) is performed based on the change polarity values of each pixel in all second images within the time period (e.g., [T1-t, T1+t]) to obtain the event frame image data. Whether the light intensity increases or decreases is not important, only whether the light intensity changes. This reduces the amount of data and improves image processing efficiency.

[0094] In a specific embodiment, combining Figure 1-8 As shown, Figure 1 and Figure 2 It is an image acquisition system composed of an RGB camera and an event camera. After the first light beam of a scene with sufficient external light enters the device, it is first split into two beams of light by the No. 1 spectroscope. The second light beam 7 is photographed by the No. 1 RGB camera to obtain a well-lit RGB image; the third light beam 8 passes through an attenuation filter 5 to attenuate and reduce the brightness of the light, and then passes through the No. 2 spectroscope to be split into two beams of light again. The fifth light beam 10 is photographed by the No. 3 RGB camera to obtain an RGB image in a dark environment, and the fourth light beam 9 is photographed by the No. 2 event camera to obtain an event image frame. The resolution of these three cameras needs to be consistent, and the degree of attenuation of light by the attenuation filter 5 can be adjusted automatically according to actual needs. The three images obtained through this data acquisition device are as follows Figure 7 The left image is an RGB image with sufficient illumination, the middle image is an RGB image in a dark environment, and the right image is an event frame image.

[0095] Figure 6 It is a joint data collection and annotation processing method for RGB cameras and event cameras in dark environments, which mainly consists of the following processes:

[0096] (1) Start the process and execute step (2);

[0097] (2) Select a scene with sufficient daylight and proceed to step (3);

[0098] (3) Install the joint data acquisition device on the mobile vehicle and execute step (4);

[0099] (4) Set all cameras to start shooting at the same time, and the shooting frequency of RGB cameras 1 and 3 is consistent. Event camera 2 records the change of light intensity of each pixel in the picture at every moment. If the light intensity exceeds a certain threshold, a polarity is assigned to the pixel (usually +1 indicates an increase in brightness, -1 indicates a decrease in brightness). The event camera asynchronously outputs a series of 4-tuples, including the pixel coordinates of the event, the timestamp of the event, and the event polarity: m =(x m ,y m ,tm ,p m ), where x, y are pixel coordinate positions, t is timestamp information, and p is polarity value. We usually call this output event stream data and execute step (5);

[0100] (5) Perform data collection and execute step (6);

[0101] (6) Determine whether sufficient data has been obtained. If not, execute step (5). If so, execute step (7).

[0102] (7) Stop data collection and execute step (8);

[0103] (8) Obtain the timestamp of the RGB image frame, fit the event data in the time period near the timestamp (the length of the time period can be specified by yourself) to obtain the event frame image, and collect all event stream data within the length of the time period. Because the event stream data only records the pixel points with changes in light intensity, the pixel points that have not changed in the time period are set to 0, and then record the polarity changes of the pixel points at the same position, calculate the average of these polarity changes, and display them according to the polarity changes of all pixels in the time period. Set the pixels with positive polarity changes to white, the pixels with negative polarity changes to black, and the pixels with no polarity changes to gray, so as to obtain the event frame image data, as shown in Figure 1. Figure 8 shown. Figure 8 Medium red indicates a positive pixel polarity, indicating an increasing light intensity at that pixel; blue indicates a negative pixel polarity, indicating a decreasing light intensity. Overlaying the timestamps according to their polarity creates a grayscale image that eliminates the effects of polarity and reflects event frame data within a specific period. By focusing on the change in light intensity, regardless of whether it increases or decreases, we can reduce data volume and improve image processing efficiency.

[0104] (9) Data annotation is performed on the RGB image with sufficient illumination (captured by camera 1). The annotation results can be directly reused on all images (including event frame images and image data captured by camera 3) because there is no pixel deviation between images. Step (10) is performed.

[0105] (10) The process ends and three image frames and the corresponding image annotation datasets are obtained.

[0106] The present application provides an image processing method, including:

[0107] Acquire low-brightness images;

[0108] Inputting the low-brightness image into a first annotation model to obtain a first annotation result, wherein the first annotation model is used to annotate the low-brightness image;

[0109] Comparing the first annotation result with the annotation dataset obtained by the above-mentioned image annotation method to obtain a first annotation error;

[0110] The first annotation model is optimized according to the first annotation error, so as to process the low-brightness image with the optimized first annotation model to obtain a first annotation result.

[0111] Among them, it can be understood that obtaining low-light images may include but is not limited to methods such as using sensors on the vehicle, or collecting through road-side equipment and then transmitting them to the vehicle; the first annotation model may include but is not limited to a model that processes and annotates images through a combination of different algorithms, such as through convolution, pooling, target detection, target tracking and other algorithms; the first annotation result may include but is not limited to the result obtained by annotating key targets, areas or features in the image; the first annotation error may include but is not limited to the difference value obtained by comparing the first annotation result with the annotation data set in the same scene, such as information such as key targets, areas or features.

[0112] Specifically, combined Figure 4 and Figure 5 As shown, in a dark environment, a low-light image of the scene in which the current vehicle is located is first acquired. Then, using a first annotation model, key objects, regions, or features in the low-light image are annotated, resulting in a first annotation result. Due to the effects of lighting on object visibility, feature clarity, and noise, the first annotation result directly obtained from the low-light image will exhibit first annotation errors when compared with a previously annotated dataset collected under full lighting conditions. Finally, based on the obtained first annotation errors, the first annotation model is trained and optimized to continuously improve its image annotation capabilities and performance in dark environments. In this way, through optical path design and the use of a brightness attenuation device, normal-light images, dark-light images, and event camera data are simultaneously acquired in a fully lit environment, with no pixel differences between the camera images. Annotation on the normal-light image is very easy and can be simultaneously reused on the other two images, reducing the cost of data acquisition and annotation in dark environments. The annotation model is continuously updated based on the annotation errors, improving its image annotation capabilities and performance in dark environments.

[0113] Furthermore, the low-brightness image is input into the first annotation model to obtain a first annotation result, including:

[0114] After feature extraction and sampling of the low-brightness image, data annotation is performed to obtain a first annotation result.

[0115] Among them, it can be understood that feature extraction can include but is not limited to through convolution operations, etc.; adoption can include but is not limited to through pooling operations, etc.; data annotation can include but is not limited to through various perception tasks such as image classification, target detection, lane line detection, instance segmentation, etc.

[0116] Specifically, in one embodiment, combining Figure 4 and Figure 5 As shown, RGB cameras capture low-brightness images in dark environments. Due to the low brightness, it's difficult to discern the target object and key information. First, the acquired low-light RGB image is subjected to convolution and pooling operations. The convolution operation extracts image features, while the pooling operation reduces the amount of data and speeds up processing. Finally, the convolved and pooled features are fed into a traditional perception algorithm for processing, yielding the initial annotation results. Traditional perception algorithms can be used for various perception tasks, such as image classification, object detection, lane detection, and instance segmentation.

[0117] The present application provides an image processing method, including:

[0118] Acquire low-light images and event frame pictures;

[0119] Inputting the low-brightness image and the event frame picture into a second annotation model to obtain a second annotation result, wherein the second annotation model is used to fuse and annotate the low-brightness image and the event frame picture;

[0120] Comparing the second annotation result with the annotation dataset obtained by the above-mentioned image annotation method to obtain a second annotation error;

[0121] The second annotation model is optimized according to the second annotation error, so as to process the low-brightness image and the event frame picture with the optimized second annotation model to obtain a second annotation result.

[0122] Among them, it can be understood that obtaining low-light images and event frame images may include but is not limited to methods such as using sensors on the vehicle, or collecting them through road-side equipment and then transmitting them to the vehicle; the second annotation model may include but is not limited to a model that processes and annotates images through a combination of different algorithms, such as through convolution, pooling, splicing, full connection operations, target detection, target tracking and other algorithms; the second annotation result may include but is not limited to the result obtained by annotating key targets, areas or features in the image; the second annotation error may include but is not limited to the difference value obtained by comparing the second annotation result with the annotation data set in the same scene, such as information such as key targets, areas or features.

[0123] Specifically, combined Figure 3 and Figure 5As shown, in a dark environment, a low-light image and an event frame image of the scene in which the vehicle is currently located are first acquired, for example, using an RGB camera and an event camera. A second annotation model is then used to fuse and extract features from the low-light image and event frame image to produce a second annotation result. Due to the effects of illumination on object visibility, feature clarity, and noise, the second annotation result directly obtained from the low-light image and event frame image will exhibit second annotation errors when compared to a previously annotated dataset collected under sufficient illumination. Finally, the second annotation model is trained and optimized based on the obtained second annotation errors, continuously improving its image annotation capabilities and performance in dark environments. In this way, through optical path design and the use of a brightness attenuation device, normal-light images, dark-light images, and event camera data are simultaneously acquired in a well-illuminated environment, with no pixel differences between the camera images. Annotation on the normal-light image is very easy and can be simultaneously reused on the other two images, reducing the cost of data acquisition and annotation in dark environments. The annotation model is continuously updated based on the annotation errors, improving its image annotation capabilities and performance in dark environments.

[0124] Furthermore, the low-brightness image and the event frame picture are input into the second annotation model to obtain a second annotation result, including:

[0125] The low-brightness image and the event frame image are fused to obtain a fusion weight coefficient;

[0126] Extract features from the low-brightness image and the event frame image respectively to obtain the first feature and the second feature;

[0127] A second labeling result is obtained according to the fusion weight coefficient, the first feature and the second feature.

[0128] It can be understood that fusion processing may include but is not limited to the process of merging multiple images into a single image, such as image stitching, deep learning fusion and other algorithms; fusion weight coefficients may include but are not limited to vectors obtained based on the fusion between RGB images and event frame images, such as one-dimensional vectors; feature extraction may include but is not limited to convolution operations, etc.

[0129] Specifically, combined Figure 3 and Figure 5As shown, in a specific embodiment, a low-brightness image taken in a dark environment and an event frame image (taken by an event camera) are used as input, and after the two are fused, they are input to the downstream perception algorithm to perform the perception task. First, the RGB image and the event frame image are directly spliced, followed by pooling and full connection operations to obtain a one-dimensional vector, which represents the fusion weight coefficient between the RGB image and the event frame image; then the RGB image and the event frame image are convolved to obtain the first feature and the second feature. After the features of the two are multiplied by the fusion weight coefficient, they are spliced again to obtain the fused features. The fused features are convolved and can then be connected to various existing perception algorithms: such as target detectors, semantic segmenters, etc. The fusion of RGB images and event camera images improves the dark environment perception capability acquisition limit; and the data fusion step is pre-placed, so the various subsequent perception algorithms can be directly reused without modification.

[0130] The present application also provides an electronic device, including:

[0131] Memory, on which computer programs / instructions are stored;

[0132] A processor is configured to execute the computer program / instructions in the memory to implement the steps of the above-mentioned image annotation method or the above-mentioned image processing method.

[0133] The embodiments of the present application further provide a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the above-mentioned image annotation method or the steps of the above-mentioned image processing method.

[0134] The embodiments of the present application further provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image annotation method or the steps of the above-mentioned image processing method.

[0135] An embodiment of the present application further provides a vehicle, characterized in that the vehicle includes the electronic device as described above, or includes a computer-readable storage medium as described above, or includes the image acquisition system as described above.

[0136] In the description of this specification, the descriptions with reference to the terms "specifically", "further", "particularly", "may be understood to" and the like mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.

[0137] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0138] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An image acquisition system, characterized in that: include: a first camera (1), a second camera (2), and a third camera (3); a first beam splitter (4), the first beam splitter (4) being used to split the first light beam (6) into a second light beam (7) and a third light beam (8), the second light beam (7) entering the first camera (1); a second beam splitter (6), the second beam splitter (6) being used to split the third light beam (8) into a fourth light beam (9) and a fifth light beam (10), the fourth light beam (9) entering the second camera (2), and the fifth light beam (10) entering the third camera (3); An attenuation filter (5) is used to attenuate the intensity of the third light beam (8) or the fifth light beam (10), so that the light beam with attenuated intensity enters the third camera (3).

2. The system according to claim 1, wherein: The attenuation filter (5) is located between the first beam splitter (4) and the second beam splitter (6), and the attenuation filter (5) is used to attenuate the intensity of the third light beam (8); or, the attenuation filter (5) is located between the second beam splitter (6) and the third camera (3), and the attenuation filter (5) is used to attenuate the intensity of the fifth light beam (10).

3. The system according to claim 1, wherein: The shooting frequencies of the first camera (1) and the third camera (3) are consistent.

4. The system according to claim 1, wherein: The first camera (1), the second camera (2) and the third camera (3) have the same resolution.

5. The system according to any one of claims 1 to 4, characterized in that The first camera (1) and the third camera (3) are RGB cameras, and the second camera (2) is an event camera.

6. The system according to claim 1, wherein: The light splitting ratio of the first beam splitter (4) and the second beam splitter (6) is 50:50, and the light attenuation of the attenuation filter (5) is 60% to 80%.

7. An image annotation method, characterized in that: include: Acquire a first image captured by a first camera, event frame image data captured by a second camera, and a third image captured by a third camera in the image acquisition system according to any one of claims 1 to 6, wherein the brightness of the third image is lower than that of the first image, and the event frame image data is used to record changes in scene brightness; annotating the first image, and multiplexing the annotated result into the event frame image data and the third image; A labeled data set is obtained according to the labeled first image, the event frame image data, and the third image.

8. The image annotation method according to claim 7, characterized in that: include: Acquire a second image captured by the second camera at each moment and recording brightness changes of the scene; determining a length of a time period according to a timestamp of the first image and / or the third image; All the second images within the time period are fitted to obtain the event frame image data, wherein the first image, the third image and the event frame image data have consistent time stamps.

9. The image annotation method according to claim 8, characterized in that: The fitting of all the second images within the time period to obtain the event frame image data includes: Obtaining a change polarity value of each pixel point in the second image at each moment in response to a brightness change of the scene; The change polarity value of each pixel point in all the second images within the time period is fitted to obtain the event frame image data.

10. An image processing method, characterized in that: include: Acquire low-brightness images; Inputting the low-brightness image into a first annotation model to obtain a first annotation result, wherein the first annotation model is used to annotate the low-brightness image; Comparing the first annotation result with an annotation dataset obtained by the image annotation method according to any one of claims 7 to 9 to obtain a first annotation error; The first annotation model is optimized according to the first annotation error, so as to process the low-brightness image using the optimized first annotation model to obtain a first annotation result.

11. The image processing method according to claim 10, wherein: Inputting the low-brightness image into a first annotation model to obtain a first annotation result includes: After performing the feature extraction and sampling on the low-brightness image, data annotation is performed to obtain the first annotation result.

12. An image processing method, characterized in that: include: Acquire low-light images and event frame pictures; Inputting the low-brightness image and the event frame picture into a second annotation model to obtain a second annotation result, wherein the second annotation model is used to fuse and annotate the low-brightness image and the event frame picture; Comparing the second annotation result with the annotation data set obtained by the image annotation method according to any one of claims 7 to 9 to obtain a second annotation error; The second annotation model is optimized according to the second annotation error, so as to process the low-brightness image and the event frame picture with the optimized second annotation model to obtain a second annotation result.

13. The image processing method according to claim 12, wherein: The step of inputting the low-brightness image and the event frame picture into a second annotation model to obtain a second annotation result includes: Fusing the low-brightness image and the event frame image to obtain a fusion weight coefficient; Extract features from the low-brightness image and the event frame image respectively to obtain a first feature and a second feature; The second labeling result is obtained according to the fusion weight coefficient, the first feature and the second feature.

14. An electronic device, characterized in that: include: Memory, on which computer programs / instructions are stored; A processor is configured to execute the computer program / instructions in the memory to implement the steps of the image annotation method according to any one of claims 7 to 9 or the steps of the image processing method according to any one of claims 10 to 13.

15. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the image annotation method according to any one of claims 7 to 9 or the steps of the image processing method according to any one of claims 10 to 13 are implemented.

16. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the image annotation method according to any one of claims 7 to 9 or the steps of the image processing method according to any one of claims 10 to 13 are implemented.

17. A vehicle, characterized in that: The vehicle includes the electronic device according to claim 14, or includes a computer-readable storage medium according to claim 15, or includes the image acquisition system according to any one of claims 1-6.

Citation Information

Cited By

  • Extremely low illumination imaging method, system and device based on event-image fusion

    CN120894244A

  • Event-image fusion-based ultra-low light imaging methods, systems, and devices

    CN120894244B