Photoelectric target image acquisition method

Through the photoelectric target image acquisition method, the photographing is performed using a drone or a camera, and the image data is processed through self-developed image processing and labeling software, the problems of sparse images and poor quality are solved, and high-quality image data sets are provided, suitable for the training of machine learning models.

CN120014218APending Publication Date: 2025-05-16HENAN ZHONGZHI KELIAN EQUIP TECH DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411938823.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the image sources are poor and the quality is uneven, making it difficult to support high-quality machine learning model training.

Method used

A method of photoelectric object image acquisition includes model provision, shooting, image processing and labeling. Shooting through drones or cameras, self-developed image processing and labeling software is used to realize video frame extraction, image compression, cropping and labeling functions, and supports a variety of model and scene requirements.

Benefits of technology

It solves the problems of sparse images, poor quality and labeling, provides high-quality image datasets, suitable for machine learning scenarios, and improves the recognition and classification performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a photoelectric target image acquisition method. The method comprises the following steps: obtaining a demand through communication with a client; an unmanned aerial vehicle or a camera is adopted for shooting, the unmanned aerial vehicle is used for shooting a model with a large model and a large scene in an external field or an open space, and the camera is used for shooting a model with a small model and a small scene in a sand table indoors; the functions of video frame extraction, picture compression, picture cutting and picture labeling are achieved through image processing and image labeling software, c # language construction software is adopted for video frame extraction, an application program ffmpeg program is called, one frame extraction is carried out on a video in one second, and picture compression and picture cutting are to call the application program ffmepg to compress or cut a picture; the picture annotation is saved as an XML file in a PASCALVOC format, the YOLO format is also supported, the problems of rare images, poor quality and annotation are solved, and the method is suitable for machine learning scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image acquisition and annotation, and in particular relates to a photoelectric target image acquisition method. Background Art

[0002] With the continuous development of machine learning technology, its application background is becoming more and more extensive. At present, machine learning has been widely used in image recognition, speech recognition, natural language processing, recommendation systems and other fields. In terms of image recognition, high-quality images can provide more detailed information, which is conducive to the model learning more accurate features, thereby improving the recognition and classification performance of the model. Machine learning models need to extract useful features from images for learning and prediction. Therefore, the features in the image are crucial to the performance of the model. In practical applications, it may be necessary to enhance the features in the image through preprocessing steps, such as image enhancement and denoising. However, the sources of images are scarce and the quality is uneven, which makes it difficult to support machine learning. Summary of the invention

[0003] In order to overcome the above problems, the present invention provides a photoelectric target image acquisition method, which covers models, shooting, image processing and image annotation, solves the problems of image scarcity, poor quality and labeling, and is suitable for machine learning scenarios.

[0004] The technical solution adopted by the present invention is:

[0005] A method for acquiring an optoelectronic target image comprises the following steps:

[0006] Step 1: Provide a model and obtain requirements through communication with customers;

[0007] Step 2: Shooting. Use a drone or camera to shoot. For models with large models and large scenes, use a drone to shoot outdoors or in an open space. For models with small models and small scenes, place them on a sand table and use a camera to shoot indoors.

[0008] Step 3: Image processing and annotation. Through image processing and image annotation software, the functions of video frame extraction, image compression, image cropping and image annotation are realized. The video frame extraction software is built using the C# language, and the application ffmpeg program is called to extract one frame of the video per second. Image compression and image cropping are to call the application ffmepg to compress or crop the image; image annotation is saved as an XML file in PASCAL VOC format (the format used by ImageNet), and YOLO format is also supported.

[0009] Among them, the needs in step one include production, purchase or rental options. For models that are rare in the market or the market cannot meet customer needs, production is selected, and for models that are used multiple times, purchase is selected, and for models that are used less frequently, rental is selected.

[0010] Among them, the models in step one are 1:1, 1:5 and 1:24.

[0011] Among them, when the drone shoots at an oblique downward angle in step 2, the height to horizontal distance is 0.1:12-0.5:12, and the target posture shooting of 0 degree - 360 degrees is completed. Specifically, the drone is equipped with an RTK navigation system. When RTK FIX is in effect, the horizontal position accuracy is 1cm+1ppm, and the vertical position accuracy is 1.5cm+1ppm. The horizontal distance and the height from the ground of the drone to the shooting target are calculated, and the shooting is performed according to the condition that the height to horizontal distance meets the 0.1:12-0.5:12 condition.

[0012] Among them, the visible light resolution of the image acquisition in step 2 is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, and the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels. The physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted by the following formula:

[0013] DPI = resolution product / physical size (length × width);

[0014] The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14×0.14cm-0.7×0.7cm;

[0015] The infrared image resolution is 640×512, and when the physical size is 5.4×4.3cm, the image size at this size is between 0.14×0.14cm-0.7×0.7cm;

[0016] The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%.

[0017] Among them, when shooting the sand table in step 2, a tape measure is used to measure the size of the position in the house, and the shooting is carried out according to the condition that the height to horizontal distance meets the ratio of 0.1:12-0.5:12. The position of the sand table camera cannot be moved when shooting the target, and the posture image of the target is collected at 0 degrees to 360 degrees by changing the height, position, and pitch angle of the sand table.

[0018] Among them, the image acquisition visible light resolution is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, and the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels. The physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted using the following formula:

[0019] DPI = resolution product / physical size (length × width);

[0020] The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14cm×0.14cm-0.7cm×0.7cm;

[0021] The infrared image resolution is 640×512, and when the physical size is 5.4cm×4.3cm, the image size at this size is between 0.14cm×0.14cm-0.7cm×0.7cm;

[0022] The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%.

[0023] The advantages of the present invention are as follows:

[0024] The present invention includes providing models, shooting, image processing and image annotation. The model sizes are divided into 1:1, 1:5, 1:24, etc., which can meet various needs. The model types are also diverse, and some large models that are difficult to find can also be manufactured; the self-developed image processing and image annotation software can realize functions such as video frame extraction, picture compression, cropping and picture annotation. The video frame extraction adopts C# language to build software, and calls the application ffmpeg program to extract one frame of the video per second. The picture processing also calls the application ffmepg to compress or crop the picture, and functions can be added at any time according to needs; in addition to normal shooting, the camera is also equipped with an infrared function to meet diverse needs; it solves the problems of scarce images, poor quality and labeling, and is suitable for machine learning scenarios. DETAILED DESCRIPTION

[0025] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0026] A method for acquiring an optoelectronic target image comprises the following steps:

[0027] Step 1: Provide a model and obtain the customer's needs through communication. The customer can choose to make, buy or rent the model. For models that are rare in the market or difficult to meet the customer's needs, choose to make them. For models that are used multiple times, choose to buy them. For models that are used less frequently, choose to rent them. The models include 1:1, 1:5 and 1:24.

[0028] Step 2: Shooting, using a drone or camera;

[0029] For models with large models and scenes as large as 1:1 and 1:5, drones are used for outdoor or open space photography. When the drone is shooting at an oblique downward angle, the height to horizontal distance is 0.1:12-0.5:12, and the target posture shooting of 0-360 degrees is completed. Specifically, the drone is equipped with an RTK navigation system. When RTK FIX is in effect, the horizontal position accuracy is 1cm+1ppm, and the vertical position accuracy is 1.5cm+1ppm. The horizontal distance and height above the ground of the drone are calculated. The shooting angles are as follows when the height to horizontal distance meets the condition of 0.1:12-0.5:12:

[0030]

[0031]

[0032] When shooting a target, the drone's ascent altitude is adjusted between the minimum and maximum altitudes according to the horizontal distance between the drone and the target, and the target's posture image is collected from 0 degrees to 360 degrees;

[0033] The image acquisition visible light resolution is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, and the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels. The physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted using the following formula:

[0034] DPI = resolution product / physical size (length × width);

[0035] The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14×0.14cm-0.7×0.7cm;

[0036] The infrared image resolution is 640×512, and when the physical size is 5.4×4.3cm, the image size at this size is between 0.14×0.14cm-0.7×0.7cm;

[0037] The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%;

[0038] Table of target sizes for photos with different resolutions:

[0039]

[0040]

[0041] For smaller models and scenes, such as 1:24 models, place them on a sand table and shoot them indoors with a camera. When shooting on the sand table, use a tape measure to measure the size of the location in the room. If the height to horizontal distance meets the condition of 0.1:12-0.5:12, the shooting angles are as follows:

[0042]

[0043] The position of the sandbox camera cannot be moved when shooting the target. The target's posture image is collected from 0 to 360 degrees by changing the height, position, and pitch angle of the sandbox.

[0044] The image acquisition visible light resolution is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, and the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels. The physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted using the following formula:

[0045] DPI = resolution product / physical size (length × width);

[0046] The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14cm×0.14cm-0.7cm×0.7cm;

[0047] The infrared image resolution is 640×512, and when the physical size is 5.4cm×4.3cm, the image size at this size is between 0.14cm×0.14cm-0.7cm×0.7cm;

[0048] The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%;

[0049] Table of target sizes for photos with different resolutions:

[0050]

[0051] Step 3: Image processing and annotation. Through image processing and image annotation software, the functions of video frame extraction, image compression, image cropping and image annotation are realized. The video frame extraction software is built using the C# language, and the application ffmpeg program is called to extract one frame of the video per second. Image compression and image cropping are to call the application ffmepg to compress or crop the image; image annotation is saved as an XML file in PASCAL VOC format (the format used by ImageNet), and YOLO format is also supported.

[0052] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A photoelectric target image acquisition method, characterized in that: The following steps are involved: Step 1: Provide a model and obtain requirements through communication with customers; Step 2: Shooting. Use a drone or camera to shoot. For models with large models and large scenes, use a drone to shoot outdoors or in an open space. For models with small models and small scenes, place them on a sand table and use a camera to shoot indoors. Step 3: Image processing and annotation. Through image processing and image annotation software, the functions of video frame extraction, image compression, image cropping and image annotation are realized. The video frame extraction software is built using the C# language, and the application ffmpeg program is called to extract one frame of the video per second. Image compression and image cropping are to call the application ffmepg to compress or crop the image; image annotations are saved as XML files in PASCAL VOC format, and YOLO format is also supported.

2. The photoelectric target image acquisition method according to claim 1, characterized in that: In the step 1, there are options of making, purchasing or leasing. For models that are rare in the market or that are difficult to meet customer needs, production is selected, for models that are used multiple times, purchase is selected, and for models that are used less frequently, leasing is selected.

3. The photoelectric target image acquisition method according to claim 1, characterized in that: The models in step 1 are 1:1, 1:5 and 1:

24.

4. The photoelectric target image acquisition method according to claim 1, characterized in that: When the drone shoots at an oblique downward angle in the step 2, the target posture shooting of 0 degree to 360 degrees is completed according to the height to horizontal distance ratio of 0.1:12-0.5:

12. Specifically, the drone is equipped with an RTK navigation system. The horizontal position accuracy is 1cm+1ppm and the vertical position accuracy is 1.5cm+1ppm in RTK FIX. The horizontal distance and the height from the ground of the drone to the shooting target are calculated, and the shooting is performed according to the condition that the height to horizontal distance meets the 0.1:12-0.5:12 condition.

5. The photoelectric target image acquisition method according to claim 4, characterized in that: In the step 2, the image acquisition visible light resolution is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels, and the physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted by the following formula: DPI = resolution product / physical size (length × width); The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14×0.14cm-0.7×0.7cm; The infrared image resolution is 640×512, and when the physical size is 5.4×4.3cm, the image size at this size is between 0.14×0.14cm-0.7×0.7cm; The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%.

6. The photoelectric target image acquisition method according to claim 1, characterized in that: In the step 2, when shooting the sand table, the size of the position in the house is measured by a tape measure, and the shooting is carried out according to the condition that the height to horizontal distance meets the condition of 0.1:12-0.5:

12. The position of the sand table camera cannot be moved when shooting the target, and the posture image of the target is collected at 0 degrees to 360 degrees by changing the height, position and pitch angle of the sand table.

7. The photoelectric target image acquisition method according to claim 6, characterized in that: In the step 2, the image acquisition visible light resolution is greater than or equal to 1024×768, the infrared resolution is greater than or equal to 640×512, the imaging range of the target in the image is between 16×16 pixels and 80×80 pixels, and the physical value of the target imaging range is calculated by resolution, screen size, and DPI conversion, and the display screen DPI is converted by the following formula: DPI = resolution product / physical size (length × width); The DPI value is calculated to be 300 DPI, and the physical size of a 1024×768 photo is 8.7cm×6.5cm. At this size, the image size is between 0.14cm×0.14cm-0.7cm×0.7cm; The infrared image resolution is 640×512, and when the physical size is 5.4cm×4.3cm, the image size at this size is between 0.14cm×0.14cm-0.7cm×0.7cm; The proportion of images with target line length ≤ 40 pixels is not less than 50%, and the proportion of images in the total number of images is greater than 50%.