Efficient Image Data Intelligent Enhancement and Transcoding Method and System Adapted to Object Detection
By introducing intelligent enhanced transcoding methods and systems for efficient image data in the field of object detection, the problem of performance bottlenecks in large-scale data set processing is solved, automatic update of comments and multi-format conversion is realized, and data processing efficiency and accuracy are significantly improved.
Patent Information
- Application Number
- CN202411603738.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-12
AI Technical Summary
In the era of big data, it is difficult for existing technologies to effectively process and analyze massive data sets, especially in terms of data augmentation and annotation updates. Traditional methods have performance bottlenecks and cannot meet the needs of large-scale data sets.
An efficient image data intelligent enhanced transcoding method and system adapted to object detection is proposed. By centralizing the random clipping, rotation, and flip results and batch processing, it realizes automatic update of image augmented data set annotations, and supports mutual conversion between three common annotation formats (XML, JSON, TXT).
It significantly improves the speed and efficiency of data processing, reduces manual intervention and resource consumption, can maintain and expand data sets more effectively, and improves the training efficiency and accuracy of image processing and object detection algorithms.
Smart Images

Figure CN119150802B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an efficient image data intelligent enhancement transcoding method and system adapted to object detection. Background Art
[0002] In the context of the big data era, the amount and complexity of data that deep learning models, especially object detection models, need to process have increased significantly. Object detection models such as YOLO (You Only Look Once) perform well in real-time detection tasks, but they are highly dependent on the dataset. To improve the generalization ability and performance of the model, data augmentation has become a crucial step. However, updating annotations after data augmentation is a technical challenge, especially for large-scale datasets. Data augmentation can not only expand the dataset but also improve the generalization ability of the model. In existing technologies, image data augmentation is mainly performed by means of random cropping, random flipping, random rotation, color processing, and random scaling. In the big data era, the scale and diversity of data are increasing continuously, which requires object detection models to be able to process larger-scale datasets. At the same time, the automation, efficiency, and accuracy of data augmentation and annotation updating become more important.
[0003] Currently, mainstream data augmentation libraries such as PyTorch and TensorFlow provide rich image enhancement functions, but their support for annotation updating is relatively limited. This means that after enhancing the image, developers need to manually update the annotations, which is not only time-consuming but also error-prone. To address this challenge, some tools and libraries such as Albumentations have started to introduce automatic annotation updating functions. However, these methods perform poorly when dealing with large-scale datasets, and speed has become a key factor restricting their practical applications.
[0004] Through the above analysis, the problems and defects of the existing technologies are as follows: In the big data era, processing and analyzing massive datasets has become an urgent need. Traditional data processing methods often face performance bottlenecks and cannot effectively cope with the rapid increase in data volume. First, the process of manually updating annotations is very cumbersome and almost infeasible for large-scale datasets. Especially when dealing with tens of millions or even hundreds of millions of images, the workload of manual annotation and updating of annotations is huge, with low efficiency, which greatly limits the application scope of data augmentation technologies.
[0005] Even though some tools such as the Albumentations library in Python provide functions for automatic annotation updates, their processing speed is still difficult to meet actual requirements. Taking the processing of ten thousand images as an example, although this quantity performs acceptably on a small-scale dataset, its speed becomes inadequate when dealing with larger-scale datasets. This is mainly because existing methods have not been fully optimized in terms of design and implementation for computational efficiency, resulting in performance bottlenecks when processing high-dimensional data and complex operations. Summary of the Invention
[0006] To overcome the problems existing in related technologies, the disclosed embodiments of the present invention provide an efficient intelligent enhancement transcoding method and system for image data adapted to object detection. The objective of the present invention is to provide a technology that is efficient and capable of automatically updating annotations for image enhancement datasets, and can achieve mutual conversion between three commonly used annotation formats (XML, JSON, TXT) in the field of object detection. This technology significantly reduces the time required for enhancing data. Through this technology, users can save a large amount of computing power and time costs, and can more effectively maintain and expand datasets, improving the training efficiency and accuracy of image processing and object detection algorithms. At the same time, this system supports flexible conversion between multiple annotation formats, greatly enhancing the versatility and convenience of its application.
[0007] The technical solution is as follows: An efficient intelligent enhancement transcoding method for image data adapted to object detection records the results of all random cropping, random flipping, and random rotation operations in a centralized text file, which is used to improve the efficiency and speed of image data enhancement in large-scale image data processing; the method specifically includes the following steps:
[0008] S1, perform image enhancement on the original dataset using five data enhancement methods: random rotation, random flipping, random cropping, color processing, and random scaling, to obtain corresponding enhanced datasets; the corresponding enhanced datasets include a random rotation enhanced dataset, a random flipping enhanced dataset, a random cropping enhanced dataset, a color processing enhanced dataset, and a random scaling enhanced dataset;
[0009] S2, export the obtained random rotation enhanced dataset, random flipping enhanced dataset, and random cropping enhanced dataset, and respectively record them in a random cropping record.txt file, a random flipping record.txt file, and a random rotation record.txt file in combination with the TXT annotation file converted by the created image dataset annotation format conversion program;
[0010] S3, update the image annotations for the random cropping record.txt file, the random flipping record.txt file, and the random rotation record.txt file respectively to obtain corresponding new annotation sets.
[0011] In step S1, the original images are used for annotating the color - processed enhanced dataset and the randomly scaled enhanced dataset.
[0012] In step S2, the created image dataset annotation format conversion programs include XML2TXT and JSON2TXT;
[0013] This image dataset annotation format conversion program extracts the bounding box information for object detection from the XML annotation files in the PASCAL VOC format and converts the XML annotation files into TXT annotation files in the YOLO format.
[0014] In step S3, the randomly cropped record.txt file is used to update the image annotations, and the new annotation set includes:
[0015] The width and height of the cropped image are respectively , and the calculation formula is:
[0016] ;
[0017] ;
[0018] In the formula, is the coordinate of the lower - right corner of the cropped image, is the coordinate of the upper - left corner of the cropped image, is the coordinate of the lower - right corner of the cropped image, is the coordinate of the upper - left corner of the cropped image;
[0019] Taking the upper - left corner of the cropped image as the origin, the coordinates of the new bounding box are offset. is the coordinate of the center point of the new bounding box, is the width and height of the new bounding box, and the calculation formula is:
[0020] ;
[0021] ;
[0022] ;
[0023] ;
[0024] In the formula, is the coordinate of the upper - left corner of the new bounding box, is the coordinate of the upper - left corner of the new bounding box, is the coordinate of the lower - right corner of the new bounding box, The coordinates of the lower right corner of the new bounding box; Coordinates;
[0025] After obtaining the new bounding box, perform scaling to obtain the annotations of the cropped image .
[0026] Obtain the annotations of the cropped image The calculation formula is:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] In the formula, is the width of the cropped image, is the height of the cropped image.
[0032] In step S3, randomly flip the record.txt file to update the image annotations, and obtain a new annotation set including:
[0033] The corresponding original annotation of the record is , where is the category, is the abscissa of the center point of the bounding box, is the ordinate of the center point of the bounding box, is the width of the bounding box, is the height of the bounding box;
[0034] is the proportional size of the bounding box relative to the original image, with the upper left corner as the origin, and the updated annotation is .
[0035] In step S3, randomly rotate the record.txt file to update the image annotations, and obtain a new annotation set including:
[0036] If the first record of the randomly rotated record.txt file is , where is the file name, is the rotation angle, is the width of the original image, is the height of the original image;
[0037] If the rotation angle of the image is counterclockwise, update the calculation of the annotations of the image after counterclockwise rotation, where , , are respectively the width and height of the rotated image, and the expressions are:
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] ;
[0045] ;
[0046] In the formula, are respectively the side lengths in the counterclockwise rotated image;
[0047] After obtaining the width and height of the new annotation box according to formula (17) and formula (18), scale the width, height and the coordinates of the center of the new annotation box according to formula (7) - formula (10), and finally update the annotation of the counterclockwise rotated image .
[0048] If the rotation angle of the image is clockwise, update the calculation of the annotation of the clockwise rotated image, where is the absolute value of the rotation angle, the clockwise rotation angle is negative, and the counterclockwise rotation angle is positive, ;
[0049] ;
[0050] .
[0051] Among them, are respectively the side lengths in the clockwise rotated image.
[0052] Another object of the present invention is to provide an efficient image data intelligent enhancement transcoding system adapted to object detection. This system implements the efficient image data intelligent enhancement transcoding method adapted to object detection. This system includes:
[0053] An image enhancement module, which is used to perform image enhancement on the original data set by using five data enhancement methods of random rotation, random flipping, random cropping, color processing, and random scaling, and obtain the corresponding enhanced data set; the corresponding enhanced data set includes a random rotation enhanced data set, a random flipping enhanced data set, a random cropping enhanced data set, a color processing enhanced data set, and a random scaling enhanced data set;
[0054] A record.txt file acquisition module, which is used to export the obtained random rotation enhanced data set, random flipping enhanced data set, and random cropping enhanced data set, and respectively record them into the random cropping record.txt file, random flipping record.txt file, and random rotation record.txt file in combination with the TXT annotation files converted by the created image data set annotation format conversion program;
[0055] A new annotation set acquisition module, which is used to update the image annotations of the random cropping record.txt file, random flipping record.txt file, and random rotation record.txt file respectively, and obtain the corresponding new annotation sets.
[0056] Furthermore, the system is carried on a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the functions in the above-mentioned efficient image data intelligent enhancement transcoding system adapted to object detection can be realized.
[0057] Combining all the above technical solutions, the beneficial effects of the present invention are as follows: By introducing batch processing technology, the present invention has significant advantages in terms of speed and versatility, and greatly improves the data processing ability. Especially in the fields of deep learning and object detection, the automatic annotation update function of the present invention significantly improves the efficiency and accuracy of data annotation, which is very crucial for training high-precision models. The automatic annotation update reduces the need for manual intervention, accelerates the data preparation process, and enables the model to learn and adapt to new data faster. Under the same hardware environment, in order to improve the versatility of the method, the present invention not only supports common TXT format annotations, but also develops the function of mutual conversion between three annotation formats (XML, TXT, JSON). This enables the method to be widely applied to various data sets in different formats, greatly facilitating the use of developers in different projects.
[0058] The present invention not only solves the performance bottleneck problem in processing large-scale data sets, but also provides strong technical support for the practical application of deep learning object detection. This is particularly important for those big data application scenarios that require real-time or near-real-time data analysis and decision-making support, such as video surveillance, autonomous vehicle systems, medical image analysis, etc., all of which will benefit from it. Such technological progress is exactly the innovative move required in the big data era, enabling data analysis and processing to be more efficient, accurate, and real-time.
[0059] The present invention realizes the efficient and automated management of image enhancement data set annotation, significantly improves the efficiency of data preparation in the fields of object detection and image processing, and provides basic work for improving the speed and accuracy of model training, enabling enterprises and research institutions to gain a competitive advantage in the development and implementation of AI vision models. The commercialization of this technology will promote the wide application of intelligent image analysis technology, such as in multiple fields including unmanned driving, security monitoring, and medical image analysis.
[0060] In terms of automatically updating image data set annotation and reducing disk I / O operations, the existing image processing tools or systems have not provided a complete solution that is both efficient and resource-saving. In addition, the flexible conversion function of multiple annotation formats provided by the present invention also solves the format dependency problem existing in the prior art, increasing the applicability and flexibility of the system.
[0061] The present invention successfully solves the problems of low efficiency and high resource consumption in processing large-scale image data. This is a long-standing challenge in the research fields of image processing and object detection, especially when dealing with high-definition and large-capacity image data. The present invention overcomes the technical bias existing in traditional image data processing, that is, the belief that efficient image data set management and annotation must rely on manual operations or inefficient single-annotation format processing. Through the technical solutions of automatic update and flexible conversion of multiple annotation formats, this technical bias is effectively solved, breaking through the limitations of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure;
[0063] Figure 1 is a flowchart of an efficient image data intelligent enhancement transcoding method adapted to object detection provided by an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of data enhancement of an efficient image data intelligent enhancement transcoding method adapted to object detection provided by an embodiment of the present invention;
[0065] Figure 3It is the flowchart for updating the flipped image annotation provided by the embodiments of the present invention;
[0066] Figure 4 It is the flowchart for updating the image annotation after random cropping provided by the embodiments of the present invention;
[0067] Figure 5 It is the flowchart for updating the image annotation after random rotation provided by the embodiments of the present invention;
[0068] Figure 6 It is the calculation flowchart for the image annotation after counterclockwise rotation provided by the embodiments of the present invention;
[0069] Figure 7 It is the calculation flowchart for the image annotation after clockwise rotation provided by the embodiments of the present invention. Detailed implementation manners
[0070] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0071] The innovation of the present invention lies in that: by centralizing the recording of the results of random cropping, random rotation, and random flipping and performing batch processing, the present invention significantly reduces disk I / O operations, thereby improving the processing speed and system performance. At the same time, this method effectively utilizes system resources and adopts a modular design, which not only optimizes resource management but also enhances the maintainability and scalability of the system. This makes the present invention more advantageous and applicable in processing large-scale image data.
[0072] Batch processing optimization: By first recording the results of all random cropping operations in a centralized text file, the present invention enables batch processing of the entire data set. This method improves efficiency and speed in processing large-scale image data.
[0073] Reducing disk I / O operations: The present invention significantly reduces the number of disk writes and only performs a single batch write after all image processing is completed. This design reduces frequent disk I / O operations, thereby improving the overall performance of the system and reducing the load on the disk.
[0074] Effective utilization of system resources: By recording and caching processing results, the present invention more effectively manages memory and storage resources. In an environment with limited resources, the resource usage can be optimized by adjusting the batch size to adapt to different operating environments.
[0075] Maintainability and scalability: Since a modular processing flow is adopted (i.e., the three stages of tailoring, recording, and updating are clearly separated), the present invention is easy to maintain and upgrade. For example, the tailoring algorithm can be easily improved or the annotation file format can be changed without large-scale modification of the entire system.
[0076] Embodiment 1, as Figure 1 shown, the efficient image data intelligent enhancement transcoding method adapted to object detection provided by the embodiment of the present invention includes:
[0077] S1, performing image enhancement on the original data set by using five data enhancement methods of random rotation, random flipping, random tailoring, color processing, and random scaling to obtain corresponding enhanced data sets; the corresponding enhanced data sets include a random rotation enhanced data set, a random flipping enhanced data set, a random tailoring enhanced data set, a color processing enhanced data set, and a random scaling enhanced data set;
[0078] S2, exporting the obtained random rotation enhanced data set, random flipping enhanced data set, and random tailoring enhanced data set, and respectively recording them into a random tailoring record.txt file, a random flipping record.txt file, and a random rotation record.txt file in combination with the TXT annotation file converted by the created image data set annotation format conversion program;
[0079] S3, respectively updating the image annotations of the random tailoring record.txt file, the random flipping record.txt file, and the random rotation record.txt file to obtain corresponding new annotation sets.
[0080] Exemplarily, in step S1, since data enhancement can improve the performance and generalization ability of the model, data enhancement is often applied in the field of object detection. Common image data enhancement methods mainly include random rotation, random flipping, random tailoring, color processing, random scaling, etc. Therefore, the present invention develops a Python program to perform corresponding annotation updates on the above five data enhancement methods.
[0081] Exemplarily, among the data enhancement methods, color processing does not affect the size of the original image. Although random scaling will change the size of the original image and the annotation box, the labels used by YOLO are the proportional sizes of the original image. Therefore, color processing and random scaling do not need to update the annotations and can both use the annotations of the original image.
[0082] Exemplarily, in step S2, as Figure 2As shown, an image dataset annotation format conversion program is created. Specifically, in the field of object detection, the commonly used image data label formats are mainly XML, JSON, and TXT. Since the TXT format has the advantages of simplicity, small file size, easy processing, and flexibility, and the label format used by the current mainstream object detection algorithm YOLO is TXT, an image dataset annotation format conversion program (label format conversion program) is created in the present invention, and the original image dataset and the image dataset label in TXT format are obtained accordingly.
[0083] The created image dataset annotation format conversion program includes XML2TXT and JSON2TXT.
[0084] This code will extract the bounding box information for object detection from an XML annotation file in PASCAL VOC format and convert it into a TXT annotation file in YOLO format. The code performs the following steps:
[0085] Define the convert function to convert the bounding box information from VOC format to YOLO format. In the YOLO format, the bounding box is represented as a class label, center coordinates (x, y), width, and height, and all values are normalized to [0, 1].
[0086] Define the convert_annotation function to process all XML files in a folder and convert each file into a corresponding TXT file.
[0087] This function first lists the XML files and creates an output TXT file name for each file.
[0088] Then, it parses the XML file to obtain the image size and information about the target objects from it.
[0089] For each target object in the XML file, this function extracts the class name, difficulty level, and bounding box coordinates.
[0090] If the target is not in the defined class list or is marked as difficult to recognize, the target will be skipped.
[0091] Otherwise, it converts the class name to a class index and converts the bounding box coordinates to YOLO format.
[0092] Finally, the data in YOLO format is written to the TXT file.
[0093] In the __main__ block, the code defines the class list classes to be converted, as well as the input and output paths of the XML and TXT files.
[0094] The code actually performs the conversion using the convert_annotation function.
[0095] In summary, this script is used to convert the annotation files in the PASCAL VOC dataset from XML format to TXT format required for YOLO training;
[0096] Exemplarily, in step S3, for random flipping, the common flipping method is to flip the image along the central y-axis (i.e., the vertical axis). The image annotation update process of the present invention is as Figure 3 shown. It includes: the original dataset is enhanced by random flipping, and the flipping direction of the image is recorded in the random flipping record.txt file while performing data augmentation. The original dataset combines with the program according to the.txt file to obtain a new annotation set.
[0097] Assume that a certain record in the random flipping record.txt is , where is the file name, is the flipping direction. The corresponding original annotation is , where is the category, is the abscissa of the center point of the bounding box, is the ordinate of the center point of the bounding box, is the width of the bounding box, is the height of the bounding box;
[0098] is the proportional size of the bounding box relative to the original image, with the upper left corner as the origin. The updated annotation is .
[0099] Exemplarily, in step S3, for random cropping, the image annotation update process after random cropping of the present invention is as Figure 4 shown.
[0100] Taking the first record in the random cropping record.txt file as an example, 00001 is the file name, is the width of the original image, is the height of the original image, is the coordinate of the upper left corner of the cropped image, is the coordinate of the upper left corner of the cropped image, is the coordinate of the lower right corner of the cropped image, is the coordinate of the lower right corner of the cropped image (note: the coordinates are all actual distances, with the upper left corner of the original image as the coordinate origin).
[0101] The original annotation set provides those with the same file name which are respectively the coordinates of the upper left corner of the original bounding box, the coordinates of the upper left corner of the original bounding box, the coordinates of the lower right corner of the original bounding box, the coordinates of the lower right corner of the original bounding box.
[0102] which are respectively the coordinates of the upper left corner of the new bounding box, the coordinates of the upper left corner, the coordinates of the lower right corner, the coordinates of the lower right corner;
[0103] After obtaining the vertex coordinates of the upper left and lower right corners of the original bounding box and the cropped image, compare the coordinate points as Figure 4 shown, and finally obtain the positions of the coordinate points of the upper left and lower right corners of the new cropping box.
[0104] Specifically, it includes:
[0105] Assume that the width and height of the cropped image are respectively , and the calculation formula is as follows:
[0106] ;
[0107] ;
[0108] In the formula, is the coordinate of the lower right corner of the cropped image, is the coordinate of the upper left corner of the cropped image, is the coordinate of the lower right corner of the cropped image, is the coordinate of the upper left corner of the cropped image;
[0109] Taking the upper left corner of the cropped image as the origin, the coordinates of the new bounding box are offset, is the coordinate of the center point of the new bounding box, is the width and height of the bounding box, and the calculation formula is:
[0110] ;
[0111] ;
[0112] ;
[0113] ;
[0114] In the formula, is the x - coordinate of the upper - left corner of the new bounding box, is the y - coordinate of the upper - left corner of the new bounding box, is the x - coordinate of the lower - right corner of the new bounding box, is the y - coordinate of the lower - right corner of the new bounding box;
[0115] After obtaining the of the new bounding box, perform scaling to obtain the annotation of the cropped image , and the calculation formula is:
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] In the formula, is the width of the cropped image, is the height of the cropped image.
[0121] It can be understood that the function of the above formula is to calculate the new annotation (the x - coordinate of the center point of the annotation box, the y - coordinate of the center point of the annotation box, the width of the annotation box, the height of the annotation box) of the image after random cropping.
[0122] Exemplarily, in step S3, for random rotation, the update process of the image annotation after random rotation is as Figure 5 shown. It includes: the original dataset is enhanced by random rotation, and various parameters of the image rotation are recorded in the random rotation record.txt file while data is enhanced. The original dataset combines with the program according to the.txt file to obtain a new annotation set.
[0123] If the first record of the random rotation record.txt file is , where is the file name, is the rotation angle, is the width of the original image, is the height of the original image;
[0124] As Figure 5 , Figure 6 shown, if the rotation angle of the image is counter - clockwise, then update the calculation of the image annotation after counter - clockwise rotation, where , , are the width and height of the rotated image respectively, and the expression is:
[0125] ;
[0126] ;
[0127] ;
[0128] ;
[0129] ;
[0130] ;
[0131] ;
[0132] ;
[0133] wherein, are respectively the side lengths in the image after counterclockwise rotation;
[0134] It can be understood that the function of the above formula is to calculate the new annotation (the x coordinate of the center point of the annotation box, the y coordinate of the center point of the annotation box, the width of the annotation box, the height of the annotation box) of the image after random rotation (counterclockwise rotation).
[0135] According to Formula (17) and Formula (18), after obtaining the width and height of the new annotation box, scale the width and height of the new annotation box and the coordinates of the center according to Formulas (7)-(10), and finally update the annotation of the image after counterclockwise rotation .
[0136] If the rotation angle of the image is clockwise, update the calculation of the annotation of the image after clockwise rotation, wherein, is the absolute value of the rotation angle, the rotation angle clockwise is negative, and the rotation angle counterclockwise is positive, ;
[0137] ;
[0138] .
[0139] wherein, are respectively the side lengths in the clockwise-rotated image.
[0140] It can be understood that the function of the above formula is to calculate the new annotation (the x coordinate of the center point of the annotation box, the y coordinate of the center point of the annotation box, the width of the annotation box, the height of the annotation box) of the image after random rotation (clockwise rotation).
[0141] As can be seen from the above embodiments, the existing method for updating image enhancement annotations uses the albumentations library in Python to perform data augmentation on single-image data and then update the annotations. Taking random cropping as an example for comparative experimental analysis using the same dataset, the execution time of the present invention is 25.24 s in total, while the execution time of the existing method is 62 s. The present invention saves 36.76 s of time and improves the efficiency by 59.29% compared with the existing method. The specific advantages are as follows:
[0142] Improve processing efficiency: By dividing the processing flow of the entire dataset into three stages - cropping, recording, and annotation update, the present invention supports batch processing and significantly improves the data processing speed. Batch execution of data cropping and annotation update can greatly reduce the total processing time.
[0143] Optimize resource utilization: Reducing disk I / O operations not only improves the system response speed but also effectively reduces the disk wear of the hardware and extends the service life of the system. By using the method of centralized storage of processing results, memory and storage space are better managed, especially showing great advantages in resource-constrained systems.
[0144] Enhance system maintainability and scalability: The modular design of the present invention enables each processing unit (cropping, recording, updating) to be independently modified and upgraded without affecting other parts, facilitating system maintenance and expansion. For example, the image processing algorithm can be easily replaced or new data format support can be introduced without affecting the stable operation of the overall system.
[0145] Reduce environmental dependence: Since the dependence on high-speed disk I / O operations is reduced, the present invention is applicable to a wider range of hardware environments without particularly emphasizing high-performance disk systems, thereby reducing the deployment cost and enhancing the flexibility and universality of the application.
[0146] The present invention is applicable to the enhancement and annotation update of massive image data in the big data era, which can effectively improve the data processing efficiency and save a large amount of computing power resources and time costs.
[0147] Embodiment 2. The embodiment of the present invention provides an efficient intelligent image data enhancement transcoding system adapted to object detection, including:
[0148] An image enhancement module, which is used to perform image enhancement on the original dataset by using five data augmentation methods of random rotation, random flipping, random cropping, color processing, and random scaling to obtain corresponding enhanced datasets; the corresponding enhanced datasets include a random rotation enhanced dataset, a random flipping enhanced dataset, a random cropping enhanced dataset, a color processing enhanced dataset, and a random scaling enhanced dataset;
[0149] The record.txt file acquisition module is used to export the obtained randomly rotated augmented dataset, randomly flipped augmented dataset, and randomly cropped augmented dataset, and respectively record them into the randomly cropped record.txt file, randomly flipped record.txt file, and randomly rotated record.txt file in combination with the TXT annotation files converted by the created image dataset annotation format conversion program;
[0150] The new annotation set acquisition module is used to update the image annotations of the randomly cropped record.txt file, randomly flipped record.txt file, and randomly rotated record.txt file respectively to obtain the corresponding new annotation sets.
[0151] To further illustrate the related effects of the embodiments of the present invention, the following experiments are carried out:
[0152] The present invention proposes a batch processing mechanism: describes how to centrally process (crop) a large number of images, and how to uniformly record the cropped results in a file to optimize the processing efficiency and reduce disk I / O operations.
[0153] Data aggregation and processing flow: details the flow, management, and optimization of data in batch processing, including how to collect image data, perform batch cropping, and effectively integrate the results into the final annotation file.
[0154] Cache and temporary storage optimization: emphasizes how to effectively manage the large amount of data generated during batch processing by using memory caching or temporary storage strategies, avoid resource exhaustion, and maintain system performance.
[0155] Efficient data synchronization and recovery mechanism: demonstrates the data synchronization mechanism in batch processing to ensure the consistency of annotation and image data, and describes the data recovery strategy in case of system failures.
[0156] The present invention redefines the following technical points:
[0157] A unified data augmentation and annotation update framework, focusing on protecting the data processing flow design, especially the implementation methods of batch data augmentation and annotation synchronization update.
[0158] Data processing and I / O optimization methods, focusing on protecting how to reduce the dependence on disk I / O during data processing, including the implementation of cache usage and the optimal writing strategy.
[0159] Use the technologies proposed by the present invention and the existing technologies (albumentations library) to perform random rotation, random flipping, and random cropping on the original dataset (15052 images) respectively. The experimental results are shown in Table 1.
[0160] Table 1 Comparison of experimental results
[0161]
[0162] As can be seen from Table 1, the efficiency improvements of random cropping, random flipping, and random rotation are 59.29%, 63.68%, and 34.63% respectively.
[0163] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. An efficient image data intelligent enhancement transcoding method suitable for target detection, characterized in that: The method records the results of all random cropping, random flipping, and random rotation operations in a centralized text file, and is used to improve the efficiency and speed of image data enhancement in large-scale image data processing. The method specifically includes the following steps: S1, performing image enhancement on the original data set using five data enhancement methods including random rotation, random flipping, random cropping, color processing, and random scaling to obtain a corresponding enhanced data set; the corresponding enhanced data set includes a random rotation enhanced data set, a random flipping enhanced data set, a random cropping enhanced data set, a color processing enhanced data set, and a random scaling enhanced data set; S2, export the obtained random rotation enhancement data set, random flip enhancement data set, and random crop enhancement data set, and combine them with the TXT annotation file converted by the created image data set annotation format conversion program, and record them into the random crop record.txt file, the random flip record.txt file, and the random rotation record.txt file respectively; S3, update the image annotations of the random cropping record.txt file, the random flipping record.txt file, and the random rotation record.txt file respectively to obtain the corresponding new annotation sets; In step S3, the record .txt file is randomly cropped to update the image annotation, and the new annotation set is obtained including: The width and height of the cropped image are C w , C h , the calculation formula is: C w =C right -C left (1) C h =C bottom -C top (2) In the formula, C right is the x coordinate of the lower right corner of the cropped image, C left is the x coordinate of the upper left corner of the cropped image, C bottom is the y coordinate of the lower right corner of the cropped image, C top is the y coordinate of the upper left corner of the cropped image; Taking the upper left corner of the cropped image as the origin, the coordinates of the new bounding box are offset, bb x ,bb y is the coordinate of the center point of the new bounding box, bbx w ,bbx h is the width and height of the new bounding box, calculated as: In the formula, is the x coordinate of the upper left corner of the new bounding box, is the y coordinate of the upper left corner of the new bounding box, is the x coordinate of the lower right corner of the new bounding box, is the y coordinate of the lower right corner of the new bounding box; Get the bbx of the new bounding box x ,bbx y ,bbx w ,bbx h After that, we scale the image to get the annotation [kx′y′ w′ h′] of the cropped image. The calculation formula for the annotation [kx′ y′ w′ h′] of the cropped image is: x′=bbx x / C w (7) y′=bbx y / C h (8) w′=bbx w / C w (9) h′=bbx h / C h (10) In the formula, C w is the width of the cropped image, C h is the height of the cropped image; In step S3, the record .txt file is randomly flipped to update the image annotation, and the new annotation set is obtained including: The original annotation corresponding to the record is [kxywh], where k is the category, x is the horizontal coordinate of the center point of the bounding box, y is the vertical coordinate of the center point of the bounding box, w is the width of the bounding box, and h is the height of the bounding box; [xywh] is the proportional size of the bounding box relative to the original image, with the upper left corner as the origin. The updated annotation is [k1-(x+w) ywh]; In step S3, the randomly rotated record .txt file is used to update the image annotations, and a new annotation set is obtained including: If the first record of the random rotation record.txt file is [0001 angle o w o h ], where 0001 is the file name, angle is the rotation angle, and o w is the width of the original image, o h is the height of the original image; If the image is rotated counterclockwise, the calculation of the image annotation after counterclockwise rotation is updated, where ∠2 = ∠5 = angle, ∠3 = ∠1-∠2, r w , r h are the width and height of the circumscribed rectangle of the rotated image, respectively, and the expression is: L1=((o w ×x) 2 +(o h ×y) 2 ) 0.5 (11) ∠1=arcsin((o h ×y) / L1)(12) L4=L1×cos(∠3)(13) <h2 style=";text-align:left;direction:ltr">L3 = tan(∠4)×(1-x)×o<h2 style=";text-align:left;direction:ltr"> w <h2 style=";text-align:left;direction:ltr"> (14) L2=o h ×y(15) L5=(L2+L3)×cos(∠5)(16) r w =o w ×cos(∠4)+o h ×sin(∠4)(17) r h =o h ×cos(∠4)+o w ×sin(∠4)(18) Where L1, L2, L3, L4, and L5 are the lengths of the sides in the counterclockwise rotated image; According to formula (17) and formula (18), the width and height of the rotated image are obtained, and the width, height and center coordinates of the new annotation box are scaled according to formula (7)-formula (10), and finally the annotation [kx′ y′ w′ h′] of the counterclockwise rotated image is updated; If the image is rotated clockwise, the calculation of the image annotation after the clockwise rotation is updated, where ∠6 is the absolute value of the rotation angle, the clockwise rotation angle is a negative number, and the counterclockwise rotation angle is a positive number, ∠8 = 90° - ∠6 - ∠1; L6=L1×cos(∠8)(19) <h2 style=";text-align:left;direction:ltr">L7=o<h2 style=";text-align:left;direction:ltr"> h <h2 style=";text-align:left;direction:ltr"> ×sin(∠6)+L6×tna(∠8)(20) Among them, L6 and L7 are the lengths of the sides in the clockwise rotated image.
2. The efficient image data intelligent enhancement transcoding method adapted for target detection according to claim 1, characterized in that: In step S1, the color processing enhanced data set and the random scaling enhanced data set are annotated using the original images.
3. The efficient image data intelligent enhancement transcoding method adapted for target detection according to claim 1, characterized in that: In step S2, the created image dataset annotation format conversion programs include XML2TXT and JSON2TXT; This image dataset annotation format conversion program will extract the bounding box information of the target detection from the XML annotation file in PASCAL VOC format and convert the XML annotation file into a TXT annotation file in YOLO format.
4. An efficient image data intelligent enhancement transcoding system suitable for target detection, characterized in that: The system implements the efficient image data intelligent enhancement transcoding method adapted for target detection as described in any one of claims 1 to 3, and the system comprises: An image enhancement module is used to perform image enhancement on the original data set by using five data enhancement methods, namely, random rotation, random flipping, random cropping, color processing, and random scaling, to obtain a corresponding enhanced data set; the corresponding enhanced data set includes a random rotation enhanced data set, a random flipping enhanced data set, a random cropping enhanced data set, a color processing enhanced data set, and a random scaling enhanced data set; The record .txt file acquisition module is used to export the obtained random rotation enhancement data set, random flip enhancement data set, and random crop enhancement data set, and at the same time, combine the TXT annotation file converted by the created image data set annotation format conversion program, and record them respectively into the random crop record .txt file, the random flip record .txt file, and the random rotation record .txt file; The new annotation set acquisition module is used to update the image annotations of the random cropping record .txt file, the random flipping record .txt file and the random rotation record .txt file respectively to obtain the corresponding new annotation set.
5. The efficient image data intelligent enhancement transcoding system adapted for target detection according to claim 4, characterized in that: The system is mounted on a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the functions of the efficient image data intelligent enhancement transcoding system suitable for target detection can be realized.
Citation Information
Patent Citations
Method for enhancing labeled samples for artificial intelligence training
CN117670717A
Target detection method and system applied to intelligent construction site security
CN118262290A