Error-Traceable Image Data Structure and Annotation Method
Through innovative storage methods of HPC and TPC file formats, the problems of error traceability and sample integration in image data set acquisition are solved, and data integration efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202310112412.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-14
AI Technical Summary
The existing image data collection method relies on manual processing, resulting in large errors and low efficiency, and the incorrect image data cannot be traced to the original batch, increasing the cost of manpower and material resources, and the integrated storage and modification of samples cannot be achieved.
The HPC file format is adopted to store image data at unit and unit levels, add traceability features, and realize lossless storage and modification of image data through HPC file annotation and modification tools. The TPC file format is used for final conversion, supporting error sample traceability and integrated sample processing.
The traceability of wrong samples to the original batch is realized, which reduces manpower waste, improves the speed and accuracy of data integration manuscripts, and supports integrated processing of large-data samples.
Smart Images

Figure CN116303237B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and relates to an image data structure and annotation method, in particular to an error-retraceable image data structure and annotation method for mass data collection tasks. Background Art
[0002] With the increasing development of the network and the booming development of the Internet of Things, it is undoubtedly an inevitable future development trend that everything can be connected to the network. During this development process, the demand for various types of image data sets is also increasing. Handwriting recognition requires a handwritten text data set, and face recognition requires a face image data set. When intelligent recognition systems gradually replace traditional manual recognition work with their high efficiency and security characteristics, the demand for various image data sets in the Internet of Things also rises accordingly. However, the current collection of image data sets in the field of computer vision depends more on manual processing by participants from start to finish. Participants need to manually classify, screen, number each piece of collected data, perform certain image processing on the images, and complete information annotation before a valid data set is considered collected. In addition to the errors caused by the fatigue of manual processing of a large number of image data, this data set collection method also has the problem that it is cumbersome to process when a batch of different types of image data in the original data goes wrong and may be omitted, which greatly reduces the reliability of a valid data set. Therefore, a data set often needs to be processed manually multiple times before it can be put into use. The human, material, and time costs required to complete the collection of a data set in this way are extremely high, and the manuscript process is slow. As a result, many data sets in the field of computer vision have gradually fallen behind the development speed of their own recognition technologies, dragging down the development pace of the computer vision field. The inability to trace back the classified incorrect image data to the original data set batch for key analysis is an important problem in the existing image data set collection. If the incorrect image data can be traced back to the original data set of the incorrect batch for key analysis, it will undoubtedly save a large amount of human and time costs, speed up the data integration process, and improve the accuracy of the data set.
[0003] The existing image data set formats for the field of computer vision generally include lossless formats such as TIFF (Tag Image File Format), BMP (Bitmap), and PNG (Portable Network Graphics). The most important feature of the image file format for the field of computer vision is losslessness, without loss during the processing. To ensure that the data set does not affect the recognition efficiency and accuracy of the recognition technology, the original resolution of the image should be retained to the greatest extent, and the image quality loss caused by image compression and other reasons should be reduced. In traditional data formats, images are independent of each other, and there is no association between two images, which is applicable to almost all fields that require image data.
[0004] The application of image data formats to the collection of massive image data sets undoubtedly faces new challenges. Massive image data generally obtains a large amount of raw data in batches. In the process of processing the raw data, the image data is often reclassified according to the label, and secondary processing is performed on this basis. When a batch of raw data is wrong, because its content has been classified under different label categories, it is an extremely wasteful choice of manpower and time to pick out and delete them one by one in the secondary processing. Another important issue is the labeling of image content labels. The labeling of traditional image data formats can generally only be stored in a classified manner using folder naming, or directly using label naming for image files, which is very inconvenient when collecting image data sets with multiple labels or massive data sets. Obviously, when collecting image data sets for computer vision, we should provide a new data format that is more suitable for the image data set collection and processing process and a corresponding system.
[0005] A storage format for handwritten Chinese character image data: CASIA offline Chinese character library gnt file format (its format is shown in Table 1). HWDB1.1 is a handwritten Chinese character library from CASIA, collected by the State Key Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences. It contains 7185 commonly used Chinese characters and 171 special symbols, and these data are stored in gnt format. According to the characteristics of the handwritten Chinese character library, this format file adds a tag bit to the header file of each character image to identify the corresponding character. In order to adapt to the storage of large amounts of data and reduce storage space as much as possible, the total length of the character image and the length and width fields of the character image are added to each character image data content to describe the basic information of the character image. The specific pixel content of the character image is output and arranged row by row to complete the storage of a character content. The subsequent character content is connected to the previous character data content in the same format, so that the plural character image data is stored in a serial arrangement. There are the following disadvantages:
[0006] (1) It is impossible to trace the erroneous image samples back to the original batch image samples.
[0007] (2) It is not possible to integrate sample modification and deletion.
[0008] (3) Sample data with the same label cannot be stored in an integrated manner.
[0009] Table 1 gnt file format
[0010]
[0011] Summary of the invention
[0012] In view of the above-mentioned drawbacks of the prior art, the present invention provides an error-retraceable image data structure and annotation method for mass data acquisition tasks.
[0013] The object of the present invention is achieved by the following technical solutions:
[0014] An error-retraceable image data structure and annotation method for mass data acquisition tasks, comprising the following steps:
[0015] Step 1: Read in image data
[0016] (1) Obtain all image sample paths and sample names under the sample storage folder path by iterating through the target directory;
[0017] (2) Store the sample paths and sample names in a vector and save them as information of files to be processed;
[0018] (3) Obtain image files according to the information of files to be processed in the vector, read them into memory for subsequent processing, and after processing each image file, remove the information of this image file from the vector and then start processing the next image file;
[0019] Step 2: Preprocess the image content
[0020] (1) Perform binarization processing on the image information using the Otsu algorithm to obtain a binary image;
[0021] (2) Detect the edge information of the printed layout by using the hough transform on the binary image to obtain the tilt angle value of the image;
[0022] (3) Perform tilt correction and preliminary noise reduction on the original image according to the tilt angle value and the binary image;
[0023] (4) Cut the rectangular image areas at the four corners of the layout in proportion, perform projection analysis respectively to confirm the area where the QR code is located, and complete the flipping process of the original image according to the position of the area where the QR code is located;
[0024] (5) Use the zxing library to obtain the information inside the QR code and get the basic information including the text order corresponding to this image file;
[0025] (6) Output the image path, image name, and basic information of the image sample as a preprocessing document;
[0026] Step 3: Convert the image data into the HPC file format
[0027] (1) Use the preprocessing document to confirm the number of units in each unit and the corresponding basic information, and generate a unit header according to the HPC file specification;
[0028] (2) Locate the effective image units through layout projection and analysis of connected domains;
[0029] (3) Segment the original image according to the located image position, obtain the unit image content, and fill it into the HPC file according to the HPC file specification;
[0030] Step 4: Labeling and modifying HPC files
[0031] (1) Display of HPC file content: The HPC file modification tool reads the HPC file in accordance with the standard format, generates an image matrix based on the read unit image content, and segments the image using a connected domain search method. The segmented connected domain information is stored in a two-dimensional blockchain array in groups of words for subsequent manual deletion. When displayed, a blue rectangular frame is used to display the connected domain outer frame for easy identification and operation;
[0032] (2) HPC file annotation modification: The user selects the connected domain that is obviously a noise part by left-clicking and dragging the mouse. The selected connected domain is indicated by turning the outer border red. When the delete button is pressed, the selected connected domain on the interface is deleted, and the deleted content is synchronously updated to the HPC file, so as to complete the modification of the HPC file image information at the same time;
[0033] Step 5: Convert HPC file format to TPC file format
[0034] (1) Write a format conversion program that will use multi-threading to simultaneously open a certain number of HPC files (the number is adjusted based on the machine configuration and actual load) that may be used based on the HPC file ID to be read to speed up processing;
[0035] (2) Read the HPC file quickly in binary mode, and according to the Chinese character label of the unit image information, divert the unit image information to the thread that has opened the corresponding TPC file, and write the unit image information into the TPC file that is consistent with its own label;
[0036] Step 6: Convert TPC file format to target format file
[0037] (1) The conversion program iterates over the target directory, storing an entire TPC file each time for processing until all TPC files in the directory have been processed;
[0038] (2) Restore the unit image information in the TPC file to an image matrix;
[0039] (3) Check the image matrix, remove the blank image, and cut the blank area around the normal image;
[0040] (4) Check the image size. If the image is larger than the required size (100×100), scale down the image proportionally to the required size.
[0041] (5) Conduct a statistical analysis of the handwriting density for the valid pixels. If the density is less than the threshold value (e.g., 110), thicken the handwriting using the method of adaptive histogram equalization enhancement.
[0042] (6) Output the processed image matrix in the target format (such as jpg, png, etc.) to the corresponding folder.
[0043] In the present invention, the specification format of the HPC file is divided into two levels, namely the unit level with a single original sample as the unit, and the unit level with a valid image sample under a single original sample as the unit. Units are independent of each other. Each unit contains at least one unit and at most 65,535 units. The units are arranged in a cascading manner within the unit, and the units within the unit are also arranged in a cascading manner. Add fields for the corresponding original sample file path, basic information of the original sample, length of the previous unit, and the number of units within the unit to the data structure of the unit header. Add fields for describing the unit length, whether the unit is valid, the unit label, and the length and width of the image to the unit header. After the unit header, input the pixel information line by line to complete the lossless storage of the image data, which can solve the drawback in the prior art that error image samples cannot be traced back to the original batch image samples.
[0044] In the present invention, the annotation and modification of the HPC file refer to the reading, display, and overwrite modification of the image data content of the HPC file. When reading and displaying the image data stored in the HPC file, record the starting pointer position of the data in the HPC file corresponding to each image data. Synchronize with the corresponding pointer position according to the content of the modification operation of the program to complete the overwrite modification of the content of the HPC file. Perform deletion or restoration operations through the valid bit field within the unit in the HPC format, which can solve the drawback in the prior art that sample modification and deletion cannot be integrated.
[0045] In the present invention, the specification format of the TPC file removes the unit level from the HPC file specification format and adds the HPC file ID field corresponding to the unit and the position field in the corresponding HPC file to the data structure at the unit level to ensure its traceability. Screen the image data that meets the requirements in the HPC according to the unit label and dump it into the TPC file. Arrange the units in a cascading manner in the TPC file, which can solve the drawback in the prior art that sample data with the same label is not stored integrally.
[0046] Compared with the prior art, the present invention has the following advantages:
[0047] 1. It can trace the wrong sample cases back to the content of the original sample batch, which helps to process the wrong samples that appear in batches.
[0048] 2. It can integrate the modification and deletion of a large amount of sample information, reduce intermediate steps, greatly reduce the time used to process the collected data, reduce the waste of manpower, and improve efficiency.
[0049] 3. It can integrate the storage of samples with the same label, which is beneficial for storage, use, and sample classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of an error-traceable handwritten Chinese character image data structure and annotation method for a massive data collection task of the present invention;
[0051] Figure 2 is the original sample template during the collection of the handwritten sample data set applied in the present invention;
[0052] Figure 3 is the image basic data document generated during the preprocessing of the handwritten sample data set applied in the present invention;
[0053] Figure 4 is an example program display for modifying and annotating HPC files in the present invention;
[0054] Figure 5 is a partial display of the TPC data set generated by the present invention;
[0055] Figure 6 is the display program interface for the content in the example TPC generated by the present invention;
[0056] Figure 7 is a partial content display of the data set generated by applying the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The technical solutions of the present invention will be further described below in conjunction with the drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered within the protection scope of the present invention.
[0058] The present invention provides an error-traceable image data structure and annotation system for a massive data collection task. The system includes four modules: HPC (hyper-pixel code, super pixel coding) file format, HPC file annotation and modification tool, TPC (traceable pixel code, traceable pixel coding) file format, and target image extraction and format conversion tool. The detailed content or process of each part is as follows:
[0059] 1. HPC File Specification
[0060] The HPC file format is the format closest to the original image sample and is optimized for the roughest and large-scale sample preparation. The HPC file format is divided into two levels: the unit level with one image of the original sample as a unit, and the unit level with one valid target image content in one image of the original sample as a unit. There should be at least one unit in a unit, and at most 65,535 units can be in a unit. The units are independent of each other, and the units within a unit are stored in a cascading arrangement, and the units between units are also stored in a cascading arrangement.
[0061] The unit header stores the length information of the previous unit, the path information of the original image sample corresponding to this unit, the basic information of this image sample, and the number of units within this unit. After the unit header, the information of each unit under this unit is stored serially one by one. The unit header stores the total length information of this unit, the valid bit information indicating whether the unit image is valid, the label information corresponding to the unit image, and the number of rows and columns of the unit image. After the unit header, the information of each pixel is stored in row order, and the high-definition original image is saved intact. The format of its unit header should be as shown in Table 2, where the format, length, and content of the HPCInformation field should be determined according to the specific project content and appropriately modified according to the information required by the project.
[0062] Table 2 HPC Unit Header Format
[0063] Field Name Meaning Lastsize(4B)(unsigned long int) Total length of the content of the previous unit Tag(120B)(char) Path of the HPC file, with spaces filled in the remaining positions HPC Information Basic information of the sample corresponding to the HPC file HPC pagesize(2B)(unsigned short int) Number of valid unit images in the HPC file
[0064] The unit format within the HPC file is as shown in Table 3, where the length and format of the label content should be changed according to the specific project requirements. The pixel information of the image is determined according to the number of color channels of the required image. If it is a grayscale image, then n = 1; if it is an RGB three-channel color image, then n = 3, and so on. In the pixel information of a multi-channel color image, the arrangement order of each channel in each pixel content should be the same.
[0065] Table 3 HPC Unit Storage Format
[0066]
[0067] 2. Annotation and Modification of HPC Files
[0068] The annotation and modification of HPC files should implement this function using a windowed interface.
[0069] (1)Content display of HPC files. Read the content of the image data of each unit stored in the HPC one by one according to the file specification of the HPC, store the number of units contained in the unit and the pointer information of the starting position of the unit header. Then, when reading the unit image data format information, increase the number of read units under this unit. If it is detected that the number of read units is equal to the number of units contained in this unit, clear the number of read units and read the unit header information of the next unit. Fill the read unit image data content into the image matrix in the memory, perform connected component analysis on the image matrix, record the information of each connected component, then display the image, and at the same time record the pointer information of the starting position of this unit in the HPC file.
[0070] (2)Annotation modification of HPC files. In the windowed interface, obtain the unit object and the modification target that need to modify the annotation by the relative position of the window clicked by the left mouse button and the drawing position of each unit image when displaying the image. According to the starting pointer position of the stored unit object and the unit data structure of the HPC file, move the pointer to the position of the valid bit identification field, and overwrite the content to be modified to this field position to achieve the purpose of annotation modification.
[0071] (3)Image content deletion and modification of HPC files. In the windowed interface, determine the selection target and the target unit object involved according to the relative coordinate information feedback by dragging the left mouse button and the coordinate information when the unit image is displayed. Further determine the content of the selected connected component according to the information of the connected component in the target unit object involved. Delete the content of the selected connected component from the connected component information stored in the memory through key interaction. According to the starting pointer position of the target unit object involved and the remaining connected component information in the corresponding memory, write the modified image data into the corresponding HPC file in a covering manner to achieve the purpose of covering and modifying the image content.
[0072] 3. TPC File Specification
[0073] The TPC format is an intermediate format used to finally convert HPC files into a classified output in the png image format. At the same time, for the convenience of tracing back upwards, it should still have the file characteristics that HPC should have. It uses the unit image as the only hierarchical basic unit, and its characteristics should be basically the same as those of HPC files. However, because it is an intermediate format for image content classification conversion, the fields used to demonstrate the feature of tracing back to the source are different from those of HPC files. It does not directly point to the source image path, but points to the ID number of the source HPC corresponding to the unit in the TPC file, and adds the corresponding position field in the source HPC file to facilitate backtracking and modification. In the TPC file, since there is only the unit image information at a single level and no cell level, the file is only arranged one by one in a waterfall manner according to the unit image information.
[0074] The TPC file specifications are shown in Table 4. Among them, the same as the HPC file specifications, the length and format of the label content should be changed according to the specific project requirements. The pixel information of the image is determined according to the number of color channels of the required image. If it is a grayscale image, then n = 1; if it is an RGB three-channel color image, then n = 3, and so on. In the pixel information of a multi-channel color image, the arrangement order of each channel in each pixel content should be the same.
[0075] Table 4 TPC File Specifications
[0076]
[0077] 4. Extraction and Format Conversion of Target Images
[0078] The extraction and format conversion of the target image are divided into three parts: converting the original image sample into an HPC file, converting the HPC file into a TPC file, and converting the TPC file into the target image dataset. Among them, the extraction of the target image is realized during the process of converting the HPC file into a TPC file.
[0079] (1) Convert from the original image samples to HPC files. Read the original images into memory in matrix form, store the reading paths and names of the original images. Set the lastsize field of the first cell in an HPC file to 0 and start recording the length of this cell. Obtain the basic information of the corresponding images and the number of units to be included according to information such as manual input or pre - processed documents, and input them as the cell header into the HPC file, recording the length of this cell header. Read the valid image content in the original samples one by one, store a piece of valid image content into a new image matrix, calculate and confirm the required information in the unit header based on the image matrix information of this unit only and certain pre - processing information, and input the image matrix information as the image data within the unit into the HPC file in a format, increase the length of this unit to the cell length record. Until the start of the next cell, input the recorded length as the data in the cell header and clear it for re - recording.
[0080] (2) Convert from HPC files to TPC files. Obtain the ID information of the HPC file according to the pre - processed document, read the unit image content according to the HPC file specification, record the starting position of this unit's reading, and determine which TPC file the content of this unit should be extracted to according to the label within the unit. Input the ID information of the corresponding source HPC file and the starting position information of this unit's reading as the TPC unit header, and then directly transfer the image data of this unit in the HPC file to the corresponding position in the TPC file.
[0081] (3) Convert from TPC files to target - format files. Generate folders with corresponding label information according to requirements, determine which folder the image content of each unit should be classified into according to the label information owned by each unit in the TPC file, construct an image matrix according to the data in the unit header of this unit, fill the image data within the unit into this image matrix, perform cutting processing on blank areas, reduction processing when the image exceeds the requirements, and enhancement processing on faint handwriting for this image matrix, and then generate image files in the specified format with the processed image matrix to achieve the purpose of the final conversion.
[0082] Example:
[0083] The technical solution of the present invention will be described in detail below in combination with the collection and annotation of the handwritten Chinese character dataset.
[0084] In this embodiment, the format and content of the handwritten Chinese character dataset to be collected are shown in Table 5. The development platform is VisualStudio 2015 and the development language is C++.
[0085] Table 5 A sample table of dataset requirements
[0086]
[0087] The original image samples collected should have the following characteristics:
[0088] 1. Scanned grayscale images with a resolution of 200 dpi or higher;
[0089] 2. It is recommended that the image storage format be a lossless format (such as tiff);
[0090] 3. The image should contain the complete sample layout content.
[0091] 4. The original image format of the handwritten sample should be as Figure 2 shown.
[0092] In the specific implementation process, the original handwritten sample images are processed according to the process Figure 1 shown, and the specific process is as follows:
[0093] 1. Read in image data
[0094] (1) Obtain all the image sample paths and sample names under the sample storage folder path by iterating through the target directory;
[0095] (2) Store the sample paths and sample names in a vector and save them as the information of the files to be processed;
[0096] (3) Obtain the image files according to the information of the files to be processed in the vector, read them into memory for subsequent processing. After processing each image file, remove the information of this image file from the vector and then start processing the next image file.
[0097] 2. Preprocess the image content
[0098] The basic image information is confirmed according to the information obtained by scanning the QR code in the lower right corner of the sample, and a preprocessing document is generated according to the basic image information for subsequent content calls.
[0099] (1) Perform binarization processing on the image information using the Otsu algorithm to obtain a binary image;
[0100] (2) Detect the edge information of the printed layout on the binary image using the hough transform method to obtain the tilt angle value of the image;
[0101] (3) Perform tilt correction and preliminary noise reduction on the original image according to the tilt angle value and the binary image;
[0102] (4) Cut the rectangular image areas at the four corners of the layout in proportion, perform projection analysis on each of them to confirm the area where the QR code is located, and complete the flipping process of the original image according to the position of the area where the QR code is located;
[0103] (5) Use the zxing library to obtain the information within the QR code, and obtain the basic information including the text order corresponding to the image file.
[0104] (6) Output the image path, image name, and basic information of the image sample as a preprocessing document, and the output content example is as Figure 3 shown.
[0105] 3. Conversion of image data into HPC document
[0106] (1) Use the preprocessing document to confirm the number of units in each cell and the corresponding basic information, and generate a cell header according to the HPC file specification.
[0107] (2) Locate the positions of valid image units through layout projection and analysis of connected components.
[0108] (3) Cut the original image according to the located image positions to obtain the unit image content, and fill it into the HPC file according to the HPC file specification.
[0109] 4. Annotation and modification of HPC file
[0110] The annotation and modification tool for HPC file is as Figure 4 shown.
[0111] (1) Read the image content according to the HPC file specification. After reading the unit image content, generate an image matrix based on the image information, perform analysis and search for connected components on the image matrix, and use a two-dimensional blockchain array to store the connected component information of each character in groups by character. Draw the image at the relative specified position in the window corresponding to each unit image according to the connected component information, and display the bounding box of the connected component with blue lines.
[0112] (2) Change the color of the displayed image according to the representation of valid bits. Red is for deleted characters (marked with a dotted box in the printed matter), and light blue is for modified characters (marked with an elliptical dotted box in the printed matter).
[0113] (3) When the user presses the left mouse button, record the mouse position LBD. When the user drags the mouse, draw a black rectangular selection box in real time according to the mouse position and LBD to facilitate the user to identify their selection range. When the user releases the left mouse button, obtain the mouse position LBU. The position information of LBD and LBU are both relative position information relative to the top left vertex of the window. According to the position information of LBD and LBU and the relative position information of the unit image, it is possible to obtain which connected components are selected, and change the color of the bounding boxes of these connected components to red lines for display (replaced by a dotted box in the printed matter).
[0114] (4) When the user presses the delete button on the premise of selecting the connected component, the selected array is thrown out from the blockchain array, and the connected component information left after throwing is integrated into a unit image matrix. The image data and the modified flag value are overwritten at the corresponding position of the original HPC file according to the stored page pointer position and the corresponding character sequence information, and the interface display is updated to complete an image modification and identification.
[0115] 5. Conversion of HPC file format to TPC file format
[0116] (1) The program reads the HPC file in binary form to obtain the HPC file ID;
[0117] (2) Open a certain number of TPC files to speed up the program processing speed;
[0118] (3) After reading the HPC file information, the program analyzes and operates on the unit image data information one by one. The corresponding number is confirmed through the label information in the unit image data, and the TPC file number into which it should be classified. The unit image information that meets the TPC format requirements is stored in the data stream of the corresponding port and written into the corresponding TPC file.
[0119] 6. Conversion of TPC file format to target format file
[0120] (1) The program reads the TPC file in binary form and reads the entire content of the TPC file into memory at one time;
[0121] (2) Analyze and process the image data information inside the TPC file one by one, and convert the corresponding image data in the TPC file into an image matrix.
[0122] (3) Perform post-processing operations on the image matrix. First, cut the blank part of the image, and use the projection method to confirm the cutting row and column numbers. If the difference between the final row and the starting row is negative, it means that the picture has no content and is directly discarded;
[0123] (4) After cutting the image into a smaller image according to the confirmed row and column numbers, compare it with 100×100. If the number of rows / columns is greater than 100, convert the number of rows / columns to 100 and scale the number of columns / rows proportionally. The image is scaled to the specified size through the scaling matrix;
[0124] (5) Finally, count the effective pixels in the image content. If its average value is less than the specified threshold, it is considered that the image needs handwriting enhancement, and the adaptive equalization enhancement method is used to thicken the handwriting to achieve a better sample effect.
[0125] Figure 5 For TPC dataset display, Figure 6For the display of TPC content, the invalid samples are shown in red and not included in the delivery quantity. Figure 7 It is a partial specified PNG format output example.
[0126] The above system was experimented on a total of more than one hundred thousand original image samples, processed a total of more than twenty million unit images, and finally obtained a dataset of 16,300,020 effective handwritten character images containing 21,003 Chinese characters under GBK encoding and 94 displayable ASCII symbols.
[0127] The technology involved in the present invention is illustrated and verified by taking the collection and annotation of a Chinese character handwritten dataset as an example, but is not limited to the collection and annotation of a Chinese character handwritten dataset, and can be applied to the collection and annotation of a large amount of image data that requires manual processing, with good results.
Claims
1. An image data structure with error traceability and an annotation method, characterized in that The method comprises the following steps: Step 1: Read in image data; Step 2: Image content preprocessing; Step 3: Convert image data into HPC file format (1) using the preprocessing document to confirm the number of units and corresponding basic information in each unit, and generating a unit header according to the HPC file specification. The HPC file format is divided into two levels, one with an image of the original sample as the unit level of the unit, and one with a valid target image content in an image of the original sample as the unit level of the unit; The unit header stores the length information of the previous unit, the original image sample path information corresponding to the unit, the basic information of the image sample and the number of units in the unit; After the unit header, the information of each unit under the unit is stored serially one by one; (2) Locate the position of effective image units through layout projection and analysis of connected domains; (3) Segment the original image according to the located image position, obtain the unit image content, and fill it into the HPC file according to the HPC file specification; Step 4: Labeling and modifying HPC files (1) Display of HPC file content: The HPC file modification tool reads the HPC file in accordance with the standard format, generates an image matrix based on the read unit image content, and segments the image using a connected domain search method. The segmented connected domain information is stored in a two-dimensional blockchain array in groups of words for subsequent manual deletion. When displayed, a blue rectangular frame is used to display the connected domain outer frame for easy identification and operation; (2) HPC file annotation modification: The user selects the connected domain that is obviously the noise part by left-clicking the mouse and dragging it. The selected connected domain is indicated by turning the outer border red. When the delete button is pressed, the selected connected domain on the interface is deleted, and the deleted content is synchronously updated to the HPC file, so as to complete the modification of the HPC file image information at the same time; Step 5: Convert HPC file format to TPC file format (1) Based on the ID of the HPC file to be read, a certain number of TPC files that may be used are opened simultaneously in a multi-threaded manner to speed up the processing; (2) Read the HPC file quickly in binary mode, and according to the Chinese character label of the unit image information, divert the unit image information to the thread that has opened the corresponding TPC file, and write the unit image information into the TPC file that is consistent with its own label; Step 6: Convert the TPC file format to the target format file.
2. The error-retraceable image data structure and annotation method according to claim 1, characterized in that The units are independent of each other. There is at least one unit in a unit and a maximum of 65535 units in a unit. The units in a unit are stored in a waterfall arrangement and between units are stored in a waterfall arrangement.
3. The image data structure and annotation method with error traceability according to claim 1, characterized in that The specific steps of step one are as follows: (1) Obtain all image sample paths and sample names under the sample storage folder path by iterating the target directory; (2) Store the sample path and sample name in a vector as the file information to be processed; (3) Obtain the image file according to the information of the file to be processed in the vector, read it into the memory for subsequent processing. After processing each image file, remove the information of this image file from the vector and then start processing the next image file.
4. The image data structure and annotation method with error traceability according to claim 1, characterized in that The specific steps of step two are as follows: (1) Perform binarization processing on the image information using the Otsu algorithm to obtain a binary image; (2) Detect the edge information of the printed layout by means of the hough transform on the binary image, and obtain the tilt angle value of the image; (3) Perform tilt correction and preliminary noise reduction on the original image according to the tilt angle value and the binary image; (4) Cut the rectangular image areas at the four corners of the layout in proportion, perform projection analysis respectively to confirm the area where the QR code is located, and complete the flipping process of the original image according to the position of the area where the QR code is located; (5) Use the zxing library to obtain the information in the QR code, and obtain the basic information including the text order corresponding to this image file; (6) Output the image path, image name, and basic information of the image sample as a preprocessing document.
5. The error-retraceable image data structure and annotation method according to claim 1, characterized in that The specific steps of step four are as follows: (1) The conversion program iterates through the target directory, and each time it stores an entire TPC file for processing until all TPC files in the directory are processed; (2) Restore the unit image information in the TPC file to an image matrix; (3) Check the image matrix, remove the blank images, and perform cutting processing on the blank areas around the normal images; (4) Check the image size. If the image is larger than the required size, scale the image down proportionally to the required size; (5) Make a statistics of the handwriting density for the valid pixels. If the density is less than the threshold, thicken the handwriting by means of adaptive equalization enhancement; (6) Output the processed image matrix in the target format to the corresponding folder.
Citation Information
Patent Citations
Sensitive file steganography and tracing method based on deep adversarial network
CN112270638A