An efficient hand-character mixed object detection method
Data sets were prepared through image synthesis and data enhancement, and a deep neural network model was designed to solve the problems of the complexity and real-timeness of the existing hand-text hybrid object detection method, and achieved efficient and concise hybrid object detection effect.
Patent Information
- Application Number
- CN202111620882.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The existing manual-text hybrid object detection method has a complex process, is not very real-time, and has high requirements for object detection algorithms.
Image synthesis and data enhancement are used to prepare hand-text mixed target data sets, design deep neural network models, and train positive and negative samples through balanced division to reduce the confidence of the model's prediction box for other areas of the image, simplifying the detection process.
It realizes efficient manual-text mixed object detection, simplifies the detection process, reduces the calculation steps, improves the real-timeness of the detection, and is suitable for mixed object detection in various real scenarios.
Smart Images

Figure CN114359885B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to an efficient hand-text hybrid object detection method. Background Art
[0002] Object detection is a rapidly developing branch in the field of artificial intelligence. The task of conventional object detection is to find all the objects of interest in an image and determine their categories and positions, which is one of the core problems in the field of computer vision.
[0003] In an actual reading scenario, when encountering obscure text words, verification and learning are required. Combining with the methods in the field of artificial intelligence, object detection of finger markings on the read text can complete the verification and learning of these words. If the target text is not marked, current text detection methods will detect all the text on the current reading page. Therefore, using a columnar object such as a finger or a pen to point out or mark the target text, making it a hand-text hybrid object detection, can meet the requirements in the real scenario.
[0004] Existing methods for object detection of the finger-like object and the text pointed by the finger-like object are all based on a deep learning network to separately detect the tip of the finger-like object and the text on the same image, and then compare the coordinate information of the two detected results, and then synchronize and overlap the coordinate information of the two through coordinate affine transformation, so as to detect and frame the text pointed by the finger in the original input image; not only is the process complex and cumbersome, but also the requirements for the object detection algorithm are relatively high, resulting in poor real-time performance of detection in the actual scenario. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: to provide an efficient hand-text hybrid object detection method for simultaneously detecting columnar objects and target text in an image.
[0006] The technical solution adopted by the present invention to solve the above technical problem is: an efficient hand-text hybrid object detection method, comprising the following steps:
[0007] S1: Prepare a hand-text hybrid object dataset including columnar objects such as fingers and pens pointing to text by means of image synthesis and data augmentation, divide the positive and negative samples of the images in the hand-text hybrid object dataset, label the ground truth boxes for the target regions of the images, and record the coordinate information of the ground truth boxes;
[0008] S2: Design an algorithm suitable for hand-text hybrid object detection to build a deep neural network model; use the intersection over union (IoU) between the predicted boxes output by the deep neural network model and the ground truth boxes as a preset threshold for balancing the division of training positive and negative samples;
[0009] S3: Iteratively train a deep neural network model using a hand-text mixed target dataset, extract features from the image data, and regress according to the given object detection principle and preset threshold to obtain mixed target detection candidate boxes close to the ground truth boxes.
[0010] S4: Use the deep neural network model with adjusted parameter weights after training to detect hand-text mixed targets in the hand-text validation set and test set in the real reading scenario.
[0011] According to the above solution, in step S1, the specific steps are as follows:
[0012] S11: Use an industrial camera or a camera to capture and collect text images in different physical reading scenarios, or select text images that meet the target requirements from existing text image datasets.
[0013] S12: Use the finger images in the Fingers - finger dataset, or take real-scene pictures of column-like objects including pens and fingers; the background of the column-like object images is a simple background for highlighting and making the column-like object images easy to identify.
[0014] S13: Synthesize the column-like object images into the text images to obtain synthesized hand-text images; the synthesized hand-text images include text and the feature specifications of the column-like objects pointing to the text, and represent the position of the text pointed to by the column-like objects in the image.
[0015] S14: Divide the positive and negative samples of the synthesized hand-text images; the positive samples are the image regions including hand-text mixed targets; the negative samples are the image regions including interference objects and not including hand-text mixed targets.
[0016] S15: Label the ground truth boxes for appropriate-sized regions of the synthesized hand-text images. The labeling methods include: automatically generating and recording the coordinate position information of the ground truth boxes to be labeled according to the location information of the synthesized regions where the text images are modified; or manually labeling the coordinate position information of the ground truth boxes using the labeling tool LabelImg.
[0017] S16: Use data augmentation methods including rotation angle, scaling, stretching, cropping, and brightness enhancement to expand the number of synthesized hand-text images.
[0018] Furthermore, in step S1, the following steps are also included:
[0019] S17: Divide the prepared hand-text mixed target dataset into a training set, a validation set, and a test set.
[0020] According to the above solution, in step S1, the column-like object images in the synthetic hand-written text images include fingers of various skin colors, nail styles, hand postures, finger orientations, or multiple types of pens; the text images in the synthetic hand-written text images do not repeat.
[0021] According to the above solution, in step S2, the specific steps are as follows:
[0022] S21: Use the One Stage object detection method for feature extraction, object classification, and localization regression of the hand-written text hybrid object; build a deep neural network model including an image tiling layer, a feature extraction layer, and a sequence output layer; the image tiling layer is used to divide the input image into several small blocks and then flatten them into a sequence; the feature extraction layer is used to input the sequence after fusing the position encoding into the encoder-decoder combination to extract features; the sequence output layer is used to output the coordinate information of the prediction box belonging to the hand-written text hybrid object in the image and the hybrid object confidence.
[0023] S22: Use the intersection over union of the area of the prediction box and the ground truth box as the preset threshold to balance the division of training positive and negative samples.
[0024] Further, in step S21, the hybrid object confidence is determined by calculating the intersection over union of the area of the prediction box and the ground truth box, and the hybrid object confidence less than the preset threshold is eliminated by the non-maximum suppression method.
[0025] Further, in step S22, take the intersection over union of the area of the prediction box and the ground truth box equal to 0.7 as the preset threshold.
[0026] Further, in step S3, the specific steps are as follows:
[0027] S31: Select a specific network for pre-training and save the pre-trained weights; perform transfer learning on the deep neural network model using the pre-trained network weights.
[0028] S32: Input the hand-written text hybrid object dataset into the deep neural network model for training.
[0029] S33: The image tiling layer divides the input image into several small blocks and then flattens them into a sequence.
[0030] S34: The feature extraction layer inputs the sequence after fusing the position encoding into the encoder-decoder combination to extract features.
[0031] S35: The sequence output layer outputs the coordinate information of the prediction box belonging to the hand-written text hybrid object in the image and the hybrid object confidence.
[0032] S36: Perform autoregression on the sequence composed of the coordinate information of the predicted bounding boxes and the confidence of the mixed objects to obtain candidate mixed object detection bounding boxes that are close to the ground truth bounding boxes;
[0033] S37: Determine whether the detection objective is met. If so, output the prediction sequence; otherwise, execute step S36 for iteration until the objective is met; the number of iterations is set according to the size of the class hand-text mixed object dataset;
[0034] S38: Adjust and optimize the parameters including the number of network layers, learning rate, number of heads of the multi-head self-attention mechanism, dropout rate, and batch size, and save the parameter weights with the best mixed object detection effect obtained by training the deep neural network model;
[0035] S39: After the deep neural network model is iteratively trained using the class hand-text mixed object dataset, verify and test the deep neural network model, and adjust the parameter weights according to the verification and test results.
[0036] Further, in step S37, the Softmax function is combined with the cross-entropy loss as the objective function, which is used to calculate the coordinate information error between the predicted bounding boxes and the ground truth bounding boxes to obtain the error of the output sequence, and is also used to calculate the confidence score of whether it is a mixed object region.
[0037] According to the above solution, in step S4, the detection objects of the class hand-text mixed object network model include class hand-text mixed objects based on images or videos that have no real-time requirements.
[0038] The beneficial effects of the present invention are as follows:
[0039] 1. An efficient class hand-text mixed object detection method of the present invention uses an image dataset with class cylinder objects such as fingers and pens and target texts to train the mixed object detection model. By evenly dividing the positive and negative samples for training, the deep network model is only interested in the regions in the image that simultaneously contain class finger tips and text words, reducing the confidence of the predicted bounding boxes generated by the deep network model in other regions of the image. Thus, the model ignores other text words on the current page in the real reading scenario and only detects the text pointed to by the finger, realizing the function of simultaneously detecting class cylinder objects and target texts in the image.
[0040] 2. The present invention avoids the process of two - stage object detection and coordinate transformation for column - like objects such as fingers and pens and text in images, avoids the affine transformation between the coordinates of finger - like fingertips and text words on the image, simplifies the detection process of the text area pointed to by column - like objects such as fingers and pens; compared with existing object detection algorithms, the present invention performs fewer object detections to achieve the same effect, has a more concise object detection logic, reduces the overall number of calculation steps required for the hybrid object detection method, simplifies the detection idea, optimizes the real - time performance of detection, has low requirements for the performance of the object detection algorithm, and has important guiding significance for realizing efficient auxiliary reading application fields.
[0041] 3. The present invention is applicable to hybrid object detection in a variety of real - world scenarios, especially suitable for the detection of hand - text hybrid objects in real reading scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the flowchart of an embodiment of the present invention.
[0043] Figure 2 is the flowchart for preparing the dataset of an embodiment of the present invention.
[0044] Figure 3 is the flowchart for model processing of an embodiment of the present invention.
[0045] Figure 4 is the flowchart for building and training the deep network model of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0046] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0047] See Figure 1 , an efficient hand - text hybrid object detection method according to an embodiment of the present invention, includes the following steps:
[0048] S1: Prepare a hybrid object detection dataset. The specific steps are as follows: Prepare an image dataset in which column - like objects such as fingers and pens point to text by means of image synthesis and data augmentation, divide the positive and negative samples of the dataset images, label the target areas of the images, and record the real - box coordinate information;
[0049] The specific steps for preparing the hand - text hybrid object dataset include:
[0050] S11: Prepare text images; Use an industrial camera or a camera to capture text images in different physical reading scenarios, or select text images that meet the target requirements from existing text image datasets; The text content in the prepared images should be clear and rich.
[0051] S12: Prepare a column-like object image including fingers and pens; use finger images in the existing Fingers - Finger dataset, or images of pens, fingers, etc. taken in real scenes; the background of the column-like object image should be simple, and the instance target should be prominent and easy to identify.
[0052] S13: Synthesize a hand - text image; use the PIL toolkit to synthesize column-like objects such as fingers and pens in a pre-processed simple background into the text image to obtain a synthesized image; the information of the synthesized image contains a large amount of text and the feature specification of the hand-like object pointing to the text, and determines the position of the indicated text in the image, that is, the text pointed to by the column-like object is a specific word.
[0053] S14: Divide the positive and negative samples of the images in the hand - text mixed target image dataset; determine the division of positive and negative samples in the dataset prepared by the above method. The sample images include: positive sample images, which include hand - text mixed targets, that is, the image area containing the hand - text mixed target is the positive sample of the dataset; negative sample images, which include interference objects. The interference objects do not belong to the hand - text mixed target, that is, the image area without the hand - text mixed target is the negative sample of the dataset, such as blank backgrounds, single texts, single column-like objects, etc.
[0054] S15: Perform Ground - Truth bounding box annotation on an appropriate size area of the synthesized dataset image. There are two annotation methods, including: one is to automatically generate and record the coordinate position information of the bounding box to be annotated according to the position information of the synthesized area modified by the text image; the other is to manually annotate the coordinate position information of the bounding box using the annotation tool LabelImg.
[0055] S16: Use data augmentation methods to expand the number of synthesized images in the prepared hand - text mixed target detection dataset, so that the pre-training and training of the deep network model are completed on a large amount of image data. Specifically, through different image processing methods, such as rotating 5 degrees, scaling, stretching, cropping, brightness enhancement, etc., the number of images in the dataset is expanded.
[0056] For the synthesized image data in the embodiments of the present invention, the styles of the column-like objects in the image are diverse. It cannot be the same pen or the same hand posture all the time. The skin color and nail styles of the fingers, hand postures, finger orientations, or the types of pens in the dataset images should be as rich as possible. Similarly, the text information in the prepared dataset images should not be the same, and the content of the text page should also be as rich as possible.
[0057] The flowchart for completely establishing the hand - text mixed target dataset in the embodiments of the present invention is as Figure 2As shown in the figure. In step S1, a scientific image synthesis method is used to prepare a certain number of effective datasets containing columnar objects such as fingers and pens and text images, and data augmentation is used to obtain a larger number of mixed target images, so as to form the hand-text mixed target detection dataset described above. After preparing the dataset of the hand-text mixed target, the training set, validation set, and test set are scientifically and reasonably divided.
[0058] S2: Design an algorithm suitable for hand-text mixed target detection, build a deep neural network model, and perform feature extraction, target classification, and localization regression of hand-text mixed targets on the basis of the OneStage target detection method; and preset a threshold for balancing the division of training positive and negative samples, where the threshold is the intersection over union of the area between the prediction box generated by the deep network and the ground truth box; the specific steps for building the deep neural network model include:
[0059] S21: Design an algorithm suitable for hand-text mixed target detection according to the scenario requirements. The process of detecting hand-text mixed targets is approximated as a binary classification problem of judging whether the target is a hand-text mixed target through the confidence of the hand-text mixed target.
[0060] S22: The regression problem in traditional target detection is specifically the autoregression of the sequence composed of the output prediction box coordinate information and the mixed target confidence in the embodiment of the present invention; the confidence score reflects the reliability of whether the prediction box contains the target by the novel neural network model and the accuracy of the prediction box; the present invention determines the mixed target confidence by calculating the intersection over union of the area between the prediction box and the ground truth box, and eliminates the mixed target confidence less than the preset threshold by the non-maximum suppression method.
[0061] S23: For any deep neural network target detection algorithm based on supervised learning, positive and negative samples need to be set for the training of the deep network so that the network can learn features from them. The present invention presets the intersection over union of the area between the ground truth box marked in the prepared dataset image and the prediction box generated by the designed deep neural network as the threshold for balancing the division of training positive and negative samples. In the embodiment of the present invention, the value of IoU (intersect over union) is taken as 0.7 as the threshold for setting positive and negative samples.
[0062] See Figure 3, the present invention provides a hybrid object detection method based on image-to-sequence, which uses a pre-trained deep neural network model to autonomously detect hand-text hybrid objects through an autoregressive method. The deep neural network model includes a combination of a convolutional neural network and an encoder-decoder connected in parallel or serially. This network framework generally has excellent ability to fuse local and global features of images and shows more significant advantages in dealing with the hand-text hybrid object detection task. The deep neural network of the present invention includes three parts: an image tiling layer, a feature extraction layer, and a sequence output layer; the image tiling layer is used to divide the input image into several small blocks and then flatten them into a sequence; the feature extraction layer is used to input the sequence after fusing the position encoding into the encoder-decoder combination for feature extraction; the sequence output layer is used to output the coordinate information and confidence of the prediction boxes belonging to the hand-text hybrid objects in the image.
[0063] S3: Use the designed object detection model to perform iterative training on the synthetically prepared hand-text image dataset, extract features from the image data, and regress to obtain candidate hybrid object detection boxes close to the ground truth boxes according to the given object detection principle and threshold;
[0064] See Figure 4 , the specific steps for training the hand-text hybrid object detection model include:
[0065] S31: Before training the hand-text hybrid object detection model, select a specific network for pre-training, save the pre-trained weights, and then use the pre-trained network weights to perform transfer learning on the built network model.
[0066] S32: Input the hand-text hybrid object dataset into the object detection network model based on deep learning for training;
[0067] S33: Adjust the optimization parameters and save the parameter weights with the best hybrid object detection effect obtained from the network model training.
[0068] During the training process, the Softmax function combined with the cross-entropy loss is used as the objective function for the entire network architecture to calculate the error of the output sequence, that is, the coordinate information error between the prediction box and the ground truth box, and the confidence score of whether it is a hybrid object area.
[0069] In this step, after the framework design suitable for hand-text hybrid object detection is completed, some hyperparameters in the deep network are adjusted and set, such as the number of network layers, learning rate, number of heads of the multi-head self-attention mechanism, dropout rate, batch size, etc.; the number of iterations is adjusted according to the size of the dataset.
[0070] After the deep network model designed in the embodiment of the present invention is iteratively trained several times on the prepared hand-written mixed target dataset, the model is verified and tested. When the detection effect has no obvious improvement compared with the ordinary target detection algorithm, the network model built by the designed target detection algorithm should be fine-tuned, and step S3 is repeated until the hand-written mixed target detection effect is improved.
[0071] S4: Use the network model after the training is completed to detect the hand-written mixed targets in the hand-written verification set and the test set.
[0072] The present invention uses the parameter weights saved in the above step 3 to detect the hand-written mixed targets based on images or videos in the real reading scenario.
[0073] The hand-written mixed target detection method of the present invention can also be applied to the video-based mixed target detection with less strict real-time requirements.
[0074] The hand-written mixed target detection method provided by the embodiment of the present invention puts forward a creative idea on the basis of target detection to meet the application needs in the real reading text scenario.
[0075] Compared with the ordinary target detection algorithm, the embodiment of the present invention performs fewer target detections to achieve the same effect, the target detection logic is simpler than the conventional target detection logic, and at the same time, the calculation steps required for the overall mixed target detection method are reduced.
[0076] The embodiment of the present invention generalizes the hand-written mixed target method proposed by the present invention, and transplants the efficient mixed target detection method and technology obtained in this example to other applicable fields as a prerequisite for downstream application development.
[0077] The above embodiments are only used to illustrate the design ideas and features of the present invention, and their purposes are to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.
Claims
1. An efficient hand-text hybrid object detection method, characterized in that: It includes the following steps: S1: Prepare a hand-text hybrid object dataset including fingers and pens pointing to text by means of image synthesis and data augmentation, divide the positive and negative samples of the images in the hand-text hybrid object dataset, annotate the ground truth boxes for the target regions of the images, and record the coordinate information of the ground truth boxes; S2: Design an algorithm suitable for hand-text hybrid object detection to build a deep neural network model; use the intersection over union (IoU) between the predicted boxes and the ground truth boxes output by the deep neural network model as a preset threshold to balance the division of training positive and negative samples; S3: Iteratively train the deep neural network model using the hand-text hybrid object dataset, extract features from the image data, and regress according to the given object detection principle and preset threshold to obtain hybrid object detection candidate boxes close to the ground truth boxes; S4: Use the deep neural network model with adjusted parameter weights after training to detect hand-text hybrid objects in the hand-text validation set and test set in the real reading scenario.
2. An efficient hand-text hybrid object detection method according to claim 1, characterized in that: In the step S1, the specific steps are: S11: Use an industrial camera or a camera to capture text images in different physical reading scenarios, or select text images that meet the target requirements from the existing text image datasets; S12: Use the finger images in the Fingers - finger dataset, or capture real-scene images of columnar objects including pens and fingers; the background of the columnar object images is a simple background to highlight and make the columnar object images easy to identify; S13: Synthesize the columnar object images into the text images to obtain synthesized hand-text images; the synthesized hand-text images include text and the feature specifications of the columnar objects pointing to the text, and represent the position of the text pointed to by the columnar objects in the image; S14: Divide the positive and negative samples of the synthesized hand-text images; the positive samples are the image regions including hand-text hybrid objects; the negative samples are the image regions including interference objects and not including hand-text hybrid objects; S15: Annotate the ground truth boxes for appropriate-sized regions of the synthesized hand-text images. The annotation method includes: automatically generating and recording the coordinate position information of the ground truth boxes to be annotated according to the location information of the synthesized regions modified from the text images; or manually annotating the coordinate position information of the ground truth boxes using the annotation tool LabelImg; S16: Use data augmentation methods including rotation angle, scaling, stretching, cropping, and brightness enhancement to expand the number of synthesized hand-text images.
3. An efficient hand-text hybrid object detection method according to claim 2, characterized in that: In the step S1, the following steps are further included: S17: Divide the prepared hand-text hybrid object dataset into a training set, a validation set, and a test set.
4. An efficient hand-text hybrid object detection method according to claim 1, characterized in that: In the step S1 described above, the column-like object images in the synthetic hand-script images include fingers of various skin colors, nail styles, hand postures, finger orientations, or multiple types of pens; the text images in the synthetic hand-script images do not repeat.
5. An efficient hand-script hybrid object detection method according to claim 1, characterized in that: In the step S2 described above, the specific steps are: S21: Use the One Stage object detection method to perform feature extraction, object classification, and localization regression of the hand-script hybrid object; build a deep neural network model including an image tiling layer, a feature extraction layer, and a sequence output layer; the image tiling layer is used to divide the input image into several small blocks and then flatten them into a sequence; the feature extraction layer is used to input the sequence after fusing the position encoding into the encoder-decoder combination to extract features; The sequence output layer is used to output the coordinate information of the prediction box belonging to the hand-script hybrid object in the image and the hybrid object confidence. S22: Use the intersection over union of the area of the prediction box and the ground truth box as a preset threshold to balance the division of training positive and negative samples.
6. An efficient hand-script hybrid object detection method according to claim 5, characterized in that: In the step S21 described above, the hybrid object confidence is determined by calculating the intersection over union of the area of the prediction box and the ground truth box, and the hybrid object confidence less than the preset threshold is eliminated by the non-maximum suppression method.
7. An efficient hand-script hybrid object detection method according to claim 5, characterized in that: In the step S22 described above, take the intersection over union of the area of the prediction box and the ground truth box equal to 0.7 as the preset threshold.
8. An efficient hand-script hybrid object detection method according to claim 5, characterized in that: In the step S3 described above, the specific steps are: S31: Select a specific network for pre-training and save the pre-trained weights; use the pre-trained network weights for transfer learning of the deep neural network model; S32: Input the hand-script hybrid object dataset into the deep neural network model for training; S33: The image tiling layer divides the input image into several small blocks and then flattens them into a sequence; S34: The feature extraction layer inputs the sequence after fusing the position encoding into the encoder-decoder combination to extract features; S35: The sequence output layer outputs the coordinate information of the prediction box belonging to the hand-script hybrid object in the image and the hybrid object confidence; S36: Perform autoregression on the sequence composed of the coordinate information of the prediction box and the hybrid object confidence to obtain a hybrid object detection candidate box close to the ground truth box; S37: Judge whether the detection target is satisfied. If so, output the prediction sequence; if not, execute step S36 for iteration until the target is satisfied; the number of iterations is set according to the size of the hand-script hybrid object dataset; S38: Adjust and optimize the parameters including the number of network layers, learning rate, number of heads of the multi-head self-attention mechanism, dropout rate, and batch size, and save the parameter weights with the best hybrid object detection effect obtained by training the deep neural network model. S39: After the iterative training of the deep neural network model using the hand-text hybrid target dataset is completed, verify and test the deep neural network model, and adjust the parameter weights according to the verification and test results.
9. An efficient hand-text hybrid target detection method according to claim 8, wherein: In step S37, the Softmax function is combined with the cross-entropy loss as the objective function, which is used to calculate the coordinate information error between the predicted box and the ground truth box to obtain the error of the output sequence, and is also used to calculate the confidence score of whether it is a hybrid target area.
10. An efficient hand-text hybrid target detection method according to claim 1, wherein: In step S4, the detection objects of the hand-text hybrid target network model include hand-text hybrid targets based on images or videos that have no real-time requirements.
Citation Information
Patent Citations
Text-independent end-to-end handwriting recognition method based on deep learning
CN105893968A
Cross-border national culture text classification method based on knowledge representation
CN111444343A