Object detection and segmentation for inking applications

Through computerized methods and neural network technology, digital ink strokes are detected and segmented, and the problems of traditional technology being sensitive to stroke order and inaccurate recognition are solved, and the accurate identification and grouping of digital ink strokes are realized, and the recognition of different ink font sizes is supported.

CN113711232BActive Publication Date: 2025-05-13MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080022660.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-20
Filing Date
2020-03-10
Publication Date
2025-05-13
Estimated Expiration
2040-03-10

AI Technical Summary

Technical Problem

Traditional handwriting recognition and stroke analysis techniques are sensitive to stroke order, and are inaccurate in the large-size handwriting content and different stroke sampling methods, and the grouping and classification results are not optimal.

Method used

Using computerized methods, digital ink strokes are detected and segmented, including writing strokes and painting strokes, using U-net and Yolo convolutional neural networks for segmentation and detection, reducing the sensitivity to stroke sorting, and realizing text recognition of different ink font sizes.

Benefits of technology

Accurate recognition and grouping of digital ink strokes is realized, the sensitivity to stroke order is reduced, the recognition of different ink font sizes is supported, and the stroke data of different sampling methods is agnostic, improving the accuracy and flexibility of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113711232B_ABST
    Figure CN113711232B_ABST
Patent Text Reader

Abstract

An ink parsing system receives ink strokes at an inking device input and plots the received ink strokes into an image in pixel space. Written strokes are detected in the image and marked. Pixels corresponding to the marked written strokes are removed from the image. Painted strokes in the image with the removed pixels are detected and marked. Written objects and painted objects corresponding to the marked written strokes and the marked painted strokes, respectively, are output. A digital ink parsing pipeline with accurate ink stroke detection and segmentation is thereby provided.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Digital inking tools allow users to create digital content such as diagrams, flow charts, notes, etc. Recognizing content created with digital ink can facilitate increased user productivity. To this end, handwriting recognition and stroke analysis are common digital inking functions, where an image or drawing is interpreted to extract specific categories of information, such as the presence and location of specific characters or shapes.

[0002] However, conventional handwriting recognition and stroke analysis have many limitations, including being sensitive to the order of strokes input, so that changes in the stroke order reduce recognition accuracy. In addition, if the size of the handwriting is large, the recognition is inaccurate. Grouping and classification are also performed using different neural networks trained separately, so that combining the results does not produce the best final result. Summary of the invention

[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0004] A computerized method for digital ink parsing includes receiving ink strokes at an inking device input, drawing the received ink strokes into an image in pixel space, and detecting written strokes in the image and marking the written strokes. The computerized method also includes removing pixels corresponding to the marked written strokes from the image and detecting painted strokes in the image with the removed pixels and marking the painted strokes. The method also includes outputting written objects and painted objects corresponding to the marked written strokes and the marked painted strokes, respectively.

[0005] As many of the attendant features become better understood by reference to the following detailed description considered in conjunction with the accompanying drawings, these features will be more readily appreciated. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present description will be better understood from the following detailed description read in light of the accompanying drawings, in which:

[0007] Figure 1 is an inked input that may be parsed according to an embodiment;

[0008] Figure 2 is a block diagram of an ink parsing engine with a parsing pipeline according to an embodiment;

[0009] Figure 3 is a block diagram of a writing detector according to an embodiment;

[0010] Figure 4illustrates a neural network according to an embodiment;

[0011] Figure 5 illustrates a painting detection process according to an embodiment;

[0012] Figure 6 illustrates a bounding box used in a painting detection process according to an embodiment;

[0013] Figure 7 illustrates parsing using a parse tree according to an embodiment;

[0014] Figure 8 is a flow chart of a process for detection and segmentation of an inking environment according to an embodiment;

[0015] Fig. 9 is another flow chart of a process for detecting and segmenting an inking environment according to an embodiment; and

[0016] Fig.10 is a block diagram of an example computing environment suitable for implementing some of the various examples disclosed herein.

[0017] Throughout the drawings, corresponding reference numerals indicate corresponding parts. In the drawings, the system is illustrated as a schematic diagram. The drawings may not be drawn to scale. DETAILED DESCRIPTION

[0018] The computing devices and methods described herein are configured to perform detection and segmentation of user input, particularly digital inked input. Object detection and segmentation techniques are implemented in a pipeline of a machine learning engine configured as a parsing pipeline. Using the configured parsing pipeline, the content of inked input, such as a chart, can be accurately identified.

[0019] In one example, a convolutional neural network architecture (e.g., U-net) is trained to segment handwriting strokes, followed by a convolutional neural network (e.g., You Look Only Once (Yolo)) that detects and classifies drawn objects. In some examples, the conversion of Yolo recognition results from image space to stroke space utilizes boosted decision trees to provide accurate stroke mapping. The pipeline allows for improved grouping and classification of digital ink strokes into strokes belonging to shapes and strokes belonging to text. In addition, the pipeline is more resilient to stroke order during recognition.

[0020] Although described in some examples with reference to U-net and Yolo, aspects of the disclosure may operate with any other neural network having the properties described herein to support the disclosed functionality.

[0021] The present disclosure thus provides a pipeline with the ability to parse free-form ink drawings (such as multiple font sizes and random stroke order drawings in a page). Therefore, the ink parsing engine is configured to perform digital ink analysis, which allows receiving a wide set of strokes, using classification techniques to segment the strokes into more fine-grained domains (e.g., text, shapes, connectors, etc.), and running each subgroup through a simpler classification algorithm adjusted for the domain. The ink parsing engine has reduced sensitivity to ink stroke ordering, enables text recognition with different ink font sizes, and is agnostic to different sampling methods of incoming ink stroke data. In this way, when the processor is programmed to perform the operations described herein, the processor is used in an unconventional manner and allows for faster and / or more accurate recognition of different inputs created with digital ink. Thereby providing a more efficient parsing process that improves the user experience.

[0022] Various examples are implemented as a digital ink conversion background application programming interface (API) that can be called by a specific application to parse inking input to identify different types of ink stroke input. It should be noted that while the examples described herein relate to inking applications, other applications such as optical character recognition (OCR) applications can also benefit from utilizing an ink parsing engine that separates ink detection into different components.

[0023] Figure 1 Illustrated is inking input 100 that may be parsed according to various examples. The inking input includes writing input 102, which in this example corresponds to words, and drawing input 104, which in this example corresponds to lines and boxes. In various examples, pixel data represents writing input 102 and drawing input 104, which correspond to inking strokes written by a human using a stylus. In one example, the parsing engine is implemented as part of an inking application, where text written by a human using a stylus is converted to text characters, and sketches drawn by a human using a stylus are converted to drawing objects by drawing all inking strokes into image (pixel) space. Output from the various examples may also be useful for further processing, such as in other applications operated by the computing device 1000, which may be useful for processing the inking input. Fig.10 Describe in more detail.

[0024] Figure 2 An ink parsing engine 220 is illustrated that is operable to parse an inked input, such as inked input 100, using a parsing pipeline 200. The components in the parsing pipeline 200 are arranged in the particular order shown to perform parsing to separate ink detection into different components that allow for more accurate detection and segmentation of the inked input. That is, the order of operations can affect the accuracy of the calculations. In the illustrated example, writing detection is performed before drawing detection in order to provide more accurate parsing for inked applications.

[0025] Parsing pipeline 200 receives input stroke data 202, such as inked input 206 including different types of input, in the illustrated example, letters, boxes, and lines. Inked strokes (e.g., letters, boxes, and lines) are converted to images (e.g., converted to pixels in image space) at 204.

[0026] Writing detection is then performed on the converted inked strokes at 208. For example, word detection is performed to identify letters 210 (e.g., A, B, YES, and NO) within the image space. For example, the writing detector component 300 is configured to perform word detection on the converted image strokes, such as Figure 3 As illustrated in the figure, the inked input 206 represents the inked input to the semantic segmentation process 302. In one example, the semantic segmentation process 302 performs semantic segmentation using pixel-by-pixel classification (e.g., writing or drawing pixels) to form a set of writing masks (regions) 304. Various different techniques can be used to perform the pixel-by-pixel classification. In one example, the pixel-by-pixel classification is performed using a custom U-net neural network configured for the pixel-by-pixel classification. That is, the neural network is trained according to the neural network technique to perform the pixel-by-pixel classification.

[0027] For example, Figure 4 As illustrated in , a custom U-net neural network 400 configured as a compressed U-net network is utilized. The custom U-net neural network 400 is configured to have a fixed number of cores (as indicated by the numbers in parentheses) on each of the layers 402, 404, 406, and 408. In this configuration, the inherent connection operation of the custom U-net neural network 400 is changed from cascading to addition, thereby providing gains with respect to inference time.

[0028] Reference again Figure 3 , the marking process 306 is configured to use the writing mask 304 (which in some examples is a predicted writing mask) to mark the writing strokes. In the illustrated example, the writing strokes marked are of word type, which are letters. That is, the writing strokes are extracted from the writing mask 304. For example, given the writing mask set {M 1 , R 2 , R 3 , ..., R M} and stroke set {S 1 , S 2 , S 3 , ..., S N}, the marking process 306 determines the written strokes and then marks them. One algorithm for performing this operation is:

[0029] 1. For i=1, 2, ..., N.

[0030] 2. For any stroke point S belonging to the writing mask i The percentage is counted and denoted as p i .

[0031] 3. If p i ≥ threshold.

[0032] 4. S i Set to writing strokes.

[0033] Thus, the written strokes of the inked input 206 are identified and marked.

[0034] return Figure 2 , the identified written strokes are removed at 212. For example, as can be seen in the modified inked input 214, the letter 210 is removed. That is, the pixels in the image space corresponding to the identified letter 210 are removed from the inked input image.

[0035] Drawing detection 216 is then performed by the drawing detector component to identify drawing input corresponding to inked strokes. Figure 5 , a painting detection process 500 is performed to identify inked strokes corresponding to a painting. More specifically, in some examples, the modified inked input 214 represents a non-written image input to a convolutional neural network (CNN) 502, which is illustrated as a Yolo convolutional network that detects and classifies painting objects. For example, the CNN 502 is trained using a neural network training technique to decode the detected object using a bounding box 504. The bounding box 504 facilitates the identification of shape objects illustrated as squares, rectangles, lines, and polylines (lines with arrows). In one example, the bounding box 504 is assigned to define a group of pixels that are classified as an object (e.g., using a class label). The bounding box 504 and the class label are assigned to all objects in the image at the same time. In some examples, the Yolo convolutional network operates using only one forward pass, which is sufficient to obtain a prediction result. Therefore, a faster detection process is performed. It should be noted that other CNNs can be implemented, such as based on the accuracy level desired in target detection.

[0036] In one example, the Yolo convolutional network is an improved Yolo convolutional network configured as a thin Yolo convolutional network. A thin Yolo convolutional network is implemented, which is a combination of a full Yolo convolutional network and a tiny Yolo convolutional network. The full Yolo convolutional network achieves high accuracy, but is slow, and the tiny Yolo convolutional network is fast, but achieves lower accuracy. In one example, the thin Yolo convolutional network is a combination of a full Yolo convolutional network and a tiny Yolo convolutional network.

[0037] Drawing detection at 216 uses different techniques to assign bounding boxes 504 to strokes. In one example, given a stroke with a corresponding predicted label {l 1 , l 2 , l 3 , ..., l M}'s bounding box set {B 1 , B 2 , B 3 , ..., B M} and non-writing stroke set {S 1 , S 2 , S 3 , ..., S N}, the painting detection 216 determines which bounding box 504 the current stroke belongs to. One algorithm for performing this operation is:

[0038] 1. For i=1, 2, ..., N.

[0039] 2. For j=1, 2, ..., M.

[0040] 3. Based on S i and B j To calculate the numerical feature f = (f 1 , f 2 , f 3 , ..., f 7 ).

[0041] 4. Input the numerical features into the binary classifier to calculate S i ∈B j The probability of ij .

[0042] 5. Set the bounding box B j* Assign to S i As S i belongs to an object such that j * = argmax(p i1 , p i2 , ..., p iM ).

[0043] More specifically, if Figure 6 The figure is based on S i 600 and B j The numerical characteristics of 602 (f 1 , f 2 , f 3 , ..., f 7 ) is calculated as follows:

[0044] 1.f 1 As Bj .

[0045] 2.f 2 As B j and for s i The IOU (intersection over union) between the bounding boxes of i ).Right now,

[0046] 3.f 3 As B j Medium S i The length ratio of the part, where

[0047] 4.f 4 As B on the diagonal of the joint rectangle 604 j and B(S i ), where

[0048] 5.f 5 As B j and B(S i ) is the logarithm of the aspect ratio, as follows:

[0049] in

[0050] 6.f 6 As B j and B(S i ) on the x-axis, where

[0051] 7.f 7 As B j and B(S i ) on the y-axis, where

[0052] Where a drawing stroke is detected and a writing stroke has been previously detected, this output 218 is provided to the next stage (eg, for further processing). It should be appreciated that the output from the writing detection performed at 208 is also provided as part of the output 218 to the next stage.

[0053] Therefore, using Figure 7 The inked stroke 700 is parsed by the parsing tree 702 illustrated in FIG. 7 , which may be the selected decision tree. The parsing tree 702 corresponds to the decision tree generated by the parsing pipeline 200 (e.g., Figure 2708. In this example, an inked stroke 700 includes different writing and drawing strokes corresponding to a recipe. A parse tree 702 illustrates the detection of writing areas 706 and drawings 708 from a root 704 (e.g., an image in pixel space of drawn image strokes), which may correspond to writing detection and drawing detection as described herein. For example, using semantic segmentation as described herein, the written strokes are identified and labeled as paragraphs 710 with lines 712, symbols 714, and words 716. The labeled written strokes are then removed from the pixel space and drawing detection is performed to identify and label the drawings 708, as described herein. Thus, the parse tree 702 allows for efficient and accurate detection and labeling of written and drawing strokes by separating ink detection into different components. The parsing pipeline can be used to retain and change neural networks that perform inked transformations according to the present disclosure.

[0054] Thus, digital ink strokes are drawn as an image, which is input to a writing detector, which identifies and labels the writing strokes, which are then removed from the image in pixel space and input to a painting detector. The painting detector then identifies and labels the painting strokes. Thereafter, the labeled writing objects and the labeled painting objects are output by various examples, such as by a parsing pipeline.

[0055] Figure 8 800 is a flowchart illustrating exemplary operations involved in detection and segmentation for inking applications and other applications. In some examples, the operations described with respect to flowchart 800 are performed by Fig.10 1000. Flowchart 800 begins at operation 802, where an ink stroke is received. For example, a user uses an inking application to input text and drawings on a touch screen device. The input device can be any inking device capable of generating inking input for an inking application. It should be appreciated that the operations performed by flowchart 800 can be applied to non-inking applications.

[0056] Operation 804 includes drawing the ink strokes into an image in pixel space. In some examples, the ink strokes are drawn as image pixel data, where the pixel data represents handwriting. That is, in various examples, the ink strokes are drawn into an image in pixel space. Image detection and segmentation techniques can now be used to process the ink strokes.

[0057] Operation 806 includes detecting written strokes and marking the detected written strokes as written objects. For example, image pixels corresponding to the written objects are marked to identify letters corresponding to the ink strokes. Then at operation 808, pixels corresponding to the marked written objects are subsequently removed from the image. For example, all pixels corresponding to the identified letters are removed from the image.

[0058] Operation 810 includes detecting a drawing stroke in an image having pixels corresponding to a marked written object removed, and marking the drawing stroke. For example, image pixels corresponding to the drawing object are marked to identify a line or shape corresponding to the ink stroke. Pixels corresponding to the line and shape are marked accordingly.

[0059] The marked written and drawn objects are output at operation 812, such as for further processing. For example, the segmented and identified, marked written and drawn objects are input to perform additional inking operations.

[0060] Fig. 9 900 is a flowchart illustrating exemplary operations involved in detection and segmentation for inking applications and other applications. In some examples, the operations described with respect to flowchart 900 are performed by Fig.10 1000. Flowchart 900 begins at operation 902, where digital ink input is received. For example, digital ink input is received at a digital ink input device.

[0061] Ink strokes corresponding to the digital ink input are parsed at operation 904. For example, using a parse tree, the writing area and the drawing area are identified and subjected to ink stroke type detection respectively. In various examples, the writing area is processed first, and then the drawing area is processed. To this end, at operation 906, it is determined whether a written stroke is detected. If a written stroke is detected, at operation 908, semantic segmentation using pixel-by-pixel classification is performed to mark the written stroke. In one example, a custom U-net neural network is used. Using this technology, the predicted writing mask is used to mark the written stroke.

[0062] At operation 910, image pixels corresponding to the marked handwritten strokes are removed and output, for example, as handwritten stroke objects. In some examples, the handwritten stroke objects are used for the next stage of processing.

[0063] If no written strokes are detected at operation 906, then at operation 912, it is determined whether a drawing stroke is detected. If no drawing strokes are detected, the operation begins again at operation 902. If a drawing stroke is detected (after removing the pixels at operation 910), then at operation 914, a modified Yolo CNN is used, where the output is decoded using the bounding box and the selected decision tree as the desired classifier. For example, as described herein, a plurality of digital features are calculated based on the non-written strokes and the corresponding bounding boxes. This results in the identification of the drawing strokes, which are marked at 916 and output as marked drawing strokes.

[0064] Thus, various examples include an ink parsing engine configured to perform image-based processing for diagram processing. For handwriting semantic segmentation classification, a compressed variant of the U-net neural network is used in some examples. For painting detection, object detection techniques are used in some examples, with a compressed variant of Yolo configured to reduce computational cost. For the conversion of painting objects to strokes, a decision tree is constructed that specifies the design features to be analyzed in some examples.

[0065] Additional Examples

[0066] Some aspects and examples disclosed herein relate to an ink parsing system, comprising: a memory associated with a computing device, the memory including a writing detector component and a drawing detector component; and a processor executing an ink parsing engine having a parsing pipeline, the parsing pipeline using the writing detector component and the drawing detector component to: receive ink strokes at an inking device input; draw the received inking into an image in pixel space; detect writing strokes in the image and mark the writing strokes using the writing detector component; remove pixels corresponding to the marked writing strokes from the image; detect drawing strokes in the image with the removed pixels using the drawing detector component and mark the drawing strokes; and output writing objects and drawing objects corresponding to the marked writing strokes and the marked drawing strokes, respectively.

[0067] Additional aspects and examples disclosed herein relate to a computerized method for digital ink parsing, comprising: receiving ink strokes at an inking device input; drawing the received ink strokes into an image in pixel space; detecting written strokes in the image and marking the written strokes; removing pixels corresponding to the marked written strokes from the image; detecting drawing strokes in the image with the removed pixels and marking the drawing strokes; and outputting written objects and drawing objects corresponding to the marked written strokes and the marked drawing strokes, respectively.

[0068] Additional aspects and examples disclosed herein relate to one or more computer storage media having computer-executable instructions for digital ink parsing, which instructions, when executed by a processor, cause the processor to at least: receive ink strokes at an inking device input; draw the received ink strokes into an image in pixel space; detect written strokes in the image and mark the written strokes; remove pixels corresponding to the marked written strokes from the image; detect drawing strokes in the image with the removed pixels and mark the drawing strokes; and output written objects and drawing objects corresponding to the marked written strokes and the marked drawing strokes, respectively.

[0069] Alternatively, or in addition to other examples described herein, examples include any combination of the following:

[0070] Perform semantic segmentation using pixel-by-pixel classification to detect handwriting strokes;

[0071] Using the predicted writing mask to mark the writing strokes;

[0072] Use a U-net neural network with multiple layers to perform semantic segmentation, with each layer having a fixed number of cores;

[0073] Use You Look Only Once (Yolo) convolutional network to perform paint stroke detection;

[0074] Decoding the detected paint strokes using a plurality of bounding boxes with corresponding predicted labels and using a decision tree as a binary classifier to detect the paint strokes; and

[0075] Although aspects of the present disclosure have been described in terms of various examples and their related operations, those skilled in the art will appreciate that combinations of operations from any number of the different examples are also within the scope of aspects of the present disclosure.

[0076] Example operating environment

[0077] Fig.10 1 is a block diagram of an example computing device 1000 for implementing various aspects disclosed herein, and is generally designated as computing device 1000. Computing device 1000 is only an example of a suitable computing environment and is not intended to imply any limitation on the functionality or scope of use of the examples disclosed herein. Computing device 1000 should also not be interpreted as having any dependency or requirement on any of the components / modules or component / module combinations illustrated. The examples disclosed herein can be described in the general context of computer code or machine-usable instructions, including computer executable instructions executed by a computer or other machine (such as, a personal data assistant or other handheld device), such as program components. Typically, program components include routines, programs, objects, components, data structures, etc., which refer to codes that perform specific tasks or implement specific abstract data types. The disclosed examples can be practiced in various system configurations, including personal computers, laptop computers, smart phones, mobile tablet computers, handheld devices, consumer electronic devices, professional computing devices, etc. When tasks are performed by remote processing devices linked by a communication network, the disclosed examples can also be practiced in a distributed computing environment.

[0078] The computing device 1000 includes a bus 1010 that directly or indirectly couples the following devices: computer storage memory 1012, one or more processors 1014, one or more presentation components 1016, input / output (I / O) ports 1018, I / O components 1020, power supply 1022, and network components 1024. Although the computer device 1000 is depicted as appearing to be a single device, multiple computing devices 1000 can work together and share the depicted device resources. For example, the computer storage memory 1012 can be distributed across multiple devices, the processor(s) 1014 can be installed on different devices, etc.

[0079] Bus 1010 may represent one or more buses (such as an address bus, a data bus, or a combination thereof). Although shown with lines for clarity, Fig.10 , but in reality, outlining the individual components is not so clear, and metaphorically, the lines are more accurately gray and fuzzy. For example, one might consider a presentation component such as a display device to be an I / O component. Furthermore, a processor has memory. Such is the nature of the art, and to reiterate, Fig.10 The diagrams of FIG. 1 illustrate only exemplary computing devices that may be used in conjunction with one or more of the disclosed examples. No distinction is made between categories such as “workstation,” “server,” “laptop,” “handheld device,” etc., as all are considered within the scope of this disclosure. Fig.10 and within the scope of references to "computing devices" herein. Computer storage memory 1012 may take the form of the following computer storage media references and is operable to provide storage of computer-readable instructions, data structures, program modules, and other data for computing device 1000. For example, computer storage memory 1012 may store an operating system, a general-purpose application platform, or other program modules and program data. Computer storage memory 1012 may be used to store and access instructions configured to perform various operations disclosed herein.

[0080] As mentioned below, computer storage memory 1012 may include computer storage media in the form of volatile and / or non-volatile memory, removable or non-removable memory, data disks in a virtual environment, or a combination thereof. And computer storage memory 1012 may include any number of memories associated with or accessible by computing device 1000. Memory 1012 may be internal to computing device 1000 (e.g., Fig.10), external to computing device 1000 (not shown), or both (not shown). Examples of memory 1012 include, but are not limited to, random access memory (RAM); read-only memory (ROM); electronically erasable programmable read-only memory (EEPROM); flash memory or other memory technology; compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical or holographic media; magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices; memory connected to an analog computing device; or any other medium for encoding the desired information and accessed by computing device 1000. Additionally or alternatively, computer storage memory 1012 may be distributed across multiple computing devices 1000, for example, in a virtualized environment where instruction processing is performed on multiple devices 1000. For purposes of this disclosure, "computer storage media," "computer storage memory," "memory," and "memory device" are synonymous terms for computer storage memory 1012, and none of these terms include carrier waves or propagating signaling.

[0081] (Multiple) processors 1014 may include any number of processing units that read data from various entities such as memory 1012 or I / O components 1020. Specifically, (multiple) processors 1014 are programmed to execute computer executable instructions for implementing various aspects of the present disclosure. Instructions may be executed by a processor, by multiple processors within computing device 1000, or by a processor external to client computing device 1000. In some examples, (multiple) processors 1014 are programmed to execute instructions such as those illustrated in the flowcharts discussed below and depicted in the accompanying drawings. In addition, in some examples, (multiple) processors 1014 represent the implementation of analog technology to perform the operations described herein. For example, the operations may be performed by analog client computing device 1000 and / or digital client computing device 1000. (Multiple) presentation components 1016 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc. Those skilled in the art will understand and appreciate that computer data may be presented in a variety of ways, such as visually in a graphical user interface (GUI), audibly through a speaker, wirelessly between computing devices 1000, through a wired connection, or otherwise. Ports 1018 allow computing device 1000 to be logically coupled to other devices including I / O components 1020, some of which may be built-in. Example I / O components 1020 include, for example, but are not limited to, microphones, joysticks, game controllers, satellite dishes, scanners, printers, wireless devices, and the like.

[0082] The computing device 1000 can operate in a networked environment via a network component 1024 using a logical connection to one or more remote computers. In some examples, the network component 1024 includes a network interface card and / or computer executable instructions (e.g., a driver) for operating a network interface card. Communication between the computing device 1000 and other devices can occur through any wired or wireless connection using any protocol or mechanism. In some examples, the network component 1024 is operable to transmit data wirelessly between devices using short-range communication technology (e.g., near field communication (NFC), Bluetooth™ communication, etc.) or a combination thereof, using a transmission protocol on a public, private, or hybrid (public and private). For example, the network component 1024 communicates with the network 1028 via a communication link 1026.

[0083] Although described in conjunction with the example computing device 1000, the examples of the present disclosure can be implemented with many other general or special computing system environments, configurations, or devices. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with aspects of the present disclosure include, but are not limited to, smartphones, mobile tablet computers, mobile computing devices, personal computers, server computers, handheld devices or laptops, multiprocessor systems, game consoles, microprocessor-based systems, set-top boxes, programmable consumer electronics, mobile phones, wearable or accessory mobile computing and / or communication devices (e.g., watches, glasses, headphones or earphones), network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, VR devices, holographic devices, etc. Such systems or devices may accept input from a user in any manner, including from an input device such as a keyboard or pointing device, via gesture input, proximity input (such as, by hovering), and / or via voice input.

[0084] Examples of the present disclosure may be described in the general context of computer executable instructions such as program modules executed by one or more computers or other devices in software, firmware, hardware or a combination thereof. Computer executable instructions may be organized into one or more computer executable components or modules. Typically, program modules include, but are not limited to, routines, programs, objects, components, and data structures that perform specific tasks or implement specific abstract data types. Aspects of the present disclosure may be implemented with such components or modules of any number and organization. For example, aspects of the present disclosure are not limited to specific computer executable instructions or specific components or modules illustrated in the accompanying drawings and described herein. Other examples of the present disclosure may include different computer executable instructions or components having more or less functionality than illustrated and described herein. In examples involving general-purpose computers, when configured to execute instructions described herein, aspects of the present disclosure convert general-purpose computers into special-purpose computing devices.

[0085] As an example and not limitation, computer-readable media include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable memory implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, etc. Computer storage media are tangible and mutually exclusive with communication media. Computer storage media are implemented in hardware and do not include carriers and propagation signals. Computer storage media for the purposes of this disclosure are not signals themselves. Exemplary computer storage media include hard disks, flash drives, solid-state memories, phase change random access memories (PRAM), static random access memories (SRAM), dynamic random access memories (DRAM), other types of RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, DVD or other optical storage, magnetic cassettes, tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information for access by computing devices. In contrast, communication media typically contain computer-readable instructions, data structures, program modules, or other content in modulated data signals such as carriers or other transmission mechanisms, and include any information delivery media.

[0086] It will be apparent to the skilled artisan that any range or device value presented herein may be expanded or altered without losing the effect sought.

[0087] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0088] It should be understood that the above benefits and advantages may relate to one embodiment or may relate to multiple embodiments. The embodiments are not limited to those embodiments that solve any or all of the problems described or those embodiments that have any or all of the benefits and advantages described. It should also be understood that reference to "an" item refers to one or more of those items.

[0089] The embodiments illustrated and described herein, as well as embodiments not specifically described herein but within the scope of the various aspects of the claims, constitute exemplary means for digital ink parsing. The illustrated one or more processors 1014 together with the computer program code stored in the memory 1012 constitute exemplary processing components for using and / or training a neural network.

[0090] The term “comprising” is used in this specification to mean including the following feature(s) or action(s), but does not exclude the existence of one or more additional features or actions.

[0091] In some examples, the operations illustrated in the figures may be implemented as software instructions encoded on a computer-readable medium, implemented in hardware programmed or designed to perform the operations, or both. For example, aspects of the present disclosure may be implemented as a system on a chip or other circuit comprising a plurality of interconnected conductive elements.

[0092] Unless otherwise noted, the order in which the operations in the examples of the present disclosure illustrated and described herein are performed or implemented is not essential. That is, unless otherwise noted, the operations may be performed in any order, and the examples of the present disclosure may include more or fewer operations than those disclosed herein. For example, it is contemplated that it is within the scope of the various aspects of the present disclosure to perform or implement a particular operation before, simultaneously with, or after another operation.

[0093] When introducing elements of aspects of the present disclosure or examples thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements in addition to the listed elements. The term "exemplary" is intended to mean "an example of..." The phrase "one or more of: A, B, and C" means "at least one of A and / or at least one of B and / or at least one of C."

[0094] Having described various aspects of the disclosure in detail, it is apparent that modifications and variations are possible without departing from the scope of various aspects of the disclosure as defined in the appended claims. As various changes can be made to the above-described constructions, products, and methods without departing from the scope of various aspects of the disclosure, it is intended that all matter contained in the above description and all matter shown in the accompanying drawings shall be interpreted as illustrative and not in a limiting sense.

Claims

1. An ink analysis system, comprising: a memory associated with a computing device, the memory including a writing detector component and a drawing detector component; as well as at least one processor executing an ink parsing engine that uses the writing detector component and the drawing detector component to: receiving ink strokes at an inking device input; drawing the received ink strokes into an image; detecting one or more handwriting strokes in the image using the handwriting detector component and marking the handwriting strokes; removing from the image one or more pixels corresponding to the marked handwritten strokes; After detecting the one or more written strokes and removing the one or more pixels corresponding to the marked written strokes, detecting one or more painted strokes in the image using the painting detector component and marking the painted strokes; as well as One or more writing objects and one or more drawing objects corresponding to the marked writing strokes and the marked drawing strokes are output respectively.

2. The ink parsing system of claim 1, wherein the at least one processor executes the ink parsing engine to perform semantic segmentation using the handwriting detector component, the semantic segmentation using pixel-by-pixel classification to detect the handwriting strokes.

3. The ink parsing system of claim 2, wherein the at least one processor executes the ink parsing engine to mark the handwriting strokes using a predicted handwriting mask.

4. The ink parsing system of claim 2, wherein the at least one processor executes the ink parsing engine to perform the semantic segmentation using a neural network having a plurality of layers, each layer having a fixed number of kernels.

5. The ink parsing system of claim 1, wherein the at least one processor executes the ink parsing engine to perform paint stroke detection utilizing the paint detector component using a convolutional network.

6. The ink parsing system of claim 5, wherein the at least one processor executes the ink parsing engine to decode the detected paint strokes using a plurality of bounding boxes with corresponding predicted labels, and also detect the paint strokes using a decision tree as a binary classifier.

7. The ink analysis system according to claim 1, wherein: The ink parsing engine has a parsing pipeline configured to perform detection using the writing detector component before performing detection using the drawing detector component, the writing detector assembly being configured to remove the one or more pixels to generate the modified inking input, The drawing detector component is configured to perform detection of the one or more drawing strokes on the modified inking input comprising a non-written image, and The modified inking input, the one or more writing objects, and the one or more drawing objects are output to a next stage for processing.

8. A computerized method for digital ink analysis, the computerized method comprising: receiving ink strokes at an inking device input; drawing the received ink strokes into an image; detecting one or more handwritten strokes in the image and marking the handwritten strokes; removing from the image one or more pixels corresponding to the marked handwritten strokes; After detecting the one or more written strokes and removing the one or more pixels corresponding to the marked written strokes, detecting one or more painted strokes in the image and marking the painted strokes; as well as One or more writing objects and one or more drawing objects corresponding to the marked writing strokes and the marked drawing strokes are output respectively.

9. The computerized method of claim 8, further comprising: Semantic segmentation is performed using pixel-wise classification to detect the written strokes.

10. The computerized method of claim 9, further comprising: The predicted writing mask is used to mark the writing strokes.

11. The computerized method of claim 9, further comprising: A neural network with multiple layers is used to perform the semantic segmentation, each layer having a fixed number of kernels.

12. The computerized method of claim 8, further comprising: Use convolutional networks to perform paint stroke detection.

13. The computerized method of claim 12, further comprising: The detected paint strokes are decoded using multiple bounding boxes with corresponding predicted labels, and a decision tree is used as a binary classifier to detect the paint strokes.

14. The computerized method of claim 8, further comprising: A parsing pipeline is used that is configured to detect written strokes before detecting drawn strokes.

15. One or more computer storage media having computer executable instructions for digital ink parsing, the computer executable instructions, when executed by at least one processor, causing the at least one processor to at least: receiving ink strokes at an inking device input; drawing the received ink strokes into an image; detecting one or more handwritten strokes in the image and marking the handwritten strokes; removing from the image one or more pixels corresponding to the marked handwritten strokes; After detecting the one or more written strokes and removing the one or more pixels corresponding to the marked written strokes, detecting one or more painted strokes in the image and marking the painted strokes; as well as One or more writing objects and one or more drawing objects corresponding to the marked writing strokes and the marked drawing strokes are output respectively.

16. The one or more computer storage media of claim 15, further having computer executable instructions that, when executed by the at least one processor, cause the at least one processor to at least: perform semantic segmentation using pixel-by-pixel classification to detect the handwritten strokes.

17. The one or more computer storage media of claim 16, further having computer executable instructions that, when executed by the at least one processor, cause the at least one processor to at least: use a predicted writing mask to mark the writing strokes.

18. The one or more computer storage media of claim 16, further comprising computer executable instructions which, when executed by the at least one processor, cause the at least one processor to at least: use a neural network having multiple layers to perform the semantic segmentation, each layer having a fixed number of cores.

19. The one or more computer storage media of claim 15, further comprising computer executable instructions which, when executed by the at least one processor, cause the at least one processor to at least: perform paint stroke detection using a convolutional network, decode the detected paint strokes using a plurality of bounding boxes with corresponding predicted labels, and detect the paint strokes using a decision tree as a binary classifier.

20. The one or more computer storage media of claim 15, further comprising computer executable instructions which, when executed by the at least one processor, cause the at least one processor to at least: use a parsing pipeline configured to detect written strokes before detecting drawn strokes.

Citation Information

Patent Citations

  • One-dimensional handwritten character input equipment and one-dimensional handwritten character input equipment

    CN105549890A

  • Ink Stroke Grouping Method And Product Based On Stroke Attributes

    CN105930763A