Display device, display system, display method, program

The display device addresses the mismatch issue by integrating handwriting and voice recognition with speaker identification to automatically align and size text data, improving user experience.

JP7838380B2Active Publication Date: 2026-04-01RICOH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-08
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Conventional display devices fail to determine whether text data recognized by handwriting and voice match, leading to inconveniences such as unnecessary operations and incorrect display positions and sizes.

Method used

A display device with a handwritten data receiving unit, character recognition unit, voice data input receiving unit, voice recognition unit, and display control unit that matches and displays text data based on speaker recognition, ensuring accurate alignment and size adjustment of handwritten and voice-recognized text.

Benefits of technology

The device accurately determines and aligns handwritten and voice-recognized text data, eliminating the need for manual specification of display positions and sizes, thus enhancing user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838380000001
    Figure 0007838380000001
  • Figure 0007838380000002
    Figure 0007838380000002
  • Figure 0007838380000003
    Figure 0007838380000003
Patent Text Reader

Abstract

To provide a display device that determines whether or not handwriting recognized text data matches voice recognized text data.SOLUTION: A display device 2 comprises: a handwriting data reception unit that receives input of handwriting data by input means; a character recognition unit that converts the handwriting data to first text data; a voice data input reception unit that receives input of voice data; a voice recognition unit that converts the voice data to second text data; and a display control unit that displays third text data converted from the voice data by the voice recognition unit when the first text data converted by the character recognition unit matches at least a part of the second text data converted by the voice recognition unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] ,

[0001] The present invention relates to a display device, a display system, a display method, and a program.

Background Art

[0002] There is known a display device that uses handwriting recognition technology to convert handwriting data into characters and display them on a display. A display device equipped with a relatively large touch panel is installed in a conference room or a public facility and used as an electronic blackboard by multiple users.

[0003] [[ID=()]] There is also known a technology for accepting input of text data obtained by voice recognition of a user's speech (see, for example Patent Document 1). Patent Document 1 discloses a technology for improving character recognition accuracy by correcting the recognition result of handwriting data using the recognition result of speech recognition.

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the conventional technology has a problem that it does not determine whether the text data recognized by handwriting and the text data recognized by voice match.

[0005] In view of the above problems, an object of the present invention is to provide a display device that determines whether the text data recognized by handwriting and the text data recognized by voice match.

Means for Solving the Problems

[0006] In view of the above problems, the present invention provides a handwritten data receiving unit that receives handwritten data input by an input means, a character recognition unit that converts the handwritten data into first text data, a voice data input receiving unit that receives voice data input, a voice recognition unit that converts the voice data into second text data, and a display control unit that displays a third text data converted from the voice data by the voice recognition unit when at least a portion of the first text data converted by the character recognition unit and the second text data converted by the voice recognition unit match. The system includes a speaker recognition unit that, within a certain time after the display control unit displays the first text data, compares feature information extracted from the audio data received by the audio data input receiving unit with feature information of the audio data registered in advance for each user to recognize the speaker of the audio data received by the audio data input receiving unit, and if the speaker recognized by the speaker recognition unit is the scribe of the first text data, the audio recognition unit converts the scribe's audio data into second text data. To provide a display device. Furthermore, the present invention provides a display device comprising: a handwriting data receiving unit that receives handwriting data input via an input means; a character recognition unit that converts the handwriting data into first text data; a voice data input receiving unit that receives voice data input; a voice recognition unit that converts the voice data into second text data; and a display control unit that displays third text data converted from voice data by the voice recognition unit when at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the voice recognition unit. The character recognition unit determines the size of the first text data based on the size of the handwriting data received by the handwriting data receiving unit, and the display control unit displays the third text data at a size based on the size of the first text data. [Effects of the Invention]

[0007] A display device can be provided that determines whether handwritten text data and speech-recognized text data match. [Brief explanation of the drawing]

[0008] [Figure 1] This diagram illustrates the general process by which a display device uses speech recognition to obtain text data from a speaker's utterance and then displays that data. [Figure 2] This figure shows an example of the hardware configuration of a display device. [Figure 3] This figure shows an example of the hardware configuration of a contact sensor. [Figure 4] This is an example of a functional block diagram that explains the functions of a display device by dividing them into blocks. [Figure 5] This figure shows an example of input data stored in the input data storage unit. [Figure 6] This figure shows an example of the initial screen displayed by the display device. [Figure 7] This is a diagram explaining how to calculate font size. [Figure 8] This diagram illustrates how to calculate the font size for cursive or block letters. [Figure 9]A diagram showing the number of pixels of "today" existing on the line slicing the rectangle for each line. [Figure 10] A diagram showing an example of a voice input mark. [Figure 11] A diagram showing an example of a voice input mark moved to the right side of "Today's Topic". [Figure 12A] An example of a flowchart diagram explaining the process of the display device receiving inputs of handwritten recognition text data and voice recognition text data (Part 1). [Figure 12B] An example of a flowchart diagram explaining the process of the display device receiving inputs of handwritten recognition text data and voice recognition text data (Part 2). [Figure 13] A diagram showing an example of a screen where the voice input mark is erased. [Figure 14] A diagram showing an example of the movement of the voice input mark. [Figure 15] An example of a diagram showing the voice input mark displayed on the right side of "(1)". [Figure 16] An example of a diagram showing the text data "Planning" and the voice input mark displayed on the right side of "(1)". [Figure 17] An example of a schematic configuration diagram of the display system. [Figure 18] A diagram showing an example of the hardware configuration of the server device. [Figure 19] An example of a functional block diagram explaining the functions of the display device and the server device by dividing them into blocks.

Embodiments for Carrying Out the Invention

[0009] Hereinafter, as an example of an embodiment for carrying out the present invention, a display device and a display method performed by the display device will be described with reference to the drawings.

Examples

[0010] <Supplementary Explanation Regarding Handwritten Input and Voice Input> Since it is a burden for the note-taker to write by hand on a display device such as an electronic blackboard (the note-taker's hand gets tired or it takes a long time), there is a desire to input characters by voice instead of handwritten input.

[0011] However, in this case, there are the following problems.

[0012] 1. When a note-taker such as a chairperson who is writing by hand on a display device tries to input characters by voice recognition, if another person speaks, the voice of the other person will be converted into text data by voice recognition and displayed on the display device.

[0013] To avoid this inconvenience, there is a method in which the note-taker displays the list of meeting participants on the display device and selects the target person for voice recognition. However, it is necessary to select the target person every time characters are input by voice, and this operation is troublesome. <…>

[0014] 2. The note-taker must specify the position where the text data converted from voice by voice recognition is to be displayed. If the note-taker does not specify the display position of the text data, it will be displayed at the default display position (for example, the upper left of the screen).

[0015] 3. When the display device displays the text data converted from voice by voice recognition, it will be displayed in the default size unless the size of the characters is specified in advance. When it is to be displayed in a size other than the default, the note-taker needs to specify the size of the characters in advance (before the speaker speaks) from a menu or the like, and the operation is troublesome.

[0016] <Outline of the process> Therefore, in this embodiment, the following method is used to realize that the note-taker inputs characters by voice instead of handwritten input.

[0017] FIG. 1 is a diagram for explaining an outline of a process of displaying text data obtained by the display device 2 performing voice recognition on the speech of a speaker.

[0018] (i) Speakers A, B, and C are each speaking. The speech recognition engine 101 of the display device 2 uses speech recognition technology to convert each speaker's speech into text data. Text data A, B, and C are the text data of speakers A, B, and C, respectively. In this embodiment, it is assumed that speakers A, B, and C do not speak at the same time, but they may speak at the same time. In this case, the display device 2 separates the speech data for each speaker.

[0019] (ii) The scribe X may be any of speakers A, B, or C, or may be someone other than speakers A, B, or C. The handwriting recognition engine 102 uses handwriting recognition technology on the handwriting data of scribe X and converts it into text data X.

[0020] (iii) The display device 2 compares the speaker feature vectors (an example of matching information) detected from the speech data of scribe X and speakers A, B, and C with the speaker feature vectors that have been registered in advance, and determines whether there are any speaker feature vectors whose similarity is equal to or greater than a threshold.

[0021] (iv) If only scribe X is speaking, the speaker feature vector registered by scribe X is identified, and scribe X's user ID is also identified. Here, assume that scribe X is speaking and speaker B's user ID has been identified (scribe X = speaker B).

[0022] (v) The display device 2 determines whether at least a portion of the text data X of the scribe X and the speech data B of the speaker B match.

[0023] (vi) If at least a portion of text data X and text data B match, the display device 2 will subsequently use the text data converted from speaker B's (i.e., scribe X's) speech data as input assistance.

[0024] In this way, the display device 2 identifies the scribe using the audio data, and uses the scribe's audio data as input assistance when at least a portion of the text data B obtained by speech recognition of the scribe's audio data matches the text data X of the scribe X. Therefore, even if someone other than the scribe speaks, the display device 2 can be prevented from displaying the text data converted from that audio data.

[0025] Furthermore, since the text data converted from speech is displayed immediately after the handwritten text data X by the scribe, the scribe does not need to specify the display position. Also, the display device 2 displays the text data converted from speech at the same size as the text data X converted from the scribe's handwritten data, so the scribe does not need to specify the font size in advance (before the speaker speaks).

[0026] <About Terminology> The input method can be any means that allows handwriting by specifying coordinates on a touch panel. Examples include an electronic pen, a person's finger or hand, or a rod-shaped object.

[0027] A stroke is a series of operations performed by a user, involving pressing an input device against the display, moving it continuously, and then lifting it from the display. Stroke data is information displayed on the screen based on the trajectory of coordinates entered by the input device. Stroke data may be interpolated as appropriate. Handwritten data is data that contains one or more stroke data. Handwritten input indicates that handwritten data is entered by the user.

[0028] The elements displayed on the screen based on stroke data are called objects. While "object" generally means a target or subject, in this embodiment it refers to the object being displayed. Examples include handwritten data, text data, tables, figures, images, and so on.

[0029] The string of characters converted from handwritten data using character recognition may include not only text data, but also data displayed based on user actions, such as stamps that appear as fixed characters or marks like "completed," shapes like circles and stars, and straight lines.

[0030] Text data refers to one or more characters handled by a computer. Essentially, text data is a character code. Text data includes numbers, letters, and symbols.

[0031] Conversion refers to the process of converting handwritten or audio data into character codes and displaying those codes using a specified font. Furthermore, conversion can also include transforming handwritten data into shapes such as lines, curves, and rectangles, or into tables.

[0032] <Example Hardware Configuration> Figure 2 shows an example of the hardware configuration of the display device 2. The display device 2 in this embodiment is equipped with a CPU (Central Processing Unit) 201, ROM (Read Only Memory) 202, RAM (Random Access Memory) 203, SSD (Solid State Drive) 204, network controller 205, and external device connection I / F (Interface) 206, and is a shared terminal for sharing information among multiple users.

[0033] Of these, the CPU 201 controls the operation of the entire display device 2. The ROM 202 stores programs used to drive the CPU 201, such as the CPU 201 and the IPL (Initial Program Loader). The RAM 203 is used as the work area for the CPU 201.

[0034] SSD204 stores various data, such as the OS and programs for display device 2. This program may be an application program that runs on an information processing device equipped with a general-purpose OS (such as Windows®, Mac OS®, Android®, iOS®, etc.). In other words, display device 2 can be a PC or a smartphone.

[0035] The network controller 205 controls communication with the network. The external device connection interface 206 controls communication with the USB (Universal Serial Bus) memory 2600 and external devices (camera 2400, speaker 2300, microphone 2200).

[0036] The display device 2 also includes a capture device 211, a GPU 212, a display controller 213, a contact sensor 214, a sensor controller 215, an electronic pen controller 216, a short-range communication unit 219, and an antenna 219a for the short-range communication unit 219.

[0037] Of these, the capture device 211 transfers still image data or video data input from the PC 10 to the GPU (Graphics Processing Unit) 212. The GPU (Graphics Processing Unit) 212 is a semiconductor chip specializing in graphics. The display controller 213 controls and manages screen display in order to output the output image from the GPU 212 to the display 220, etc.

[0038] Figure 3 shows the hardware configuration of the contact sensor 214. In this figure, the infrared-emitting LEDs and phototransistors in a row are spaced equally apart, and each infrared-emitting LED and phototransistor is positioned opposite the other. Note that the figure shows an example where 20 infrared-emitting LEDs and 20 phototransistors are arranged horizontally and 15 vertically, but for sizes larger than 40 inches, many more LEDs and phototransistors will actually be required.

[0039] The contact sensor 214 outputs the number of the phototransistor that is blocked by an object, i.e., a phototransistor that cannot detect light, to the sensor controller 215, which then identifies the coordinate position of the object's contact. The electronic pen controller 216 communicates with the electronic pen 2500 to determine whether the pen tip or the pen end is touching the display 220. The short-range communication unit 219 is a communication circuit such as NFC or Bluetooth®.

[0040] Furthermore, the electronic whiteboard 200 is equipped with a bus line 210. The bus line 210 is an address bus, data bus, etc., for electrically connecting each component such as the CPU 201.

[0041] Furthermore, the contact sensor 214 is not limited to the infrared blocking method. Various detection means may be used, such as a capacitive touch panel that identifies the contact position by detecting changes in capacitance, a resistive touch panel that identifies the contact position by voltage changes between two opposing resistive films, or an electromagnetic induction touch panel that identifies the contact position by detecting electromagnetic induction caused by contact between an object and the display. In addition, the electronic pen controller 216 may be configured to determine whether or not there is a touch not only on the tip and end of the electronic pen 2500, but also on the part of the electronic pen 2500 that the user holds and other parts of the electronic pen.

[0042] <About the features> Next, the functions of the display device 2 will be explained using Figure 4. Figure 4 is an example of a functional block diagram that explains the functions of the display device 2 in a block-like manner. The display device 2 has a handwriting data receiving unit 21, a drawing data generation unit 22, a character recognition unit 23, a display control unit 24, a data recording unit 25, a network communication unit 26, an operation receiving unit 27, a speech recognition unit 28, a speaker recognition unit 29, a recognition result verification unit 30, a speech data input receiving unit 31, and a storage unit 40. Each function of the display device 2 is a function or means realized by any of the components shown in Figure 2 operating according to instructions from the CPU 201 that follow a program deployed from the SSD 204 onto the RAM 203.

[0043] The handwriting data receiving unit 21 detects the coordinates of the position where the electronic pen 2500 makes contact with the contact sensor 214. The drawing data generation unit 22 obtains the coordinates of the contact point of the pen tip of the electronic pen 2500 from the handwriting data receiving unit 21. The drawing data generation unit 22 generates stroke data by interpolating this sequence of coordinate points.

[0044] The character recognition unit 23 performs character recognition processing on one or more stroke data (handwritten data) written by the writer and converts them into character codes. The character recognition unit 23 recognizes characters (not only Japanese but also multiple languages ​​such as English), numbers, symbols (%), $, &, etc., shapes (lines, circles, triangles, etc.) in parallel with the writer's pen operation. Various algorithms have been devised for recognition methods, but in this embodiment, known technology can be used, so the details are omitted.

[0045] The display control unit 24 displays handwritten data, text data converted from handwritten data, and operation menus for the scribe on the display 220. The data recording unit 25 stores handwritten data written on the display device 2, text data converted from handwritten data, screen data input from a PC, and files in the storage unit 40. The network communication unit 26 connects to a network such as a LAN and transmits and receives data with other devices via the network.

[0046] The voice data input receiving unit 31 encodes the voice data input from the microphone 2200 using PCM (pulse code modulation). The PCM-encoded voice data input from the microphone 2200 is temporarily stored in the RAM 203 and used by both the voice recognition unit 28 and the speaker recognition unit 29.

[0047] The speaker recognition unit 29 extracts acoustic features (an example of feature information) from the PCM (pulse code modulation) encoded audio data input from the microphone 2200 at short intervals of several tens of milliseconds, and converts these values ​​into an acoustic feature vector. The speaker recognition unit 29 calculates a speaker feature vector from the acoustic feature vector using a UBM (Universal Background Model) model or a speaker feature extraction model. The speaker recognition unit 29 then compares these multiple speaker feature vectors with speaker feature vectors for each speaker (registered speaker feature vectors) that have been pre-stored in the memory unit 40 to determine the similarity. If this similarity is, for example, 60% or higher, the speaker of the audio data is determined to be a registered speaker.

[0048] Before the meeting, each meeting participant uses a PC or smartphone to speak for at least 10 seconds and sends the audio data (PCM encoded data) file and their user ID (which can be a character code for kanji / hiragana, e.g., Shift JIS) to the display device 2 via the network. For transmission, the PC or smartphone uses a protocol such as HTTP. When the display device 2 receives the audio data from each meeting participant, the speaker recognition unit 29 extracts acoustic features from the audio data at short intervals of tens of milliseconds and calculates an acoustic feature vector by representing these values ​​as vectors. The speaker recognition unit 29 then calculates a speaker feature vector from the acoustic feature vector using a UBM model or a speaker feature extraction model. The speaker recognition unit 29 associates these multiple speaker feature vectors with the received user ID and stores them in the storage unit 40 as user information 42.

[0049] The speech recognition unit 28 receives PCM-encoded speech data from the microphone 2200, extracts speech features, identifies phoneme models, and identifies words using a pronunciation dictionary. It then outputs text data of the identified words. The pronunciation dictionary data is stored in the SSD 204. Any method can be used for speech recognition. For example, a speech recognition method using an RNN (Recurrent Neural Network) is known.

[0050] The recognition result matching unit 30 compares the text data generated by the character recognition unit 23 through handwriting recognition with the text data generated by the speech recognition unit 28 through speech recognition to determine whether at least a portion of them match.

[0051] Furthermore, the display device 2 has a storage unit 40 built on an SSD 204 or RAM 203 as shown in Figure 2, and an input data storage unit 41 is built on the storage unit 40. In addition, user information 42 is stored in the storage unit 40 at the start of the meeting or before, in which speaker feature vectors are associated with user IDs for each user (for example, meeting participants).

[0052] Figure 5 shows an example of input data stored in the input data storage unit 41. The input data storage unit 41 stores the input objects. Each item in the input data storage unit 41 will be explained below.

[0053] The object ID is the identification information of the object displayed by the display device 2.

[0054] • The type indicates the type of object, such as text data, stroke data, image, file, or table. In the case of character recognition-generated text data, each conversion unit converted by character recognition is considered one object. In the case of speech recognition-generated text data, each conversion unit converted by the speech recognition unit (divided by periods of silence exceeding a certain duration) is considered one object. In the case of stroke data, the period from when the writer starts writing until the input stops is considered one object.

[0055] The coordinates represent the display position of the object on display 220. These coordinates can be, for example, the position of the top-left vertex of the object's bounding rectangle.

[0056] • The size refers to the dimensions of the object. For text data, it is the size of a single character; for stroke data, it is the height of the bounding rectangle of the entire object.

[0057] • The input source indicates how the object was entered. For text data, this could include handwriting input, voice input, file input, etc.

[0058] The specific speaker speech recognition mode is an input mode in which speech-recognized text data is entered consecutively after handwritten-recognized text data. In Figure 5, the object with object ID=2 is Y, indicating that it was entered using the specific speaker speech recognition mode.

[0059] • The inputter is the identification information of the person who input the object. In the case of handwritten text data, it is determined based on the matching result of the speaker feature vector of the speech data recognized immediately after the handwritten data was recognized. In the case of speech text data, it is determined based on the matching result of the speaker feature vector generated from the speech data.

[0060] <Example of a text data input screen> Next, with reference to Figure 6, we will explain the screen operated by the scribe using voice input. Figure 6 shows an example of the initial screen 400 displayed by the display device 2. The initial screen 400 is the screen that is displayed immediately after the power of the display device 2 is turned on or immediately after logging in. The initial screen 400 displays a handwriting input icon 401 for switching to handwriting input mode, a drawing icon 402 for switching to drawing mode, and a voice input transition icon 403 for switching to voice input mode, which inputs text data by comparing the text data of the recognition results of character recognition and voice recognition.

[0061] The scribe (for example, the chairperson) selects the handwriting input icon 401 and the voice input transition icon 403, and writes, for example, "today" at any position on the screen, and then speaks "kyo" after writing. The handwriting and speaking can be done almost simultaneously (in parallel). The character recognition unit 23 then compares the handwritten data (word) "today" with the dictionary data, and if it confirms that there is a matching string (word), it outputs the string "today" as text data.

[0062] As shown in Figure 7, the character recognition unit 23 determines the size of the handwritten data and outputs the font size along with the text data. Figure 7 is a diagram illustrating the method for calculating the font size. The character recognition unit 23 identifies the boundaries between characters based on the distance between each stroke, etc., and decomposes them into individual characters. From the average of the widths W1, W2 and the heights H1, H2 of two characters, the character recognition unit 23 can determine the font size. The character recognition unit 23 may also adopt the larger or smaller of the two characters as the font size.

[0063] Figure 8 illustrates the method for calculating the font size of cursive or block letters. In the English cursive shown in Figure 8(a), it can be difficult for the character recognition unit 23 to identify the boundaries between characters based on the distance between each stroke. Therefore, if a predetermined time has elapsed since the user started (pen down) and finished (pen up) writing a character, and there is no start of the next handwriting (pen down), the character recognition unit 23 calculates the font size for the handwriting data from pen down to pen up. If the next handwriting (pen down) occurs within the predetermined time after the end of handwriting (pen up), it is considered as if handwriting is in progress. Pen up refers to the change in the state when the contact sensor 214 stops detecting light obstruction. Pen down refers to the change in the state when the contact sensor 214 stops detecting light obstruction. The condition of a predetermined time or longer is to avoid considering the horizontal bar of "t" and the dots (superscript dots) of "i" and "j" as individual characters or strings of characters.

[0064] As shown in Figure 8(a), when a user writes, for example, "today" in cursive, a pen up and a pen down occur between the start of writing "t" (pen down) and the end of writing "y" (pen up). However, the time between these pen up and pen down is shorter than a predetermined time, so the character recognition unit 23 determines the font size from the entire string "today".

[0065] Next, as shown in Figure 8(b), the character recognition unit 23 determines a rectangle 50 that encloses the string "today".

[0066] Next, the character recognition unit 23 sets lines 51 that slice the rectangle 50 horizontally at regular intervals and determines the number of pixels of "today" that exist on the lines 51. The character recognition unit 23 estimates the font size based on the range with a large number of pixels on the lines 51. Here, the pixels may be the display pixels of the display 220. The regular interval may be measured in pixels or in length. In Figure 8(b), the intervals between the lines 51 are large for the sake of drawing, but as an example, we will explain the case where a line 51 is set for every pixel.

[0067] Figure 9 is a table showing the number of pixels of "today" on each line 51 that slices the rectangle horizontally. In Figure 9, it is assumed that the vertical number of pixels in the rectangle is 30. Therefore, the table in Figure 9 has 30 lines, from the 1st line (top edge of rectangle 50) to the 30th line (bottom edge of rectangle 50).

[0068] In Figure 9, the lines are broadly divided into those with a large number of pixels (lines 8 to 21) and those with a small number of pixels (lines 1 to 7 and lines 22 to 30). The character recognition unit 23 then determines the font size based on the number of pixels on line 51. In the example in Figure 9, the number of lines from line 8 to line 21 is considered to represent the height of the character, and the font size is determined accordingly. The character recognition unit 23 considers the horizontal size of the character to be the same as the vertical size.

[0069] Furthermore, clustering can be used to determine the boundary lines (the 8th and 22nd lines). Clustering is a form of machine learning that groups data based on the similarity between them. In Figure 9, the data is grouped into two groups based on the number of pixels: 4 or less and 4 or more. Clustering methods include k-means, group average, Ward's method, shortest distance method, and longest distance method. The number of pixels per line does not follow a normal distribution, but this allows for the determination of appropriate boundaries.

[0070] Let's return to Figure 8(c) for explanation. Figure 8(c) shows a size frame 52 in which the number of lines from the 8th to the 21st line is considered as the height of the character. In this way, the character recognition unit 23 can determine the font size of the handwritten cursive data. Note that the size frame 52 results in a slightly smaller font size, so the character recognition unit 23 may consider the size frame 52 to be slightly larger, or determine a font size that is slightly larger than the font size determined by the size frame 52.

[0071] Figure 8 uses cursive writing as an example, but the font size can be determined similarly for printed type (also called block letters). Printed type refers to a typeface where each character is independent.

[0072] Furthermore, whether the character recognition unit 23 determines the font size as shown in Figure 7 or as shown in Figure 8 depends on whether the user has set a language for handwriting (the language to be converted to text). If no language is set, the character recognition unit 23 should automatically determine the language. As for how to automatically determine the language, it could be the language with the highest accuracy when text is converted to several languages, or a correspondence model could be generated by machine learning to match stroke data with languages, and the determination could be made using this correspondence model.

[0073] Let's return to Figure 6 for explanation. The CPU 201 converts the string "today" into font data of the specified font size and issues an instruction to display it at the handwritten position of "today," and the display control unit 24 displays this on the display (an example of the first text data). The display control unit 24 erases the handwritten "today" (handwritten data) and then displays the font data of "today."

[0074] Meanwhile, the speech recognition unit 28 extracts speech features, identifies a phoneme model, and identifies the word using a pronunciation dictionary from the speech "today" (PCM-encoded speech data) input from the microphone 2200, and outputs text data of the identified word "today" (an example of the second text data).

[0075] Furthermore, the speaker recognition unit 29 extracts acoustic features from the "today" audio data at short intervals of several tens of milliseconds and converts these values ​​into an acoustic feature vector. The speaker recognition unit 29 calculates a speaker feature vector from the acoustic feature vector using a UBM model or a speaker feature extraction model. The speaker recognition unit 29 then compares this speaker feature vector with the speaker feature vectors for each speaker registered in the user information 42 to determine the similarity. If the similarity between the two speaker feature vectors is above a threshold (for example, 60% or more) and there is a speaker feature vector registered in the user information 42, the speaker recognition unit 29 determines that the user ID associated with this speaker feature vector is the scribe (for example, the chairperson). The data recording unit 25 associates this speaker's user ID with the text data "today" and saves it in RAM 203.

[0076] Next, the recognition result matching unit 30 compares the text data converted by the character recognition unit 23 with the text data converted by the speech recognition unit 28, and if at least a part of them match, it decides to switch to specific speaker speech recognition mode. When in specific speaker speech recognition mode, the display control unit 24 displays a speech input mark 404 at the right end of the string "today".

[0077] Figure 10 shows an example of a voice input mark 404. The voice input mark 404 is displayed to the right of "Today" which is displayed as text data (i.e., at the end of the text data in the text input direction). The voice input mark 404 indicates that the matching of the scribe's voice is complete and that the text data in which the scribe's voice was recognized will be displayed from this position.

[0078] Next, when the scribe (for example, the chairperson) speaks "nogidai," the speech recognition unit 28 extracts speech features, identifies phoneme models, and identifies words using a pronunciation dictionary from the "nogidai" speech (PCM-encoded speech data) input from the microphone 2200, and outputs the text data "nogidai."

[0079] Furthermore, the speaker recognition unit 29 extracts acoustic features from the "Nogidai" audio data at short intervals of tens of milliseconds and generates an acoustic feature vector by representing these values ​​as vectors. The speaker recognition unit 29 calculates a speaker feature vector from the acoustic feature vector using a UBM or speaker feature extraction model. The speaker recognition unit 29 then compares this speaker feature vector with the speaker feature vectors for each speaker that have been previously stored in the user information 42 to determine the similarity. When the speaker recognition unit 29 confirms that the similarity with the speaker feature vector of the scribe of "Today" is above a threshold (for example, 60% or more), it outputs information indicating that the speaker of "Nogidai" is the scribe (the same speaker as "Today").

[0080] When CPU201 confirms that the speaker is the chairperson (the same speaker as "today"), it displays "the agenda" to the right of the string "today" in the same font size as "today" via the display control unit 24, and moves the voice input mark 404 to the right of "the agenda" (an example of third text data).

[0081] Figure 11 shows an example of the voice input mark 404 being moved to the right of "Today's Agenda." In this way, the voice input mark 404 is moved to the right edge of the text data entered by speech recognition.

[0082] If someone other than the scribe speaks, the speaker recognition unit 29 calculates a speaker feature vector from the audio data input from the microphone 2200 in the same manner as described above. The speaker recognition unit 29 then compares this speaker feature vector with the speaker feature vectors registered for each speaker in advance to determine the similarity. If the speaker recognition unit 29 confirms that the similarity with the speaker feature vector of someone other than the scribe is above a threshold (for example, 60% or more), it outputs information indicating that the audio was not spoken by the scribe (the same speaker as "today").

[0083] The speech recognition unit 28 extracts speech features, identifies phoneme models, and identifies words using a pronunciation dictionary from the speech of a person other than the scribe (PCM-encoded speech data) input from the microphone 2200, and outputs the converted text data. If the speaker of "the agenda" is not the same as the speaker of "today" (because they are not the chairperson), the CPU 201 does not display this text data. Alternatively, it may be displayed in a fixed location such as the right edge of the display 220, but not immediately following the handwritten recognized text data ("today").

[0084] <Processing procedure for display device> Figures 12A and 12B are flowcharts illustrating the process by which the display device 2 accepts handwritten text data and speech-recognized text data. The process in Figures 12A and 12B starts, for example, from the initial screen 400.

[0085] The operation reception unit 27 detects that the handwriting input icon 401 and the voice input transition icon 403 have been selected (S1).

[0086] Next, the character recognition unit 23 determines whether or not the characters were handwritten (S2).

[0087] When characters are entered, the character recognition unit 23 recognizes the characters and converts them into text data, and the display control unit 24 displays the text data (S3). The character recognition unit 23 automatically performs character recognition after a certain period of time has elapsed since the writer lifted the input means from the touch panel (after the pen was raised). The character recognition unit 23 may also perform character recognition based on the writer's operation. Furthermore, as shown in Figure 7, the character recognition unit 23 determines the size of the characters and then determines the size of the text data.

[0088] The speech recognition unit 28 starts a timer to measure the time elapsed since the scribe's text data was displayed (S4).

[0089] The voice recognition unit 28 then monitors the voice data detected by the microphone 2200 and determines whether or not voice input has been received (S5).

[0090] If the timer times out without any voice input (Yes in S6), the process returns to step S2. If the timer does not time out without any voice input (No in S6), the process returns to step S5.

[0091] If voice input is received before the timer times out (Yes in S5), the voice recognition unit 28 converts the voice into text data through voice recognition processing (S7). For explanatory purposes, this text data is referred to as the "first text data".

[0092] Furthermore, the speaker recognition unit 29 calculates the speaker's speaker feature vector through speaker recognition processing, and compares the calculated speaker feature vector with the speaker feature vectors for each speaker stored in the user information 42 to determine the similarity (S8).

[0093] The speaker recognition unit 29 determines whether there is a speaker feature vector with a similarity above a threshold (for example, 60% or more) (S9). If multiple participants in the meeting speak without overlapping time, the similarity to the first spoken audio data after the timer starts is calculated. This is because the scribe often speaks first. Even if multiple participants in the meeting speak with overlapping time, it is thought that the similarity to the scribe's audio data can be calculated by using a certain period of time from the beginning of the audio data for matching. Even if the speaker feature vector of a participant who is not the scribe is matched with the speaker feature vector of the memory unit 40 and the similarity is above the threshold, the character-recognized text data and the speech-recognized text data often do not match, and are rejected in step S11.

[0094] If there is a speaker feature vector with a similarity above a threshold (e.g., 60% or more) (Yes in S9), the speaker recognition unit 29 stores the user ID associated with the speaker feature vector with a similarity above a threshold (e.g., 60% or more) in the user information 42 as the input to the input data storage unit 41 (S10). In other words, the scribe's identification information is stored.

[0095] Next, the recognition result matching unit 30 determines whether at least a portion of the character-recognized text data matches the speech-recognized text data (first text data) used when identifying the person as a writer (S11). In this determination, a portion of the character-recognized text data may be included in the speech-recognized text data, and a portion of the speech-recognized text data may be included in the character-recognized text data.

[0096] If at least part of the recognized text data does not match the recognized speech text data, the scribe and speaker are different, and the process returns to step S2.

[0097] If at least a portion of the character-recognized text data matches the speech-recognized text data, the speech recognition unit 28 switches to the specific speaker speech recognition mode (S12).

[0098] When the system enters speaker-specific speech recognition mode, the display control unit 24 displays a speech input mark 404 at the right edge of the text data displayed by character recognition (S13). This allows the scribe to understand that speech input is possible.

[0099] Next, the speech recognition unit 28 sets the variable N to "2" (S14). The variable N becomes the identification number of the text data to be recognized by speech.

[0100] When the speech recognition unit 28 receives voice input (Yes in S15), it converts the voice into the Nth text data through speech recognition processing (S16).

[0101] Next, the speaker recognition unit 29 calculates the speaker's speaker feature vector through speaker recognition processing, and compares the calculated speaker feature vector with the speaker-specific speaker feature vector stored in the user information 42 to determine the similarity (S17).

[0102] The speaker recognition unit 29 determines whether or not there is a speaker feature vector with a similarity score above a threshold (for example, 60% or higher) (S18).

[0103] If there is a speaker feature vector with a similarity above a threshold (e.g., 60% or more) (Yes in S18), the speaker recognition unit 29 determines whether this speaker is the same as the speaker identified in step S10 (S19). If the speakers are different, the speech-recognized text data should not be displayed immediately after the text data displayed on the display, so the process returns to step S15.

[0104] If the speaker is the same (Yes in S19), the display control unit 24 converts the Nth text data into font data of the same size as the first handwritten recognized text data and displays it at the position of the voice input mark 404 (following the N-1th text data) (S20). Note that the Nth text data and the first handwritten recognized text data do not have to be exactly the same size; the Nth text data may be made larger or smaller, for example, depending on the volume of the voice.

[0105] The display control unit 24 moves the voice input mark 404 to the right of the Nth text data (S21).

[0106] CPU201 increases the variable N by one, and the process returns to step S15.

[0107] Thus, the display device 2 of this embodiment can append and display speech-recognized text data to the handwritten-recognized text data by having the writer speak before the timeout occurs after inputting the handwritten-recognized text data.

[0108] <Ending the specific speaker speech recognition mode> As shown in Figure 10, in the specific speaker speech recognition mode, a speech input mark 404 is displayed at the right edge of the text data. When the scribe clicks the speech input mark 404, the operation reception unit 27 receives the input, and the character recognition unit 23 cancels the specific speaker speech recognition mode. The display control unit 24 erases the speech input mark 404 and returns the speech input transition icon 403 to its initial display state (for example, returning the inverted display to the original display).

[0109] Figure 13 shows the screen with the voice input mark 404 removed. Since the handwriting input icon 401 remains on, the character recognition unit 23 can recognize the handwritten data entered by the writer, and the display control unit 24 can display the handwritten recognized text data.

[0110] Furthermore, if the scribe clicks the voice input mark 404 twice in a row (double-click), the CPU 201 may deactivate the specific speaker voice recognition mode.

[0111] <Main effects> As described above, the display device 2 of the present embodiment collates the text data recognized by voice and the text data recognized by handwriting, and when they match, the display device 2 displays the text data recognized by voice. Therefore, even if a person other than the note-taker speaks, it is possible to suppress the display device 2 from displaying the text.

[0112] In addition, since the text data recognized by voice is displayed following the text data recognized by handwriting, there is no need for the note-taker to specify the display position. Further, since the display device 2 displays the text data recognized by voice in the same size as the text data recognized by handwriting, there is no need for the note-taker to specify the character size in advance (before the speaker speaks).

Example

[0113] In this example, in the case of the specific speaker voice recognition mode, the display device 2 will be described, in which the note-taker moves the voice input mark 404 and causes the text data recognized by voice to be displayed from the position where the voice input mark 404 is moved to. Note that Example 1 and Example 2 can be implemented in combination as appropriate.

[0114] On the screen shown in FIG. 10, the note-taker (for example, the chairperson) can move the voice input mark 404 to an arbitrary position by dragging and dropping the voice input mark 404 with the input means (touch the voice input mark 404 with the input means and move it while keeping contact with the display, and then release the input means from the display).

[0115] FIG. 14 shows an example of the movement of the voice input mark 404. In FIG. 14, the voice input mark 404 is moved under the character "now".

[0116] For example, when the note-taker (for example, the chairperson) says "cool", the voice recognition unit 28 extracts the feature amount of the voice of "cool" input from the microphone (voice data encoded in PCM), specifies the phoneme model, and specifies the word using the pronunciation dictionary, and outputs the text data of "(1)".

[0117] Furthermore, the speaker recognition unit 29 extracts acoustic features from the "kakkoichi" audio data at short intervals of several tens of milliseconds, and generates an acoustic feature vector by representing these values ​​as vectors. The speaker recognition unit 29 then calculates a speaker feature vector from the acoustic feature vector using a UBM model or a speaker feature extraction model.

[0118] The speaker recognition unit 29 then compares this speaker feature vector with the speaker feature vectors for each speaker that have been previously stored in the user information 42 to determine the similarity. When the speaker recognition unit 29 finds a speaker feature vector with a similarity above a threshold (for example, 60% or more) and confirms that it is the speaker feature vector of the scribe, it outputs information indicating that the speaker of "kakkoichi" is the scribe (the same speaker as "kyo").

[0119] When CPU201 confirms that the speaker is a scribe (the same speaker as "today"), display control unit24 displays "(1)" in the same font size as "Today's agenda" at the position of voice input mark 404, and moves voice input mark 404 to the right of this character.

[0120] Figure 15 shows the voice input mark 404 displayed to the right of "(1)". In this way, the voice input mark 404 can be moved to the right edge of the speech-recognized text data.

[0121] Next, when the speaker (for example, the chairman) says "keikakuritsuan", the CPU 201 performs the same processing as above, and the display control unit 24 displays "(1)" to the right of the character "(1)" in the same font size as "(1)", and moves the voice input mark 404 to the right of this string.

[0122] Figure 16 shows the text data "Planning" displayed to the right of "(1)" and the voice input mark 404. In this way, the voice input mark 404 can be moved successively to the right edge of the voice-recognized text data.

[0123] In this embodiment, the scribe moved the voice input mark 404 by drag and drop, but the voice input mark 404 may also be moved by voice input of a command. For example, if the command "newline" is pre-registered, and the text data converted from the scribe's voice data matches "newline", the display control unit 24 moves the voice input mark 404 to the beginning of the line (inserts a newline).

[0124] <Main effects> According to the display device 2 of this embodiment, in addition to the effects of Embodiment 1, the scribe can change the display position of the voice-recognized text data by moving the voice input mark 404. [Examples]

[0125] In this embodiment, a display system 19 in which a server device 12 performs character recognition and speech recognition will be described. Note that Embodiment 3 and Embodiments 1 and 2 can be combined as appropriate.

[0126] Figure 17 shows an example of a schematic configuration diagram of the display system 19. The display device 2 and the server device 12 are connected via a network such as the Internet.

[0127] Figure 18 shows an example of the hardware configuration of server device 12. Server device 12 includes a CPU 301, ROM 302, RAM 303, HD (Hard Disk) 304, HDD (Hard Disk Drive) 305, recording media 306, media I / F 307, display 308, network I / F 309, keyboard 311, mouse 312, CD-ROM (Compact Disc Read Only Memory) drive 314, and bus line 310.

[0128] Of these components, the CPU 301 controls the operation of the entire server device 12. The ROM 302 stores programs used to drive the CPU 301, such as the IPL. The RAM 303 is used as the work area for the CPU 301. The HD 304 stores various data, such as programs. The HDD 305 controls the reading or writing of various data to the HD 304 according to the control of the CPU 301.

[0129] The media interface 307 controls the reading or writing (storage) of data to or from the recording medium 306, such as flash memory. The display 308 displays various information such as cursors, menus, windows, characters, or images. The network interface 309 is an interface for data communication using a network.

[0130] The keyboard 311 is a type of input device equipped with multiple keys for inputting characters, numbers, and various instructions. The mouse 312 is a type of input device for selecting and executing various instructions, selecting processing targets, and moving the cursor. The CD-ROM drive 314 controls the reading or writing of various data to the CD-ROM 313, which is an example of a removable recording medium. The bus line 310 is an address bus, data bus, etc., for electrically connecting each component, such as the CPU 301 shown in Figure 18.

[0131] Figure 19 is an example of a functional block diagram that explains the functions of the display device 2 and the server device 12 by dividing them into blocks. The functions of the display device 2 are a handwritten data receiving unit 21, a drawing data generation unit 22, a display control unit 24, a network communication unit 26, an operation receiving unit 27, and a voice data input receiving unit 31.

[0132] The server device 12 includes a character recognition unit 23, a data recording unit 25, a speech recognition unit 28, a speaker recognition unit 29, a recognition result matching unit 30, and a network communication unit 26-2. The functions of the server device 12 are functions or means realized by any of the components shown in Figure 18 operating according to instructions from the CPU 301 in accordance with a program deployed from the HD 304 onto the RAM 303.

[0133] The network communication unit 26 of the display device 2 transmits handwritten data and voice data to the server device 12. The server device 12 performs the same processing as shown in the flowcharts in Figures 12A and 12B, and transmits the text data entered by the writer via voice to the display device 2.

[0134] Thus, the display system 19 allows the display device 2 and the server device 12 to interactively display text data.

[0135] <Other application examples> Although the best mode for carrying out the present invention has been described above using examples, the present invention is not limited in any way to these examples, and various modifications and substitutions can be made without departing from the spirit of the present invention.

[0136] For example, although this embodiment describes a case where there is one scribe, multiple scribes can write in parallel. After the identification of scribes based on the matching of speaker feature vectors and text data (S2-S11 in Figure 12A) has been performed separately for the first scribe and the second scribe, even if the first and second scribes write and speak in parallel, the text data can be input via speech immediately following the handwritten text data of each scribe.

[0137] Furthermore, although this embodiment describes a display device that can be used as an electronic whiteboard, the display device only needs to be able to display images, and could be, for example, a digital signage device. Alternatively, a projector could be used instead of a display. In this case, although the coordinates of the pen tip were detected by detecting the coordinates of the pen tip on a touch panel in this embodiment, the display device 2 may detect the coordinates of the pen tip using ultrasound. The pen emits ultrasound along with light emission, and the display device 2 calculates the distance based on the arrival time of the ultrasound. The display device 2 can determine the position of the pen based on its direction and distance. The projector draws (projects) the trajectory of the pen as a stroke.

[0138] Furthermore, although an electronic blackboard was described as an example in this embodiment, any information processing device having a touch panel can be suitably applied. Devices having similar functions to an electronic blackboard are also called electronic whiteboards, electronic information boards, interactive boards, etc. Examples of information processing devices equipped with a touch panel include projectors (PJ), output devices such as digital signage, head-up displays (HUD), industrial machinery, imaging devices, sound collection devices, medical equipment, networked home appliances, personal computers (notebook PCs), mobile phones, smartphones, tablet terminals, game consoles, PDAs, digital cameras, wearable PCs, or desktop PCs.

[0139] Furthermore, the configuration examples shown in Figure 4 and other figures are divided according to their main functions to facilitate understanding of the processing performed by the display device 2. The present invention is not limited by the way the processing units are divided or the names of those units. The processing of the display device 2 can be further divided into many more processing units depending on the processing content. It can also be divided so that one processing unit includes even more processing.

[0140] The functions of server device 12 may be distributed and maintained across separate servers, or there may be multiple server devices 12 that cooperate to process information.

[0141] Furthermore, each function of the embodiments described above can be realized by one or more processing circuits. Hereinafter, "processing circuit" as used herein includes processors programmed to execute each function by software, such as processors implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (digital signal processors), FPGAs (field programmable gate arrays), and conventional circuit modules designed to execute each of the functions described above.

[0142] <Additional Claims> [Claim 1] A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the aforementioned handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. A display device. [Claim 2] The display device according to claim 1, wherein if at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data following the first text data. [Claim 3] The display control unit displays the first text data, and within a certain time period thereafter, the speaker recognition unit compares the feature information extracted from the audio data received by the audio data input receiving unit with the feature information of the audio data that has been pre-registered for each user, thereby recognizing the speaker of the audio data received by the audio data input receiving unit. If the speaker recognized by the speaker recognition unit is the scribe of the first text data, The display device according to claim 1 or 2, wherein the speech recognition unit converts the voice data of the scribe into the second text data. [Claim 4] The display device according to claim 3, in which, after the speech recognition unit converts the speech data into the second text data, if the speech data input receiving unit determines that the speech data received is the speech data of the scribe recognized by the speaker recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data, following the first text data. [Claim 5] The character recognition unit determines the size of the first text data based on the size of the handwritten data received by the handwritten data receiving unit. The display device according to any one of claims 1 to 4, wherein the display control unit displays the third text data at the same size as the first text data. [Claim 6] The display device according to claim 3, wherein if at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays a mark at the end of the first text data. [Claim 7] The display device according to claim 6, wherein the display control unit displays the third text data following the first text data, and displays the mark at the end of the third text data. [Claim 8] The input means includes an operation receiving unit that accepts the movement of the mark to any position on the display, The display device according to claim 6 or 7, wherein the display control unit displays text data converted by the speech recognition unit from the speech data of the scribe recognized by the speaker recognition unit at the position of the moved mark. [Explanation of Symbols]

[0143] 2 Display device 19 Display System 23 Character recognition section 28. Voice Recognition Unit 29 Speaker Recognition Unit 30 Recognition Result Verification Unit 31. Voice Data Input Reception Unit [Prior art documents] [Patent Documents]

[0144] [Patent Document 1] Japanese Patent Publication No. 2019-074898

Claims

1. A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. The system includes a speaker recognition unit that, within a certain time after the display control unit displays the first text data, compares feature information extracted from the audio data received by the audio data input receiving unit with feature information of the audio data that has been registered in advance for each user, and recognizes the speaker of the audio data received by the audio data input receiving unit. If the speaker recognized by the speaker recognition unit is the scribe of the first text data, A display device in which the voice recognition unit converts the voice data of the aforementioned scribe into the second text data.

2. A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, The system includes a display control unit that displays a third text data converted from speech data by the speech recognition unit when at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit. The character recognition unit determines the size of the first text data based on the size of the handwritten data received by the handwritten data receiving unit. The display control unit is a display device that displays the third text data in a size based on the size of the first text data.

3. The display device according to claim 1 or 2, wherein if at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data following the first text data.

4. The display device according to claim 1, in which, after the speech recognition unit converts the speech data into the second text data, if the speech data input receiving unit determines that the speech data received is the speech data of the scribe recognized by the speaker recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data, following the first text data.

5. The display device according to claim 1, wherein if at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays a mark at the end of the first text data.

6. The display device according to claim 5, wherein the display control unit displays the third text data following the first text data, and displays the mark at the end of the third text data.

7. The input means includes an operation receiving unit that accepts the movement of the mark to any position on the display, The display device according to claim 6, wherein the display control unit displays text data converted by the speech recognition unit from the speech data of the scribe recognized by the speaker recognition unit at the position of the moved mark.

8. A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. The system includes a speaker recognition unit that, within a certain time after the display control unit displays the first text data, compares feature information extracted from the audio data received by the audio data input receiving unit with feature information of the audio data that has been registered in advance for each user, and recognizes the speaker of the audio data received by the audio data input receiving unit. If the speaker recognized by the speaker recognition unit is the scribe of the first text data, A display system in which the voice recognition unit converts the voice data of the aforementioned scribe into the second text data.

9. A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, The system includes a display control unit that displays a third text data converted from speech data by the speech recognition unit when at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit. The character recognition unit determines the size of the first text data based on the size of the handwritten data received by the handwritten data receiving unit. The display control unit is a display system that displays the third text data in a size based on the size of the first text data.

10. The handwritten data receiving unit receives handwritten data input via an input means, The character recognition unit performs the step of converting the handwritten data into first text data, The voice data input receiving unit takes the step of receiving voice data, The speech recognition unit performs the step of converting the speech data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. The speaker recognition unit includes the step of, within a certain time after the display control unit displays the first text data, comparing the feature information extracted from the voice data received by the voice data input receiving unit with the feature information of the voice data that has been registered in advance for each user, to recognize the speaker of the voice data received by the voice data input receiving unit. If the speaker recognized by the speaker recognition unit is the scribe of the first text data, A display method in which the speech recognition unit converts the voice data of the aforementioned scribe into the second text data.

11. The handwritten data receiving unit receives handwritten data input via an input means, The character recognition unit performs the step of converting the handwritten data into first text data, The voice data input receiving unit performs the step of receiving voice data input, The speech recognition unit performs the step of converting the speech data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. The character recognition unit includes the step of determining the size of the first text data based on the size of the handwritten data received by the handwritten data receiving unit, A display method comprising the display control unit displaying the third text data in a size based on the size of the first text data.

12. Display device, A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit displays the third text data converted by the speech recognition unit from the speech data. Within a certain period of time after the display control unit displays the first text data, the display control unit compares the feature information extracted from the audio data received by the audio data input receiving unit with the feature information of the audio data that has been registered in advance for each user, and functions as a speaker recognition unit that recognizes the speaker of the audio data received by the audio data input receiving unit. If the speaker recognized by the speaker recognition unit is the scribe of the first text data, A program in which the speech recognition unit converts the voice data of the aforementioned scribe into the second text data.

13. Display device, A handwritten data receiving unit that accepts handwritten data input via an input means, A character recognition unit that converts the handwritten data into first text data, A voice data input receiving unit that accepts voice data input, A speech recognition unit that converts the aforementioned audio data into second text data, If at least a portion of the first text data converted by the character recognition unit matches the second text data converted by the speech recognition unit, the display control unit will function as a display control unit that displays the third text data converted by the speech recognition unit from the speech data. The character recognition unit determines the size of the first text data based on the size of the handwritten data received by the handwritten data receiving unit. The display control unit is a program that displays the third text data in a size based on the size of the first text data.

Citation Information

Patent Citations

  • Absolute value circuit

    JP1989095376A

  • Information processing device and information processing program

    JP2019074898A

  • JPP4565658B

  • JPP6859667B

  • JPP6870242B