Display device

The display device addresses the limitation of voice output to selected objects by analyzing and outputting voice data for all touched items, enhancing user interaction and operability by providing auditory feedback.

JP2025158286APending Publication Date: 2025-10-17KYOCERA DOCUMENT SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024060676
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing display technologies limit voice output to only the object selected by a pointer, failing to provide auditory feedback for all items related to a touched position on the display screen.

Method used

A display device comprising a display unit, operation unit, operation position detection, display content detection, voice synthesis, and voice output units, which analyzes and outputs voice data based on touch operations on the display screen, allowing all items related to the touched position to be read aloud.

Benefits of technology

Enables the output of all display items at a touched position by voice, enhancing user interaction by providing auditory confirmation and improving operability through touch operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025158286000001_ABST
    Figure 2025158286000001_ABST
Patent Text Reader

Abstract

To provide a display device which analyzes items associated with a position at which touch operation has been performed on a display screen to output the entire items in voice.SOLUTION: A display device comprises a display unit 21, an operation unit 22, an operation position detection unit 32, a display content detection unit 33, a voice synthesis unit 34, and a voice output unit 27. The display unit 21 has a display screen 21a which displays an image. The operation unit 22 receives touch operation to the display screen 21a. The operation position detection unit 32 detects an operation position on the display screen 21a of the touch operation. The display content detection unit 33 detects a display content of a region including the operation position in the display screen 21a to convert it into character data Dc. The voice synthesis unit 34 synthesizes voice data Ds on the basis of the character data Dc. The voice output unit 27 outputs voice S on the basis of the voice data Ds.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a display device. [Background technology]

[0002] The patent document 1 focuses on selective speech synthesis, which manipulates and loads text-to-speech conversion when an object is selected by a pointer. The selected object is then read aloud by speech synthesis software, conveying the selected data to the listener. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] U.S. Patent No. 10,402,159 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the technology of Patent Document 1, the object to be read aloud by the speech synthesis software is limited to the object selected by the pointer.

[0005] An object of the present invention is to provide a display device that analyzes an item related to a touched position on a display screen and outputs the entire item by voice. [Means for solving the problem]

[0006] The display device of the present invention comprises a display unit, an operation unit, an operation position detection unit, a display content detection unit, a voice synthesis unit, and a voice output unit. The display unit has a display screen that displays an image. The operation unit accepts a touch operation on the display screen. The operation position detection unit detects the operation position on the display screen of the touch operation. The display content detection unit detects display content in an area including the operation position on the display screen and converts it into character data. The voice synthesis unit synthesizes voice data based on the character data. The voice output unit outputs voice based on the voice data. [Effects of the Invention]

[0007] According to the display device of the present invention, it is possible to output by voice all of the items related to the position on the display screen that has been touched. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram showing an image forming apparatus 10 equipped with a display device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the functional configuration of the image forming apparatus 10. [Figure 3] 5 is a flowchart showing main processing by the display device according to the present embodiment. [Figure 4] 10A and 10B are explanatory diagrams illustrating an example of the operation of the display device during a touch operation. [Figure 5] 2 is an explanatory diagram illustrating an example of the display content of a display screen 21a during a touch operation. FIG. [Figure 6] 1 is an explanatory diagram showing an example of use of the display device according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Embodiments of the present disclosure will be described with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference numerals and description thereof will not be repeated.

[0010] 1. Configuration of an image forming apparatus 10 equipped with a display device according to this embodiment First, with reference to Figures 1 and 2, the configuration of an image forming device 10 equipped with a display device that outputs by voice all items related to a position touched on a display screen will be described. Figure 1 is a diagram showing an image forming device 10 equipped with a display device according to an embodiment of the present invention. Figure 2 is a block diagram showing the functional configuration of the image forming device 10. Note that, hereinafter, this display device may also be referred to as a Screen View Audio Reader (SVAR).

[0011] The image forming apparatus 10 is, for example, a copier, a printer, or a multifunction peripheral. In the following, as an example, a case will be described in which the image forming apparatus 10 is a multifunction peripheral having a printer function, a copy function, a facsimile function, and a network communication function.

[0012] 1 and 2, the image forming apparatus 10 includes an image reading unit 11, an image forming unit 12, an operation display unit 23, a FAX communication unit 24, a network communication unit 25, a microphone 26, an audio output unit 27, a storage unit 28, and a control unit 31. Hereinafter, the network communication unit 25 may be referred to as an NW communication unit 25.

[0013] The image reading unit 11 reads an image from an original document M. The image reading unit 11 has an original document table 120 and an original document transport unit 110. The image reading unit 11 reads an image formed on the original document M and generates read data. Specifically, the image reading unit 11 reads an image formed on the original document M transported by the original document transport unit 110, or an image formed on the original document M placed on the original document table 120. Specifically, the image reading unit 11 is composed of an automatic original document feeder, an original document image scanning device (scanner), etc.

[0014] The image forming unit 12 includes an image forming section 220 , a paper feed cassette 230 , a conveying section 240 , and a discharge section 270 .

[0015] The image forming unit 220 forms an image on the recording medium P. For example, the image forming unit 220 forms an image on the recording medium P based on the read data. The image forming unit 220 includes a plurality of toner containers.

[0016] The plurality of toner containers are detachably attached to the image forming apparatus 10. Each of the plurality of toner containers contains toner of a different color. The toner in the toner container is supplied to the image forming unit 220.

[0017] The image forming unit 220 includes an exposure unit, a photosensitive drum, a charging unit, a developing unit, a primary transfer roller, a cleaning unit, an intermediate transfer belt, a secondary transfer roller, and a fixing unit.

[0018] A recording medium P for printing is accommodated in paper feed cassette 230. The recording medium P is transported by transport unit 240. When printing is performed, the recording medium P in paper feed cassette 230 passes through image forming unit 220 and is discharged from discharge unit 270.

[0019] The conveying section 240 conveys the fed recording medium P to the image forming section 220. After the image forming section 220 forms an image on the recording medium P, the conveying section 240 further conveys the recording medium P from the image forming section 220 and discharges the recording medium P to the outside of the image forming apparatus 10.

[0020] The operation display unit 23 is used to allow a user to operate the image forming apparatus 10. The operation display unit 23 includes a display unit 21 and an operation unit 22.

[0021] The display unit 21 has a display screen 21a that displays images. Here, the images include text such as letters and symbols. The display unit 21 is configured with a display such as an LCD (Liquid Crystal Display) or ELD (Electro Luminescence Display) that has a touch panel function.

[0022] The operation unit 22 accepts touch operations on the display screen 21a. In this embodiment in which a touch panel functions as the operation display unit 23, the display unit 21 and a part of the operation unit 22 may be integrated. Furthermore, the operation unit 22 of the operation display unit 23 may have a touch panel and physical buttons. Note that, hereinafter, the operation display unit 23 may be referred to as a panel.

[0023] The FAX communication unit 24 transmits or receives FAX images. Specifically, the FAX communication unit 24 transmits and receives image data and the like to and from other image forming devices and facsimile devices (neither of which are shown) via a network. The received FAX images are printed on a recording medium P by the image forming unit 12. The FAX images are also written as read data into a storage area of ​​the storage unit 28.

[0024] The NW communication unit 25 is capable of communicating with electronic devices equipped with communication devices that use the same communication method (protocol). Specifically, the NW communication unit 25 communicates with other electronic devices via a network such as a LAN (Local Area Network). The NW communication unit 25 is, for example, a communication interface equipped with a communication module such as a LAN board.

[0025] Microphone 26 receives audio input by picking up sounds generated around image forming apparatus 10. Microphone 26 is arranged, for example, in the same area as operation display unit 23. Microphone 26 outputs a signal indicating the input audio.

[0026] The audio output unit 27 outputs audio S based on the audio data Ds. Examples of the audio output unit 27 include, but are not limited to, a speaker, an earphone, and a headphone.

[0027] The storage unit 28 is, for example, a hard disk drive (HDD) or a solid state drive (SSD). The storage unit 28 may include a random access memory (RAM) and a read only memory (ROM). The storage unit 28 stores various data and a control program for controlling the operation of each unit of the image forming apparatus 10. The control program is executed by the control unit 31. In addition, the image read by the image reading unit 11 is written as read data to a predetermined data area of ​​the storage unit 28.

[0028] The control unit 31 is a hardware circuit configured by a processor such as a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), etc. The control unit 31 controls the operation of each operating unit of the image forming apparatus 10 by having the processor read and execute a control program stored in the storage unit 28.

[0029] The control unit 31 includes an operation position detection unit 32, a display content detection unit 33, a voice synthesis unit , an execution control unit 35, and a display control unit .

[0030] The operation position detector 32 detects the operation position on the display screen 21a where the touch operation is performed.

[0031] The display content detection unit 33 detects the display content X in the area R including the operation position on the display screen 21a and converts it into character data Dc.

[0032] The voice synthesis unit 34 synthesizes voice data Ds based on the character data Dc.

[0033] Therefore, the display content X of the area R including the touched position is output as sound S. As a result, it becomes possible to easily and reliably perform the touch operation while confirming it auditorily.

[0034] When the touch operation received by the operation unit 22 is the first touch operation T1, the voice output unit 27 outputs the voice S. When the touch operation is a second touch operation T2 different from the first touch operation T1, the voice output unit 27 does not output the voice S. Therefore, when the first touch operation T1 is performed among various touch operations, the display content X of the region R including the operation position is output as the voice S, and in the case of the second touch operation T2, the voice S is not output. As a result, it is possible to avoid the display content X being output as the voice S when not intended.

[0035] When the operation unit 22 receives the first touch operation T1 on the region R including the operation position again during or after the output of the voice S, the execution control unit 35 causes the function corresponding to the display content X to be executed. Therefore, by performing the same touch operation again at substantially the same position during or after the output of the voice S of the display content X of the region R including the touched position, the function corresponding to the display content X is executed. As a result, since the corresponding function can be immediately executed even during the output of the voice S, the operability is improved.

[0036] Alternatively, the execution control unit 排除 the period before the elapse of the first predetermined time from the start of the output of the voice S and the period after the elapse of the second predetermined time longer than the first predetermined time from the start of the output, and when the operation unit 22 receives the first touch operation T1 on the region R including the operation position again, the execution control unit 35 causes the function corresponding to the display content X to be executed. Therefore, by excluding the period before the elapse of the first predetermined time t1 from the start of the output of the voice S of the display content X of the region R including the touched position and the period after the elapse of the second predetermined time t2 from the start of the output (where t1 < t2), and performing the same touch operation again at substantially the same position, the function corresponding to the display content X is executed. As a result, since the corresponding function can be immediately executed even during the output of the voice S, the operability is improved.

[0037] The first touch operation T1 may include a single tap operation. Therefore, when a single tap operation is performed on the display screen 21a, the display content X of the region R including the touched position is output as the voice S. As a result, it becomes possible to perform the touch operation more easily and surely by hearing.

[0038] When the touch operation received by the operation unit 22 is the second touch operation T2, the execution control unit 35 may execute a function corresponding to the display content X. Therefore, the function corresponding to the display content X is executed by the second touch operation T2 on the display screen 21a. As a result, when auditory confirmation is not required, the function corresponding to the display content X can be immediately executed by the second touch operation T2, thereby improving operability.

[0039] The second touch operation T2 may include a double-tap operation. Therefore, when a double-tap operation is performed on the display screen 21a, a function corresponding to the display content X is executed. As a result, when auditory confirmation is not required, the function corresponding to the display content X can be immediately executed by a double-tap operation, thereby improving operability.

[0040] The display content X includes the name or character string C of the operation area B that is displayed closest to the operation position. Therefore, the display content X includes the name or character string C of the operation area B that is displayed closest to the operation position. As a result, the operation position may be somewhat inaccurate and a range designation operation is not required, further improving operability.

[0041] The display control unit 36 ​​changes the display mode of the display content X on the display screen 21a. Specifically, the display control unit 36 ​​may display the name or character string C of the operation area B on the display screen 21a so as to be emphasized. Therefore, the display mode of the display content X in the area R including the position where the touch operation was performed is changed. As a result, the detected name or character string C of the operation area B can be visually confirmed.

[0042] The display content detection unit 33 converts the display content X into character data Dc using optical character recognition (OCR) and natural language processing (NLP). Therefore, when detecting the display content X and converting it into character data Dc, not only optical character recognition but also natural language processing is used. As a result, the accuracy of the conversion into character data Dc is improved.

[0043] Optical character recognition refers to a technology that optically reads handwritten or printed characters using an image scanner or digital camera and converts them into digital character data that can be used by a computer. Natural language processing refers to a technology that allows a computer to process natural language that is used on a daily basis. Combining optical character recognition with natural language processing, as in this embodiment, significantly improves the character recognition rate.

[0044] 2. Operation and Main Processing of the Display Device According to the Present Embodiment Next, the operation and main processing of the display device according to this embodiment will be described with reference to Figs. 3 to 5. Fig. 3 is a flowchart showing the main processing by the display device according to this embodiment. Fig. 4 is an explanatory diagram illustrating the operation of the display device during a touch operation. Fig. 5 is an explanatory diagram illustrating the display contents of the display screen 21a during a touch operation. As mentioned above, the panel in the following description corresponds to the operation display unit 23.

[0045] This display device is configured to analyze the text in the panel frame and read the scanned text aloud for transmission to the display device. This display device is realized by a service with the SVAR function described above running in the background. However, this realization method is not limited to this.

[0046] In step S1, the display device acquires screen data from all areas on the display screen 21a where there is readable text, and the panel sends that screen data to the display device.

[0047] In step S2, the display device prepares everything and officially starts the SVAR service for the user. After that, the display device waits for the user to select by tapping text, buttons, etc. on the display screen 21a. Prior to this, when the user touches the panel (step S11), the panel notifies the user of the service (step S12).

[0048] In step S3, when the user taps on the display screen 21a, the display device analyzes the tapped area and scans and highlights the text near the tapped area. Note that the analysis of the tapped area uses optical character recognition and natural language processing, as described above.

[0049] In step S4, once the display device has analyzed the text, it then sends the analyzed text to the voice output unit 27. The voice output unit 27 reads it out loud to the user. Note that the SVAR service runs in the background for this reading (step S13).

[0050] For example, as shown in Fig. 4, when the user taps the "Job Box" button on the display screen 21a, the voice output unit 27 outputs the voice "Job Box." Also, as shown in Fig. 5, when the user taps the "Status / Job Cancel" button on the display screen 21a, the voice output unit 27 outputs the voice "Status Job Cancel." At this time, the "Status / Job Cancel" button can be highlighted by inverting the button, for example, to make it possible to visually confirm the response to the touch operation.

[0051] In step S5, it is determined whether the same item is selected again during reading, and if the result of the determination is YES, the process proceeds to step S7, and if not, the process proceeds to step S6.

[0052] In step S6, the display device waits for the user to tap again to indicate that they want to tap this item. This tap should occur several seconds after dictation. The display device then determines whether the same item was selected again some time after reading. If the determination is YES, the process proceeds to step S7; otherwise, the process returns to step S1.

[0053] In step S7, in order to execute a tap on the panel, a tap is sent to the panel (operation display unit 23) to instruct the user to select that item.

[0054] In step S8, the display device determines whether to repeat this service. If the service is to be repeated because the user did not tap again, the process returns to step S1; otherwise, the process ends.

[0055] If the panel changes the data (step S10), the process proceeds to step S1. However, if the keyboard is used or a protected screen is accessed (step S14), the process proceeds to step S9 and the SVAR service is suspended for privacy reasons, as it is not intended to be read aloud.

[0056] The user can also close the service as desired by tapping "Close floating window." In this way, when the SVAR service termination is instructed (step S15), the process ends.

[0057] 3. Example of use of the display device according to this embodiment Next, a usage example of the display device according to this embodiment will be described with reference to Fig. 6. Fig. 6 is an explanatory diagram showing a usage example of the display device according to this embodiment. This usage example shows the start of a voice navigation service. However, the usage example is not limited to this.

[0058] First, the user U instructs the display device to start the SVAR service (step S61). In response to this instruction, the SVAR service returns a status (whether the service start was successful or unsuccessful) (step S62) and displays a panel view of the enabled service to the user (step S63).

[0059] When the SVAR service acquires service access (audio output unit 27 and reading) (steps S64 and S65), permission for the functions required by the SVAR is granted. The audio output unit 27 and reading service return the status of the access to the service and whether the function has accessed the SVAR (step S66).

[0060] The SVAR service obtains access to the panel (step S67), and the panel returns the access status (step S68) and the first frame of the panel (step S69).

[0061] The present invention can be embodied in various other forms without departing from its spirit or main features. While the above description focuses on the reservation of a "conference" in order to clarify the gist of the invention, similar applications are possible for the preparation time for work reservations, etc. Furthermore, the above-described embodiments are merely illustrative in all respects and should not be interpreted as limiting. The scope of the present invention is defined by the claims and is not bound by the specification. Furthermore, all modifications and variations within the scope of the claims are within the scope of the present invention. [Industrial Applicability]

[0062] The contents of the present disclosure can be used for display devices. [Explanation of symbols]

[0063] 10 Image forming device 21 Display section 21a Display screen 22 Control section 23 Operation display section 28 Memory section 31 Control Unit 32 Operation position detection unit 33 Display content detection unit 34 Speech synthesis unit 35 Execution Control Unit 36 Display control unit

Claims

1. a display unit having a display screen for displaying an image; an operation unit that accepts a touch operation on the display screen; an operation position detection unit that detects an operation position on the display screen of the touch operation; a display content detection unit that detects display content in an area including the operation position on the display screen and converts the display content into character data; a voice synthesis unit that synthesizes voice data based on the character data; an audio output unit that outputs audio based on the audio data; A display device comprising:

2. 2. The display device according to claim 1, wherein when the touch operation received by the operation unit is a first touch operation, the audio output unit outputs the audio, and when the touch operation is a second touch operation different from the first touch operation, the audio output unit does not output the audio.

3. 3. The display device according to claim 2, further comprising: an execution control unit that, when the operation unit again receives the first touch operation on the area including the operation position during or after output of the sound, executes a function corresponding to the display content.

4. 3. The display device according to claim 2, further comprising an execution control unit that executes a function corresponding to the display content when the operation unit again receives the first touch operation on the area including the operation position, excluding a period before a first predetermined time has elapsed since the start of output of the audio and a period after a second predetermined time longer than the first predetermined time has elapsed since the start of output.

5. The display device according to claim 1 , wherein the display content includes a name or a character string of an operation area displayed closest to the operation position.

6. a display control unit that changes a display mode of the display content on the display screen; The display device according to claim 5 , wherein the display control unit displays the name of the operation area or the character string on the display screen in an emphasized manner.

7. The display device according to claim 1 , wherein the display content detection unit converts the display content into the character data by optical character recognition and natural language processing.

8. The display device according to claim 2 , wherein the first touch operation includes a single tap operation.

9. The display device according to claim 3 , wherein when the touch operation received by the operation unit is the second touch operation, the execution control unit causes a function corresponding to the display content to be executed.

10. The display device according to claim 9 , wherein the second touch operation includes a double tap operation.

Citation Information

Patent Citations

  • Audible user interface system

    US10402159B1