Processing system
The processing system addresses the usability challenges of voice operation systems by using attribute information to manage voice control operations between distant voice control and image processing devices, thereby improving user interaction and system efficiency.
Patent Information
- Application Number
- JP2021094539
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-04
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2041-06-04
AI Technical Summary
Existing voice operation systems for external multi-function devices face usability issues when the voice control device is not in close proximity to the image forming apparatus, making it difficult for users to manage and execute commands effectively.
A processing system that includes an image processing apparatus communicable with a voice control apparatus, featuring receiving means for voice information, generating means for commands, and outputting means for response information. This system outputs specific attribute information indicating whether the voice control device is attached to the image processing device, allowing for appropriate control and user interaction.
The system improves the overall usability of the information processing system by enabling effective voice control operations even when the voice control device is not physically close to the image forming apparatus, enhancing user convenience and system efficiency.
Smart Images

Figure 0007696759000001 
Figure 0007696759000002 
Figure 0007696759000003
Abstract
Description
Technical Field
[0001] The present invention relates to a processing system that accepts voice operations, and particularly to a processing system that accepts voice operations for an external multi-function device. to Mu to Mu
Background Art
[0002] Conventionally, as methods for performing input operations on an information processing apparatus, there are voice operation devices such as smart speakers and manual operation devices using an LUI (Local User Interface).
[0003] Also, in various information processing apparatuses that accept a plurality of input operations, a technique is known for improving user usability during input operations by preventing other input operations from being accepted while any one input operation is being executed (Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the technique of Patent Document 1 is a technique premised on both the manual operation device and the voice operation device being in the vicinity of the user, and may not provide appropriate control when applied in a case where the voice operation device is installed separately from the image forming apparatus having the manual operation device.
[0006] For example, a user may voice-input a command such as "Print today's documents" to a voice-operated device, where multiple files are targeted. At this time, if there is an image forming apparatus equipped with a manual operation device near the voice-operated device, the user can view a list of the targeted files on the screen of the manual operation device. However, such an image forming apparatus may be located at a position away from the voice-operated device installed at the user's desk. In such a case, even if a list is displayed on the manual operation device of the image forming apparatus in response to the user input of a command targeting multiple files as described above, the user cannot view the screen.
[0007] Therefore, the present invention , place provides a processing system capable of improving the usability of the entire user system. Mu
Means for Solving the Problem
[0008] The processing system according to claim 1 of the present invention is a processing system having an image processing apparatus and being communicable with a voice control apparatus, and an information processing device comprising: receiving means for receiving voice information generated based on voice received from a user from the voice control apparatus; generating means for generating a command based on the voice information; and outputting means for outputting response information as a response to the command, The information processing device is wherein the outputting means outputs first type of response information of first attribute information, the attribute information set in the voice control device and and the outputting means outputs second type of response information of second attribute information different from the first attribute information, the first which is characterized in that. the first in the voice control device when the attribute information is in the case as a response to the command when the attribute information is in the case as a response to the command types of The attribute information is information indicating whether the voice control device is attached to the image processing device. The first attribute information is information indicating that the voice control device is attached to the image processing device, and the second attribute information is information indicating that the voice control device is not attached to the image processing device
Advantages of the Invention
[0010] According to the present invention, the usability of the entire information processing system can be improved.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the present invention according to the claims, and not all combinations of the features described in the present embodiments are essential for the solution means of the present invention.
[0013] (Embodiment 1) <Overall Configuration of the Information Processing System> FIG. 1 is an overall configuration diagram of the information processing system according to this embodiment.
[0014] As shown in FIG. 1, the information processing system includes an MFP 101 (image processing device), a smart speaker 102 (voice control device), a cloud server 103 (information processing device), and a print server 106. Further, the MFP 101, the smart speaker 102, and the cloud server 103 can communicate via a network 104 and a gateway 105.
[0015] The MFP 101 is a multifunction device (usable) having a plurality of functions such as a copy function, a scan function, a print function, a FAX function, etc., but may be a printer or a scanner having a single function. Details of the hardware configuration of the MFP 101 will be described later with reference to FIG. 2.
[0016] The smart speaker 102 acquires the voice of the user 107 with a microphone 308 (FIG. 3), encodes the acquired voice into voice data (voice information), and then transmits it to the cloud server 103 via the network 104 and the gateway 105. Further, when the smart speaker 102 receives voice synthesis data from the cloud server 103 via the network 104 and the gateway 105, it plays back the voice synthesis data with a speaker 310 (FIG. 3). Details of the hardware configuration of the smart speaker 102 will be described later with reference to FIG. 3.
[0017] The cloud server 103 performs voice recognition on the voice data of the user 107 transmitted from the voice control device 100, such as voice data of "job execution", "job setting", etc., and generates job information based on the voice recognition result. Thereafter, the cloud server 103 transmits the generated job information to the MFP 101 via the network 104 and the gateway 105. Thereafter, the cloud server 103 generates voice synthesis data for notifying the user 107 that the job information has been transmitted to the MFP 101, and transmits it to the smart speaker 102 via the network 104 and the gateway 105.
[0018] The cloud server 103 communicates with the MFP 101 and the smart speaker 102 using an IP address and a MAC address.
[0019] The network 104 connects the MFP 101, the smart speaker 102, the cloud server 103, and the gateway 105 to each other. As a result, various data such as voice data acquired by the smart speaker 102 and job information such as print jobs and scan jobs generated by the cloud server 103 are transmitted and received via the network 104.
[0020] The gateway 105 is, for example, a wireless LAN router compliant with the IEEE802.11 standard series. However, it may have the ability to operate according to other wireless communication methods. Also, instead of a wireless LAN router, it may be a wired LAN router compliant with the Ethernet standard typified by 10BASE-T, 100BASE-T, 1200BASE-T, etc., and may have the ability to operate according to other wired communication methods. Note that the IEEE802.11 standard series includes a series of standards belonging to IEEE802.11 such as IEEE802.11a and IEEE802.11b.
[0021] Note that the following embodiments do not limit the invention according to the claims, and not all combinations of features described in the embodiments are essential for the solution means of the invention.
[0022] The print server 106 is a server that manages print data, and transmits print data in response to a request from the MFP 101 via the network 104 to the MFP 101.
[0023] <Configuration of MFP> FIG. 2 is a block diagram showing the hardware configuration of the MFP 101.
[0024] As shown in FIG. 2, the MFP 101 includes a controller unit 200, an operation panel 209, a print engine 211, and a scanner 213.
[0025] The controller unit 200 includes a CPU 202, a RAM 203, a ROM 204, a storage 205, a network I / F 206, a display controller 207, an operation I / F 208, a print controller 210, and a scan controller 212. These components are communicably connected to each other by a system bus 201.
[0026] The CPU 202 is a central processing unit that controls the overall operation of the controller unit 200. The CPU 202 reads out a control program stored in the ROM 204 or the storage 205 and performs various controls such as reading control and printing control.
[0027] The RAM 203 is a volatile memory used as the main memory of the CPU 202. It is used as a work area and also as a temporary storage area for expanding various control programs stored in the ROM 204 and the storage 205.
[0028] The ROM 204 is a non-volatile memory that stores control programs executable by the CPU 202.
[0029] The storage 205 is a storage device with a larger capacity compared to the RAM 203. Print data, image data, various programs, and various setting information are stored in the storage 205.
[0030] In this embodiment, the MFP 101 has one CPU 202 execute each process shown in the flowchart described below using one memory (RAM 203), but other modes are also possible. For example, it is also possible to have a plurality of CPUs, RAMs, ROMs, and storages cooperate to execute each process shown in the flowchart described below. Further, some processes may be executed using a hardware circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).
[0031] The network I / F 206 is an interface for causing the MFP 101 to communicate with an external device via the network 104. The MFP 101 analyzes print data received via the network I / F 206 by a software module (PDL analysis unit, not shown) for analyzing the print data stored in the storage 205 or the ROM 204. The PDL analysis unit generates data for printing by the print engine 211 based on print data expressed in various types of page description languages.
[0032] The display controller 207 is connected to the operation panel 209 composed of an LCD touch panel, and performs display control of the screen of the operation panel 209 in accordance with an instruction from the CPU 202.
[0033] The operation I / F 208 is connected to the operation panel 209. When the user 107 operates the operation panel 209 according to the screen displayed on the operation panel 209, the operation I / F 208 detects an event corresponding to the user operation and transmits the detected event to the CPU 202.
[0034] The print controller 210 is connected to the print engine 211. The control command and the image data to be printed are transferred to the print engine 211 via the print controller 210.
[0035] The print engine 211 forms an image on a sheet based on the received control command and the image data to be printed. The printing method of the print engine 211 may be an electrophotographic method or an inkjet method. In the case of the electrophotographic method, an electrostatic latent image is formed on a photoreceptor, then developed with toner, the toner image is transferred to a sheet, and the transferred toner image is fixed to form an image. On the other hand, in the case of the inkjet method, ink is ejected to form an image on a sheet.
[0036] The scan controller 212 is connected to the scanner 213 and receives the image data generated by the scanner 213 reading the image on the sheet from the scanner 213. The image data generated by this scanner 213 is stored in the storage 205. Also, the MFP 101 has a copy function of forming an image on a sheet by the print engine 211 based on the image data generated by the scanner 213.
[0037] The scanner 213 has a document feeder (not shown) and can read while transporting the sheets placed on the document feeder one by one.
[0038] <Configuration of the smart speaker> FIG. 3 is a block diagram showing the hardware configuration of the smart speaker 102.
[0039] As shown in FIG. 3, the smart speaker 102 includes a controller unit 300, a microphone 308 as an audio input device, a speaker 310 as an audio output device, and an LED 312 as a notification device.
[0040] The controller unit 300 includes a CPU 302, a RAM 303, a ROM 304, a storage 305, a network I / F 306, a mic I / F 307, an audio controller 309, and a display controller 311. These components are communicably connected to each other by a system bus 301.
[0041] The CPU 302 is a central processing unit that controls the operation of the entire controller unit 300. The CPU 302 expands the control program stored in the storage 305 into the RAM 303 and performs various controls such as audio input control and audio output control.
[0042] The RAM 303 is a volatile memory used as the main memory of the CPU 202. The RAM 303 is used as a work area and also as a temporary storage area for expanding various control programs stored in the storage 305.
[0043] The ROM 304 is a non-volatile memory and stores the startup program for the CPU 302.
[0044] The storage 305 is a large-capacity storage device (e.g., an SD card) compared to the RAM 303. The storage 305 stores the control program for the smart speaker 102 that the CPU 302 executes. Note that the storage 305 may be replaced with a flash ROM other than an SD card, or may be replaced with another storage device having a function equivalent to that of an SD card.
[0045] When starting up, such as when power is turned on, the CPU 302 executes the startup program stored in the ROM 304. This startup program is for reading out the control program stored in the storage 305 and expanding it onto the RAM 303. When the CPU 302 executes the startup program, it then executes the control program expanded onto the RAM 303 to perform control. Also, the CPU 302 stores and reads / writes data used when executing the control program onto the RAM 303. Various settings and the like necessary when executing the control program can be further stored on the storage 305 and read / written by the CPU 302. Further, the CPU 302 communicates with external devices via the network I / F 306 and the network 104.
[0046] The network I / F 306 is configured to include a circuit and an antenna for connecting to the network 104 according to a wireless communication method compliant with the IEEE802.11 standard series to enable the smart speaker 102 to communicate with an external device. However, it may be a wired communication method compliant with the Ethernet standard instead of a wireless communication method, and is not limited to a wireless communication method.
[0047] The microphone I / F 307 is connected to the microphone 308, converts the voice uttered by the user 107 input from the microphone 308 into encoded voice data, and holds it in the RAM 303 according to the instruction of the CPU 302.
[0048] The microphone 308 is a small MEMS microphone mounted on, for example, a smartphone or the like, but may be replaced with other devices as long as it can acquire the voice of the user 107. Further, it is preferable to arrange and use three or more microphones 308 at predetermined positions so that the arrival direction of the voice emitted by the user 107 can be calculated. However, the present embodiment can be realized even if there is one microphone 308, and the present invention does not insist on three or more.
[0049] The audio controller 309 is connected to the speaker 310, converts voice data into an analog voice signal according to an instruction from the CPU 302, and outputs voice through the speaker 310.
[0050] The speaker 310 reproduces a response sound of a device indicating that the smart speaker 102 is responding to the voice of the user 107, and a voice synthesized by the cloud server 103. The speaker 310 may use a general-purpose device for reproducing voice.
[0051] The display controller 311 is connected to the LED 312 and controls the display of the LED 312 according to an instruction from the CPU 302. In the present embodiment, the display controller 311 mainly performs lighting control of the LED 312 for indicating that the smart speaker 102 correctly inputs the voice of the user 107. The LED 312 is an LED that emits light in a color (for example, blue) that the user 107 can recognize lighting and extinguishing. The LED 312 is a general-purpose device. Note that, instead of the LED 312, an LUI capable of displaying characters and pictures may be used.
[0052] <Configuration of Cloud Server> FIG. 4 is a block diagram showing the hardware configuration of the cloud server 103.
[0053] As shown in FIG. 4, the cloud server 103 includes a CPU 402, a RAM 403, a ROM 404, a storage 405, and a network I / F 406. These components are communicably connected to each other by a system bus 401.
[0054] The CPU 402 is a central processing unit that controls the operation of the entire cloud server 103. The CPU 402 expands the control program stored in the storage 405 into the RAM 403 and executes voice recognition processing and the like.
[0055] The RAM 403 is a volatile memory used as the main memory of the CPU 402. The RAM 403 is used as a work area and also as a temporary storage area for expanding various control programs stored in the storage 405.
[0056] The ROM 404 is a non-volatile memory in which a startup program for the CPU 402 is stored.
[0057] The storage 405 is a large-capacity storage device (for example, a hard disk drive: HDD) compared to the RAM 403. The storage 405 stores a control program for the cloud server 103 that is executed by the CPU 402. Note that the storage 405 may be a solid state drive (SSD) or the like, and may be replaced with another storage device having a function equivalent to that of a hard disk drive.
[0058] When starting up, such as when power is turned on, the CPU 402 executes the startup program stored in the ROM 404. This startup program is for reading out the control program stored in the storage 405 and expanding it onto the RAM 403. When the CPU 402 executes the startup program, it then continues to execute the control program expanded onto the RAM 403 to perform control. Also, the CPU 402 stores and reads / writes the data used when executing the control program onto the RAM 403. Various settings necessary when executing the control program can be further stored on the storage 405 and are read and written by the CPU 402. In addition, the CPU 402 communicates with other devices on the network 104 via the network I / F 406 and the gateway 105.
[0059] <Configuration of Print Server> FIG. 5 is a block diagram showing the hardware configuration of the print server 106.
[0060] As shown in FIG. 5, the print server 106 includes a CPU 502, a RAM 503, a ROM 504, a storage 505, a network I / F 506, a RIP processing unit 507, and an encoding unit 508. These components are communicably connected to each other by a system bus 501.
[0061] The CPU 502 is a central processing unit that controls the overall operation of the print server 106. The CPU 502 expands the control program stored in the storage 505 onto the RAM 503 and executes management of print data and the like.
[0062] The RAM 503 is a volatile memory used as the main memory of the CPU 502. The RAM 503 is used as a work area and also as a temporary storage area for expanding various control programs stored in the storage 505.
[0063] The ROM 504 is a non-volatile memory and stores the startup program of the CPU 502.
[0064] Storage 505 is a large-capacity storage device (e.g., hard disk drive: HDD) compared to RAM 503. The storage 505 stores a control program for the print server 106 and print data that are executed by the CPU 502. Note that the storage 505 may be a solid state drive (SSD) or the like, and may be replaced with another storage device having a function equivalent to that of the hard disk drive.
[0065] When starting up such as when power is turned on, the CPU 502 executes a startup program stored in the ROM 504. This startup program is for reading out the control program stored in the storage 505 and expanding it onto the RAM 503. When the CPU 502 executes the startup program, it subsequently executes the control program expanded onto the RAM 503 to perform control. Also, the CPU 502 stores and reads / writes data used when executing the control program onto the RAM 503. Various settings necessary when executing the control program can be further stored on the storage 505 and read / written by the CPU 502. Further, the CPU 502 communicates with external devices via the network I / F 506 and the network 104.
[0066] The RIP processing unit 507 generates raster data from the PDL data received from an external device.
[0067] The encoding unit 508 changes the raster data generated by the RIP processing unit 507 into print data or a data format in a form supported by the MFP 101.
[0068] <Functional Configuration of the Device Control Program of the MFP> FIG. 6 is a block diagram showing the functional configuration of a device control program 600 executed by the MFP 101.
[0069] The device control program 600 of the MFP 101 is stored in the ROM 204 as described above, and is expanded and executed onto the RAM 203 by the CPU 202 at startup.
[0070] The device control program 600 includes a data transmission / reception unit 601, a data analysis unit 602, a job control unit 603, a data management unit 604, a display unit 605, an operation target determination unit 606, a scan unit 607, and a printer unit 608. As shown in FIG. 6, the job control unit 603 is connected to the scan unit 607 and the printer unit 608, and the data analysis unit 602 is connected to the data transmission / reception unit 601, the job control unit 603, the data management unit 604, the display unit 605, and the operation target determination unit 606.
[0071] The data transmission / reception unit 601 performs data transmission and reception with other devices on the network 104 via the network I / F 206 using TCP / IP. Also, the data transmission / reception unit 601 performs data transmission and reception with the cloud server 103 via the gateway 105 on the network 104. Specifically, it receives device operation data generated by the cloud server 103 and transmits various notifications to the cloud server 103. Here, the various notifications include a screen update notification indicating that the screen display content showing the job execution result or the response result to the device operation data has been updated, and a job execution status notification indicating the status of the job. The details of the content of the screen update notification and the job execution status notification will be described in the sequence diagrams of FIGS. 10 and 20 below.
[0072] The data analysis unit 602 converts the device operation data received by the data transmission / reception unit 601 into commands for each module within the device control program 600 to communicate. Then, the data analysis unit 602 transmits it to either the job control unit 603, the data management unit 604, or the display unit 605 according to the content of the command.
[0073] The job control unit 603 issues control instructions for the scan unit 607 and the printer unit 608. For example, when the user 107 presses the start key while the display unit 605 is displaying the copy function screen, the job control unit 603 receives the parameters of the copy job and the job start instruction from the operation target determination unit 606. Thereafter, the job control unit generates the scan job parameters and the print job parameters from the received parameters of the copy job, and transmits the scan job parameters to the scan unit 607 and the print job parameters to the printer unit 608. Thereby, the job control unit 603 controls the scan unit 607 and the printer unit 608 to print the image data read by the scanner 213 on the sheet by the print engine 211. Note that since the mechanisms of scan and print control are not the main subject, further description thereof will be omitted.
[0074] The data management unit 604 stores and manages various data such as the work data generated during the execution of the device control program 600 and the setting parameters required for each device control in predetermined areas on the RAM 203 and the storage 205. For example, the data management unit 604 stores and manages job data composed of combinations of each setting item and setting value of the job executed by the job control unit 603 described later, language settings which are information on the language to be displayed on the operation panel 209, and the like. In addition, the data management unit 604 stores and manages the authentication information necessary for communication with the gateway 105, the device information necessary for communication with the cloud server 103, and the image data of the object to be imaged by the MFP 101. Further, the data management unit 604 stores and manages the screen control information used by the display unit 605 for screen display control and the operation target determination information used by the operation target determination unit 606 for determining the operation target for each screen displayed by the display unit 605.
[0075] The display unit 605 controls the operation panel 209 via the display controller 207. More specifically, the display unit 605 displays UI components (buttons, pull-down lists, check boxes, etc.) operable by the user 107 on the operation panel 209, or updates the screen of the operation panel 209 based on screen display control information. For example, the display unit 605 acquires a language dictionary corresponding to the language setting stored in the data management unit 604 from the storage 205, and displays text data based on the language dictionary on the screen of the operation panel 209.
[0076] The operation target determination unit 606 acquires the touched coordinates on the operation panel 209 via the operation I / F 208, and determines the UI components operable by the user 107 currently displayed on the operation panel 209 as the operation targets. Further, the operation target determination unit 606 reads out the screen display control information corresponding to the UI components determined as the operation targets, and determines the processing content at the time of accepting an operation based on the information. For example, the operation target determination unit 606 issues an instruction to update the display content of the screen to the display unit 605, or transmits the parameters of the job set by the user operation and the start instruction of the job to the job control unit 603.
[0077] The scan unit 607 executes a scan with the scanner 213 via the scan controller 212 based on the scan job parameters transmitted from the job control unit 603, and stores the read image data in the data management unit 604.
[0078] The printer unit 608 executes printing of the image data stored in the data management unit 604 with the print engine 211 via the print controller 210 based on the print job parameters transmitted from the job control unit 603.
[0079] <Functional Configuration of the Voice Control Program of the Voice Control Device> FIG. 7 is a block diagram showing the functional configuration of a voice control program 700 executed by the smart speaker 102.
[0080] The voice control program 700 of the smart speaker 102 is stored in the storage 305 as described above, and is expanded and executed on the RAM 303 when the CPU 302 is activated.
[0081] The voice control program 700 includes a data transmission / reception unit 701, a data management unit 702, a voice control unit 703, a voice acquisition unit 704, a voice playback unit 705, a display unit 706, a voice operation start detection unit 707, and an utterance end determination unit 708. As shown in FIG. 6, the voice control unit 703 is connected to all of the other modules of the voice control program 700.
[0082] The data transmission / reception unit 701 transmits and receives data by TCP / IP with other devices on the network 104 via the network I / F 306. In addition, the data transmission / reception unit 701 transmits and receives data with the cloud server 103 via the gateway 105 on the network 104. Specifically, the data transmission / reception unit 701 transmits the voice data of the voice uttered by the user 107 acquired by the voice acquisition unit 704 described later to the cloud server 103, or receives the voice synthesis data generated on the cloud server 103, which is a response to the user 107.
[0083] The data management unit 702 stores and manages various data such as work data generated during the execution of the voice control program 700 in a predetermined area on the storage 305. For example, volume setting data of the voice to be played back by the voice playback unit 705 described later, authentication information necessary for communication with the gateway 105, MFP 101, and each device information necessary for communication with the cloud server 103 are stored and managed.
[0084] The voice acquisition unit 704 converts the analog voice of the user 107 near the smart speaker 102 acquired by the microphone 308 into voice data and temporarily stores it. The voice of the user 107 is converted into voice data in a predetermined format such as MP3, for example, and temporarily stored on the RAM 303 as encoded voice data for transmission to the cloud server 103. The start and end timings of the processing of the voice acquisition unit 704 are managed by the voice control unit 703. Also, the encoding of the voice data may be in a general-purpose streaming format, and the encoded voice data may be sequentially transmitted by the data transmission / reception unit 701.
[0085] The voice playback unit 705 plays back the voice synthesis data (voice message) received by the data transmission / reception unit 701 through the audio controller 309 using the speaker 310. The voice playback timing of the voice playback unit 705 is managed by the voice control unit 703.
[0086] The display unit 706 controls the display of the LED 312 via the display controller 311. For example, when the voice operation start detection unit 707 detects that there is a voice operation, the display unit 706 controls the LED 312 to light up. The display timing of the display unit 706 is managed by the voice control unit 703.
[0087] The voice operation start detection unit 707 detects the wake word uttered by the user 107 or the pressing of an operation start key (not shown) of the smart speaker 102, and sends an operation start notification to the voice control unit 703. Here, the wake word is a predetermined voice word. The voice operation start detection unit 707 constantly detects the wake word from the analog voice of the user 107 near the smart speaker 102 acquired by the microphone 308. The user 107 can operate the MFP 101 by speaking the wake word and then speaking what he / she wants to do. The voice processing after the voice operation start detection unit 707 detects the wake word will be described later.
[0088] The speech end determination unit 708 determines the end timing of the processing in the voice acquisition unit 704. For example, when the voice of user 107 stops for a predetermined time (e.g., 3 seconds), it is determined that the speech of user 107 has ended, and a speech end notification is sent to the voice control unit 703. Note that the determination of the end of speech may be made not based on the time without speech (hereinafter referred to as blank time), but from a predetermined phrase spoken by user 107 (e.g., "yes", "no", "OK", "cancel", "end", "start", "begin", etc.). Also, when user 107 speaks a predetermined phrase, the speech end determination unit 708 may determine that the speech has ended without waiting for the predetermined time. Further, the determination of the end of speech may be made not by the smart speaker 102 but by the cloud server 103. In this case, the cloud server 103 may determine the end of speech from the meaning and context of the speech content of user 107.
[0089] The voice control unit 703 is the center of control and controls other modules in the voice control program 700 to operate in cooperation with each other. Specifically, it controls the start and end of the processing of the voice acquisition unit 704, the voice playback unit 705, and the display unit 706. Also, after voice data is acquired by the voice acquisition unit 704, it controls the voice data to be transmitted to the cloud server 103 by the data transmission / reception unit 701. Further, after receiving the voice synthesis data from the cloud server 103 by the data transmission / reception unit 701, it controls the voice playback unit 705 to play back the voice synthesis data.
[0090] Here, the start and end timing of the processing of the voice acquisition unit 704, the voice playback unit 705, and the display unit 706 will be described.
[0091] When the voice control unit 703 receives an operation start notification from the voice operation start detection unit 707, it starts the processing of the voice acquisition unit 704. Also, when it receives an utterance end notification from the utterance end determination unit 708, it ends the processing of the voice acquisition unit 704. For example, assume that user 107 utters a wake word and then utters "want to copy". At this time, the voice operation start detection unit 707 detects the analog voice of the wake word and sends an operation start notification to the voice control unit 703. When the voice control unit 703 receives the operation start notification, it controls to start the processing of the voice acquisition unit 704. The voice acquisition unit 704 acquires the subsequent analog voice of "want to copy", converts it into voice data, and temporarily stores it. When the utterance end determination unit 708 determines that there is a blank time for a predetermined time after the utterance of "want to copy" by user 107, it sends an utterance end notification to the voice control unit 703. When the voice control unit 703 receives the utterance end notification, it ends the processing of the voice acquisition unit 704. Hereinafter, the state from when the voice acquisition unit 704 starts processing to when it ends is called the utterance processing state of the smart speaker 102. The display unit 706 turns on and displays the LED 312 while the smart speaker 102 is in the utterance processing state.
[0092] After an utterance end notification is sent from the utterance end determination unit 708, a dialogue session with the cloud server 103 is started. Specifically, the voice control unit 703 reads out the voice data temporarily stored by the voice acquisition unit 704, controls the data transmission / reception unit 701 to send the voice data to the cloud server 103, and then waits for a response from the cloud server 103. The response from the cloud server 103 is, for example, a response message consisting of a header part indicating that it is a response and voice synthesis data. When the voice control unit 703 receives the response message by the data transmission / reception unit 701, it controls the voice playback unit 705 to play back the voice synthesis data as a response process. The voice synthesis data is, for example, "The copy screen will be displayed". Hereinafter, the state from when the utterance end determination unit 708 determines that the utterance has ended to when the playback of the voice synthesis data by the voice playback unit 705 ends is called the response processing state of the smart speaker 102. The display unit 706 blinks and displays the LED 312 while the smart speaker 102 is in the response processing state.
[0093] After the response process, while the dialogue session with the cloud server 103 continues, the user 107 can continuously speak out what he / she wants to do without uttering the wake word. When the cloud server 103 determines that the dialogue session has ended, it sends a dialogue session end notification to the smart speaker 102. When the voice control unit 703 receives the dialogue session end notification, it ends the dialogue session with the cloud server 103. Hereinafter, the state from the end of the dialogue session until the next dialogue session starts is called the standby state of the smart speaker 102. Also, the state until the smart speaker 102 receives an operation start notification from the voice operation start detection unit 707 is called the always-on standby state of the smart speaker 102. The display unit 706 turns off the LED 312 while the smart speaker 102 is in the standby state or the always-on standby state.
[0094] <Functional Configuration of the Voice Data Conversion Control Program of the Cloud Server> FIG. 8 is a block diagram showing the functional configuration of a voice data conversion control program 800 executed by the cloud server 103.
[0095] The voice data conversion control program 800 of the cloud server 103 is stored in the storage 405 as described above, and is expanded and executed on the RAM 403 when the CPU 402 is started.
[0096] The voice data conversion control program 800 includes a data transmission / reception unit 801, a data management unit 802, a device operation data generation unit 803, and a voice data conversion unit 808. The voice data conversion unit 808 includes a voice recognition unit 804, a morphological analysis unit 805, and a voice synthesis unit 807. As shown in FIG. 8, each module of the voice data conversion control program 800 is connected to each other.
[0097] The data transmission / reception unit 801 transmits and receives data via TCP / IP with other devices on the network 104 through the network I / F 406 and the gateway 105. Specifically, the data transmission / reception unit 801 receives the voice data of the voice uttered by the user 107 from the smart speaker 102, or transmits the text data determination result generated by the voice recognition process in the voice recognition unit 804.
[0098] The data management unit 802 stores and manages various data such as the work data generated in the execution of the voice data conversion control program 800 and the parameters required for the voice recognition process in the voice recognition unit 804 in a predetermined area on the storage 405. For example, the voice recognition unit 804 stores and manages the acoustic model and language model for converting the voice data received by the data transmission / reception unit 801 into text data in a predetermined area on the storage 405. Also, the data management unit 802 stores and manages the dictionary for performing morphological analysis of text data in the morphological analysis unit 805 in a predetermined area on the storage 405. Also, the data management unit 802 stores and manages the voice database for performing voice synthesis in the voice synthesis unit 807 in a predetermined area on the storage 405. Also, the data management unit 802 stores and manages each device information necessary for communicating with the smart speaker 102 and the MFP 101 in a predetermined area on the storage 405.
[0099] The device operation data generation unit 803 generates device operation data based on the morphological analysis result of the voice recognition data output from the voice data conversion unit 808.
[0100] When the voice data of the voice uttered by the user 107 received by the data transmission / reception unit 801 is input from the data transmission / reception unit 801, the voice recognition unit 804 performs voice recognition processing for converting the input voice data into voice recognition data which is text data. The voice recognition processing uses an acoustic model to convert the input voice data into phonemes, and further uses a language model to convert the phonemes into actual text data. Note that there may be a plurality of languages for the input voice data. Therefore, the voice recognition processing may determine the language of the input voice data and use a first voice recognition method for converting it into text data conforming to the language. Also, the voice recognition processing may use a second voice recognition method for converting the input voice data into phonemes using acoustic models of a plurality of languages, and then converting and outputting it into text data for each language using the corresponding language model. In the case of the second voice recognition method, in order to convert it into text data of a plurality of languages, the voice recognition unit 804 outputs voice recognition data composed of text data and its language setting as a result of the voice recognition processing.
[0101] In this embodiment, the languages of the input voice data are Japanese and English. The voice recognition data in Japanese consists of text data composed of one or more kana and its language setting "Japanese". Also, the voice recognition data in English consists of text composed of one or more alphabets and its language setting "English". However, other methods may be used as the voice recognition processing for converting voice data into voice recognition data, and it is not limited to the above-described method. Since the details of the voice recognition processing are not the main point, further explanation is omitted here.
[0102] The morpheme analysis unit 805 performs morpheme analysis on the voice recognition data converted by the voice recognition unit 804 in accordance with its language setting. Morpheme analysis derives a morpheme sequence from a dictionary having information such as the grammar and part-of-speech of the language, and further discriminates the part-of-speech of each morpheme. The morpheme analysis unit 805 can be realized, for example, using known morpheme analysis software such as JUMAN, Chatan, and MeCab. Since the morpheme analysis software is a known technology, detailed explanation here is omitted.
[0103] The voice synthesis unit 807 generates voice synthesis data for notifying the user 107 of various notifications. This voice synthesis data is transmitted to the smart speaker 102 via the data transmission / reception unit 801.
[0104] <Functional Configuration of Print Data Control Program of Print Server> FIG. 9 is a block diagram showing the functional configuration of a print data control program 900 executed in the print server 106.
[0105] The print data control program 900 of the print server 106 is stored in the storage 505 as described above, and is expanded and executed on the RAM 503 when the CPU 502 is activated.
[0106] The print data control program 900 includes a data transmission / reception unit 901, a control unit 902, and a print data storage unit 903. As shown in FIG. 9, each module of the print data control program 900 is connected to each other.
[0107] The data transmission / reception unit 901 transmits and receives data by TCP / IP to and from other devices on the network 104 via the network I / F 506. The data transmission / reception unit 901 receives a job list reception command from the MFP 101.
[0108] The control unit 902 is a central processing unit for controlling the print server 106. By executing processing based on the program stored in the print data storage unit 903, the processing related to the functions of the print server 106 is realized.
[0109] The print data storage unit 903 is a storage device such as a hard disk or an SSD, and stores various programs, print jobs, etc. Further, the print data storage unit 903 also functions as an auxiliary storage device for the control unit 902.
[0110] <Control Sequence at the Time of Voice Input by the User 107 to a Smart Speaker Not Near the MFP> Figure 10 is a sequence diagram showing the interaction between control programs executed by each device constituting the information processing system in this embodiment when the smart speaker 102 is not in the vicinity of the MFP 101.
[0111] In the sequence shown in Figure 10, it is assumed that the smart speaker 102, the MFP 101, the cloud server 103, and the print server 106 are in a state where they can communicate with each other. Also, it is assumed that the MFP 101 is displaying a home screen that can call functions such as copy, scan, print, and FAX after being powered on and started up.
[0112] First, in step S1001, when the voice operation start detection unit 707 of the voice control program 700 detects that the user 107 has given a voice operation start instruction to the smart speaker 102, the process proceeds to step S1002. The voice operation start instruction by the user 107 is given by the user 107 uttering (inputting) a wake word toward the microphone 308 of the smart speaker 102 or by pressing the operation start key of the smart speaker 102.
[0113] In step S1002, the voice playback unit 705 of the voice control program 700 plays back the synthesized voice data notifying the start-up state and notifies the user 107 that the voice operation has started.
[0114] In step S1003, the display unit 706 of the voice control program 700 turns on the LED 312 to indicate that the smart speaker 102 has entered the speech processing state (the dialogue session has started). At the same time, the processing of the voice acquisition unit 704 is started.
[0115] In step S1004, after the voice acquisition unit 704 of the voice control program 700 detects a job execution instruction from the user 107 to the smart speaker 102, when the blank time has elapsed for a predetermined time and the speech end determination unit 708 determines that the speech has ended, the process proceeds to step S1005. Here, the job execution instruction refers to data obtained by converting analog voices such as "Print today's materials" or "Tell me the remaining toner level" uttered by the user 107 following the input of the wake word in step S1001 into digital voice data by the voice acquisition unit 704.
[0116] In step S1005, the display unit 706 of the voice control program 700 blinks the LED 312 to indicate that the speech has ended and the smart speaker 102 has entered the response processing state. At the same time, the processing of the voice acquisition unit 704 is terminated.
[0117] In step S1006, the data transmission / reception unit 701 (first attribute information notification means) of the voice control program 700 transmits the attribute information held by the data management unit 702 of the voice control program 700 to the cloud server 103. In this process, the attribute information is information indicating whether the smart speaker 102 is attached to the MFP 101. Note that FIG. 10 shows an example in which the attribute information indicating that the smart speaker 102 is not attached to the MFP 101 is transmitted here.
[0118] In step S1007, the data transmission / reception unit 701 (job notification means) of the voice control program 700 transmits the job execution instruction detected in step S1004 to the cloud server 103.
[0119] In step S1008, at the cloud server 103, the data transmission / reception unit 801 of the voice data conversion control program 800 performs processing according to the received attribute information and voice data.
[0120] In steps S1009a and S1009b, the device control program 600 of the MFP 101 and the print data control program 900 of the print server 106 cooperate to perform processing according to the attribute information and job information received from the cloud server 103. Although the details will be described later, through this processing, a job list information response including the information of the job list generated by the print server 106 is transmitted from the MFP 101 to the cloud server 103.
[0121] In step S1010, in the cloud server 103, the voice synthesis unit 807 of the voice data conversion control program 800 performs processing corresponding to the job list information response received from the MFP 101 and generates voice synthesis data. The details of steps S1008, S1009a, and S1010 will be described using the flowcharts of FIGS. 11, 12, and 13.
[0122] In step S1011, the data transmission / reception unit 801 transmits to the smart speaker 102 a voice synthesis data generated in step S1010 and a dialogue session end notification for notifying the end of the dialogue session with the user 107.
[0123] In step S1012, the voice playback unit 705 plays back the voice synthesis data received in step S1011. Thereby, for example, the voice synthesis data "Print the latest files today" generated in step S1010 is played back through the speaker 310.
[0124] In step S1013, in response to the dialogue session end notification transmitted from the cloud server 103 in step S1011, the display unit 706 of the voice control program 700 turns off the LED 312 to indicate that the smart speaker 102 has entered the standby state.
[0125] In step S1014, in response to the dialogue session end notification transmitted from the cloud server 103 in step S1011, the voice control unit 703 of the voice control program 700 ends the dialogue session with the cloud server 103. As a result, the smart speaker 102 shifts to the standby state.
[0126] In addition, in the sequence of FIG. 10, even when the smart speaker 102 is in the response processing state, that is, even when the LED 312 is blinking, the wake word can always be input. As a result, the user 107 can forcibly end the dialogue session by speaking "cancel" or "abort" etc. following the utterance of the wake word during the continuation of the dialogue session.
[0127] <Flowchart of the process of step S1008 in FIG. 10 in the cloud server> FIG. 11 is a detailed flowchart of the process of step S1008 in FIG. 10, that is, the process for the attribute information and voice data received by the data transmission / reception unit 801 of the voice data conversion control program 800 in the cloud server 103.
[0128] In step S1101, the CPU 402 receives the attribute information and voice data notified from the smart speaker 102 by the data transmission / reception unit 801.
[0129] In step S1102, the CPU 402 (inquiry notification means) generates an information acquisition job command to be notified to the MFP 101 from the voice data received in step S1101. Specifically, first, the voice recognition unit 804 of the voice data conversion control program 800 converts the voice data into phonemes. Next, the voice recognition unit 804 determines the language of the voice data, and converts the phonemes into text data using the language model of the determined language. Then, the morphological analysis unit 805 of the voice data conversion control program 800 performs morphological analysis of the text data to determine the content of the operation instructed by the user 107. After that, the device operation data generation unit 803 of the voice data conversion control program 800 generates an information acquisition job command (job information: device operation data) based on the determination result.
[0130] For example, when the voice job execution instruction from the user 107 to the smart speaker 102 is an instruction such as "Print today's materials", the device operation data generation unit 803 generates a print job list acquisition command as the information acquisition job command. Specifically, first, a keyword for searching the print data managed by the print server 106 is extracted. After that, a print job list acquisition command for inquiring the print server 106 about the print jobs searched with the keyword is generated. Note that the information acquisition job command is not limited to the print job list acquisition command exemplified here. For example, depending on the job execution instruction from the user 107 to the smart speaker 102, a FAX job list acquisition command for inquiring the FAX jobs in the storage 205 to the MFP 101 may be generated as the information acquisition job command. Note that in this embodiment, the FAX job is a job for printing the received FAX data.
[0131] In step S1103, the CPU 402 (second attribute information notification means) notifies the MFP 101 of the attribute information received in step S1101 and the information acquisition job command generated in step S1102.
[0132] <Flowchart of the processing of step S1010 in FIG. 10 in the cloud server> FIG. 12 is a detailed flowchart of the processing corresponding to the job list information response received by the data transmission / reception unit 801 of the voice data conversion control program 800 in the cloud server 103, that is, the processing of step S1010 in FIG. 10.
[0133] In step S1201, the CPU 402 (control change means) refers to the attribute information received in step S1101 and checks whether the smart speaker 102 is attached to the MFP 101. If it is not attached (NO in step S1201), the smart speaker 102 determines that it is not near the MFP 101 and transitions to step S1202. On the other hand, if it is attached (YES in step S1201), the smart speaker 102 determines that it is near the MFP 101 and transitions to step S1207.
[0134] In step S1202, the CPU 402 acquires the information of the first job in the job list included in the job list information response transmitted from the MFP 101 (hereinafter referred to as "first job information").
[0135] In step S1203, the CPU 402 generates job information based on the first job information acquired in step S1202. Here, a print job command, a FAX job command, etc. are generated as job information according to the first job information.
[0136] In step S1204, the CPU 402 notifies the MFP 101 of the job information generated in step S1203. The job control unit 603 of the device control program 600 performs the execution process of the job corresponding to this job information (step S1009a). In addition, when the job information notified here is another job command for executing a function other than the FAX function and the print function of the MFP 101, the attribute information received in step S1101 is re-notified.
[0137] In step S1205, the CPU 402 performs a voice synthesis process to generate voice synthesis data for notifying the user 107 that the print job command has been notified. For example, voice synthesis data consisting of a message such as "Among the files of today's date, the newest file has been printed" is generated. Then, the process proceeds to step S1207. When the CPU 402 notifies the smart speaker 102 of the voice synthesis data generated in step S1205, this process ends.
[0138] In step S1206, the CPU 402 performs a voice synthesis process to generate voice synthesis data according to the process. For example, voice synthesis data for notifying the user 107 that the job list has been displayed on the operation panel 209 of the MFP 101, such as "Please select the file to be printed from the operation panel for the files of today's date", is generated. Then, the process proceeds to step S1207. In step S1207, when the CPU 402 notifies the smart speaker 102 of the voice synthesis data generated in step S1206, this process ends.
[0139] <Flowchart of the process of step S1009a in FIG. 10 in the MFP> FIG. 13 is a detailed flowchart of the process of step S1009a in FIG. 10, that is, the process executed every time the data transmission / reception unit 601 of the device control program 600 in the MFP 101 receives job information.
[0140] In step S1301, the CPU 202 receives job information from the cloud server 103.
[0141] In step S1302, the CPU 202 checks whether the job information received in step S1301 is a print job list acquisition command for inquiring about a print job. If it is a print job list acquisition command (YES in step S1302), the process proceeds to step S1303. If it is another command (NO in step S1302), the process proceeds to step S1308.
[0142] In step S1303, the CPU 202 (print job list acquisition means) notifies the print server 106 of a job list acquisition instruction to inquire about the job list of the corresponding job in response to the print job list acquisition command received in step S1301. In response to this job list acquisition instruction, the print server 106 generates a job list (step S1009b in FIG. 10) and transmits a job list information response including the job list to the MFP 101.
[0143] In step S1304, the CPU 202 receives the job list information response from the print server 106 from the print server 106.
[0144] In step S1305, the CPU 202 checks whether the smart speaker 102 is installed in the MFP 101 based on the attribute information (step S1103 in FIG. 11) transmitted together with the print job list acquisition command (information acquisition job command). If it is not installed (NO in step S1305), the process proceeds to step S1306. If it is installed (YES in step S1305), the process proceeds to step S1307.
[0145] In step S1306, the CPU 202 (response transmission means) notifies the cloud server 103 of the job list information response (command response) acquired from the print server 106 in step S1304 and ends this process.
[0146] In step S1307, the CPU 202 (display means) displays the job list included in the job list information response acquired in step S1304 on the operation panel 209 in a list and ends this process.
[0147] In step S1308, the CPU 202 checks whether the job information received in step S1301 is a print job command. If it is a print job command (YES in step S1308), the process proceeds to step S1309. If not (NO in step S1308), the process proceeds to step S1310.
[0148] In step S1309, the CPU 202 starts the printing process according to the print job command received in step S1301 and ends this process.
[0149] In step S1310, the CPU 202 checks whether the job information received in step S1301 is a FAX job list acquisition command for inquiring about FAX jobs. In the case of a FAX job list acquisition command (YES in step S1310), it proceeds to step S1311; otherwise (NO in step S1310), it proceeds to step S1314.
[0150] In step S1311, the CPU 202 checks whether the smart speaker 102 is installed in the MFP 101 based on the attribute information (step S1103 in FIG. 11) transmitted together with the FAX job list acquisition command (information acquisition job command). If it is not installed (NO in step S1311), it proceeds to step S1312; if it is installed (YES in step S1311), it proceeds to step S1313.
[0151] In step S1312, the CPU 202 (job list generation means) generates a job list from the FAX jobs held in the storage 205, notifies the cloud server 103 of the job list information response including the job list, and ends this process.
[0152] In step S1313, the CPU 202 generates a preview image of the received FAX data that is the printing target in each FAX job held in the storage 205, displays the preview images in a list on the operation panel 209, and ends this process.
[0153] In step S1314, the CPU 202 checks whether the job information received in step S1301 is a FAX job command. If it is a FAX job command (YES in step S1314), the process proceeds to step S1315. If it is another job command for executing functions other than the FAX function and the printing function of the MFP 101 (NO in step S1314), the process proceeds to step S1316.
[0154] In step S1315, the CPU 202 reads out the FAX data stored in the storage 205 according to the FAX job command received in step S1301, starts the FAX process for printing, and ends this process.
[0155] In step S1316, the CPU 202 checks whether the smart speaker 102 is installed in the MFP 101 based on the attribute information (refer to step S1204 in FIG. 12) transmitted together with the other job command received in step S1301. If it is not installed (NO in step S1316), the process proceeds to step S1317. If it is installed (YES in step S1316), the process proceeds to step S1318.
[0156] In step S1317, the CPU 202 notifies that the received other job command cannot be executed and ends this process.
[0157] In step S1318, the CPU 202 displays on the operation panel 209 a function setting screen for executing the received other job command and ends this process.
[0158] <Control sequence during voice input by user 107 to smart speaker near MFP> FIG. 20 is a sequence diagram showing the interaction between control programs executed by each device constituting the information processing system in this embodiment when the smart speaker 102 is near the MFP.
[0159] In the sequence shown in FIG. 20, it is assumed that the smart speaker 102, the MFP 101 located in its vicinity, the cloud server 103, and the print server 106 are in a state where they can communicate with each other. Also, it is assumed that the MFP 101 is displaying a home screen that can call functions such as copy, scan, print, and FAX after being powered on and activated.
[0160] Hereinafter, among the sequences in FIG. 20, steps different from the sequence in FIG. 10 will be described. Specifically, in FIG. 20, step S2001 is executed instead of steps S1008 and S1010 in FIG. 10, step S2002 is executed instead of step S1009a in FIG. 10, and step S2003 is executed instead of step S1011. Also, in FIG. 20, step S2004 is executed instead of steps S1012 to S1014 in FIG. 10, and step S2005 is also executed.
[0161] Since the other steps in FIG. 20 are the same as those in FIG. 10, duplicate explanations using the same step numbers are omitted.
[0162] In this sequence, first, when the processes up to steps S1001 to S1007 described above in FIG. 10 are executed, the process proceeds to step S2001.
[0163] In step S2001, in the cloud server 103, the voice data conversion unit 808 performs processing according to the attribute information received by the data transmission / reception unit 801 of the voice data conversion control program 800 and the voice data. The details of this processing are almost the same as the flowchart in FIG. 11. However, this processing ends after performing the processing in step S1206 in FIG. 12 and generating voice synthesis data according to the processing.
[0164] In steps S2002 and S1009b, the device control program 600 of the MFP 101 and the print data control program 900 of the print server 106 cooperate to perform processing according to the attribute information and job information received from the cloud server 103. The details of this processing are almost the same as the flowchart in FIG. 13. However, in this processing, it directly proceeds from step S1304 to step S1307, and also directly proceeds from step S1311 to step S1313, and after displaying a list of the job list on the operation panel 209 of the MFP 101, this processing ends. Also, execution instructions such as those in steps S1308 and S1314 are performed when the user 107 operates the operation panel 209 of the MFP 101 in step S2005. For this reason, when the result in step S1302 is NO, it directly proceeds to step S1310, and when the result in step S1310 is NO, this processing ends as it is.
[0165] In step S2003, the data transmission / reception unit 801 transmits the voice synthesis data generated in step S2001 to the smart speaker 102. It is received from the cloud server 103. Also, at this time, a dialogue session end notification for notifying the smart speaker 102 to end the dialogue session with the user 107 is also transmitted to the smart speaker 102.
[0166] In step S2004, in the smart speaker 102, the voice playback unit 705 of the voice control program 700 plays back the voice synthesis data received in step S2003. Thereby, for example, the smart speaker 102 notifies the user 107 of a voice message saying "Please select the file to be printed".
[0167] In step S2005, the operation target determination unit 606 of the device control program 600 detects whether the user 107 has selected / executed a job on the job list that is displayed in a list on the operation panel 209 of the MFP 101 in step S2002. When such detection is made, the operation target determination unit 606 notifies the job control unit 603 of the device control program 600 of the detection result. Based on this detection result, the job control unit 603 performs job execution processing (step S2002).
[0168] As described above, the operation of the information processing system by voice control is changed according to the attribute information, that is, the information indicating whether the smart speaker 102 is attached to the MFP 101.
[0169] For example, when the user 107 performs a voice operation of "Print today's documents" on the smart speaker 102, the operation of the information processing system is different depending on whether the smart speaker 102 is near the MFP 101 or not.
[0170] As shown in FIG. 21(A), when the smart speaker 102 is near the MFP 101, the user 107 is in a state where they can immediately operate the operation panel 209 of the MFP 101. Therefore, the information processing system displays the job list in a list on the operation panel 209 of the MFP 101 and plays a voice message of "Please select the file to be printed" on the smart speaker 102.
[0171] On the other hand, as shown in FIG. 21(B), when the smart speaker 102 is not near the MFP 101, the user 107 is in a state where they cannot immediately operate the operation panel 209 of the MFP 101. Therefore, the information processing system causes the MFP 101 to print the top job of the job list and plays a voice message of "The latest file has been printed." on the smart speaker 102.
[0172] Thereby, it becomes possible to improve the usability of the entire information processing system.
[0173] (Example 2) Next, Example 2 will be described.
[0174] The hardware configuration of this example is different from that of Example 1 in that the smart speaker 102 is provided with an LUI that accepts touch input from the user 107 in addition to the hardware configuration shown in FIG. 3. The software configuration of this example is the same as that of Example 1.
[0175] Also, in Example 1, the data transmitted from the smart speaker 102 to the cloud server 103 together with the voice data was only the attribute information (Steps S1006 and S1007 in FIG. 10). On the other hand, in this example, the data includes not only the attribute information but also the presence / absence information (device configuration information) of the LUI of the smart speaker 102.
[0176] Therefore, hereinafter, the same reference numerals will be given to the same hardware configuration and software configuration as those in Example 1, and duplicate explanations will be omitted.
[0177] <Control Sequence at the Time of Voice Input by the User to a Smart Speaker Having an LUI but Not in the Vicinity of the MFP> FIG. 15 is a sequence diagram showing the interaction between the control programs executed by the respective devices constituting the information processing system in this example when the smart speaker 102 is not in the vicinity of the MFP 101.
[0178] In the sequence shown in FIG. 15, it is assumed that the smart speaker 102, the MFP 101, the cloud server 103, and the print server 106 are in a communicable state with each other. Also, it is assumed that the MFP 101 is in a state where it displays a home screen that can call functions such as copy, scan, print, and FAX after starting up with the power ON.
[0179] Next, among the sequences in FIG. 15, the steps different from the sequence in FIG. 10 will be described. Specifically, in FIG. 15, device configuration information is transmitted (step S1501) at the same timing as the transmission of attribute information and voice information (steps S1006 and S1007 in FIG. 10). Also, step S1502 is executed instead of steps S1008 and S1010 in FIG. 10, and step S1503 is executed instead of step S1009a in FIG. 10. Further, steps S1504 to S1508 are executed before step S1011.
[0180] For the other steps in FIG. 15, since they are the same as those in FIG. 10, duplicate explanations using the same step numbers are omitted.
[0181] In this sequence, first, when the processes up to steps S1001 to S1005 described above in FIG. 10 are executed, the process proceeds to step S1501.
[0182] In step S1501, the data transmission / reception unit 701 (device configuration information notification means) of the voice control program 700 transmits device configuration information indicating that the smart speaker 102 has an LUI to the cloud server 103. Note that the device configuration information is held by the data management unit 702 of the voice control program 700.
[0183] Similar to FIG. 10, in steps S1006 and S1007, the data transmission / reception unit 701 of the voice control program 700 transmits attribute information and voice data to the cloud server 103. In this process, the attribute information is information indicating whether the smart speaker 102 is attached to the MFP 101. Note that in FIG. 15, an example is shown in which attribute information indicating that the smart speaker 102 is not attached to the MFP 101 is transmitted here.
[0184] In step S1502, in the cloud server 103, the voice data conversion unit 808 performs processing according to the device configuration information, attribute information, and voice data received by the data transmission / reception unit 801 of the voice data conversion control program 800. Details of the processing in step S1502 will be described later with reference to FIG. 16.
[0185] In steps S1503 and S1009b, the device control program 600 of the MFP 101 and the print data control program 900 of the print server 106 cooperate to perform processing according to the attribute information and job information received from the cloud server 103. The details of steps S1503 and S1009b will be described later with reference to the flowcharts in FIGS. 16 and 17. As a result, a job list information response including the information of the job list generated by the print server 106 is transmitted from the MFP 101 to the cloud server 103. Also, the voice synthesis unit 807 of the voice data conversion control program 800 performs processing corresponding to the job list information response received from the MFP 101 and generates voice synthesis data (step S1502).
[0186] In step S1504, the data transmission / reception unit 801 transmits the job list included in the job list information response transmitted from the print server 106 to the smart speaker 102.
[0187] In step S1505, the data transmission / reception unit 801 transmits the voice synthesis data generated by the processing corresponding to the job list information response generated by the cloud server 103 to the smart speaker 102. Also, the data transmission / reception unit 801 transmits a dialogue session end notification for notifying the end of the dialogue session with the user 107 to the smart speaker 102.
[0188] In step S1506, in the smart speaker 102, the voice playback unit 705 of the voice control program 700 plays back the voice synthesis data transmitted in step S1505. For example, the voice synthesis data "Print the latest files today" generated in step S1009b is played back through the speaker 310. Also, the display unit 706 displays the job list transmitted in step S1504 on the LUI of the smart speaker 102.
[0189] In step S1507, in the smart speaker 102, the user 107 receives an instruction for job selection / execution from the job list displayed in a list on the display unit 706 using the LUI.
[0190] In step S1508, in the smart speaker 102, the data transmission / reception unit 701 transmits the execution instruction received in step S1512 to the cloud server 103. The cloud server 103 performs processing according to this execution instruction and generates synthesized voice data according to the content of the processing (step S1502).
[0191] After that, steps S1011 to S1014 are executed in the same manner as in FIG. 10.
[0192] Note that, similar to the sequence in FIG. 10, in the sequence in FIG. 15 as well, even when the smart speaker 102 is in the response processing state, that is, even when the LED 312 is blinking, the wake word can always be input. Thus, the user 107 can forcibly end the dialogue session by uttering "cancel" or "abort" etc. after uttering the wake word during the continuation of the dialogue session.
[0193] <Flowchart of the processing of step S1502 in FIG. 15 in the cloud server> FIG. 16 is a detailed flowchart of the processing of step S1502 in FIG. 15, that is, the processing corresponding to the device configuration information, attribute information, and voice data received by the data transmission / reception unit 801 of the voice data conversion control program 800 in the cloud server 103.
[0194] In step S1601, the CPU 402 first performs the processing of steps S1101 to S1103 in FIG. 11 and checks whether the information acquisition job command generated in step S1102 is a print job list acquisition command. If it is a print job list acquisition command (YES in step S1601), the process proceeds to step S1602, and if it is another command (NO in step S1601), the process proceeds to step S1613.
[0195]
[0195] In step S1602, the CPU 402 obtains the job list included in the job list information response transmitted from the MFP 101 by sending a print job list acquisition command to the MFP 101 (step S1103). Specifically, the data transmitted to the MFP 101 here is the print job list acquisition command and the data 1401 including attribute information shown in FIG. 14.
[0196] In step S1603, the CPU 402 refers to the attribute information obtained from the smart speaker 102 and checks whether the smart speaker 102 is installed in the MFP 101. If it is not installed (NO in step S1603), the process proceeds to step S1604; otherwise, the process proceeds to step S1612.
[0197] In step S1604, the CPU 402 refers to the device configuration information obtained from the smart speaker 102 and checks whether the LUI is installed in the smart speaker 102. If it is installed (YES in step S1604), the process proceeds to step S1605; otherwise, the process proceeds to step S1608.
[0198] In step S1605, the CPU 402 notifies the smart speaker 102 of the job list obtained in step S1602. Then, the LUI of the smart speaker 102 displays the job list notified from the cloud server 103 in step S1605 (step S1506 in FIG. 15).
[0199] In step S1606, the CPU 402 performs a voice synthesis process to generate voice information to be notified to the smart speaker 102, for example, voice synthesis data such as "The job list has been displayed. Please select the file to be printed." Then, it proceeds to step S1607, notifies the smart speaker 102 of the voice synthesis data generated in step S1606, and ends this process. Thereafter, the voice playback unit 705 of the smart speaker 102 plays back the voice synthesis data notified from the cloud server 103 in step S1607 (step S1506 in FIG. 15).
[0200] Thereby, the user 107 can easily select the file to be printed from the job list displayed on the LUI in response to the voice notification from the smart speaker 102 such as "The job list has been displayed. Please select the file to be printed."
[0201] In step S1608, the CPU 402 acquires information of the first job in the job list acquired in step S1602 (hereinafter referred to as "first job information").
[0202] In step S1609, the CPU 402 generates a print job command based on the first job information acquired in step S1608.
[0203] In step S1610, the CPU 402 notifies the MFP 101 of the print job command (job information) generated in step S1609.
[0204] In step S1611, the CPU 402 performs a voice synthesis process to generate voice synthesis data for notifying the user 107 that the print job command has been notified. For example, voice synthesis data consisting of a message such as "The newest file among the files of today's date has been printed" is generated. Then, it proceeds to step S1607, notifies the smart speaker 102 of the voice synthesis data generated in step S1611, and ends this process.
[0205] In step S1612, the CPU 402 performs a voice synthesis process to generate voice synthesis data to be notified to the smart speaker 102. For example, voice synthesis data consisting of messages such as "The job list has been displayed. Please select the file to print" is generated. Then, the process proceeds to step S1607, where the voice synthesis data generated in step S1612 is notified to the smart speaker 102, and this process ends.
[0206] In step S1613, the CPU 402 checks whether the information acquisition job command generated in step S1102 is a FAX job list acquisition command. If it is a FAX job list acquisition command (YES in step S1613), the process transitions to step S1614; otherwise, it transitions to step S1624.
[0207] In step S1613a, the CPU 402 obtains the job list included in the job list information response transmitted from the MFP 101 by transmitting a FAX job list acquisition command to the MFP 101 (step S1103). Specifically, the data transmitted to the MFP 101 here is the data 1901 including the FAX job list acquisition command, attribute information, and device configuration information shown in FIG. 19.
[0208] In step S1614, the CPU 402 refers to the attribute information obtained from the smart speaker 102 and checks whether the smart speaker 102 is installed in the MFP 101. If it is not installed (NO in step S1614), the process transitions to step S1615; otherwise, it transitions to step S1623.
[0209] In step S1615, the CPU 402 refers to the device configuration information obtained from the smart speaker 102 and checks whether the LUI is installed in the smart speaker 102. If it is installed (YES in step S1615), the process transitions to step S1616; otherwise, it transitions to step S1619.
[0210] In step S1616, the CPU 402 acquires, from the MFP 101, a preview image of the received FAX data that is the print target of each FAX job in the job list acquired in step S1613a. Specifically, the CPU 402 transmits a preview image acquisition command to the MFP 101. As a result, a preview image response including the preview image of the received FAX data that is the print target of each FAX job is transmitted from the MFP 101.
[0211] In step S1617, the CPU 402 notifies the smart speaker 102 of the preview image acquired in step S1616 through the data transmission / reception unit 701. Then, the LUI of the smart speaker 102 displays the preview image notified from the cloud server 103 in step S1617.
[0212] In step S1618, the CPU 402 performs a voice synthesis process to generate voice information to be notified to the smart speaker 102, for example, voice synthesis data such as "The preview image of the received FAX has been displayed. Please select the FAX to be printed." Then, the process proceeds to step S1607, where the voice synthesis data is notified to the smart speaker 102 and this process ends. Then, the voice playback unit 705 of the smart speaker 102 plays back the voice synthesis data notified from the cloud server 103 in step S1607.
[0213] As a result, the user 107 can easily select the FAX data to be printed from the preview image displayed on the LUI in response to the voice notification from the smart speaker 102 such as "The preview image of the received FAX has been displayed. Please select the FAX to be printed."
[0214] In step S1619, the CPU 402 acquires information of the first job in the job list acquired in step S1613a (hereinafter referred to as "first job information").
[0215] In step S1620, the CPU 402 generates a FAX job command based on the head job information acquired in step S1619.
[0216] In step S1621, the CPU 402 notifies the MFP 101 of the FAX job command (job information) generated in step S1620.
[0217] In step S1622, the CPU 402 performs a voice synthesis process to generate voice synthesis data for notifying the user 107 that the FAX job command has been notified. For example, voice synthesis data consisting of a message such as "The fax has been printed" is generated. Then, the process proceeds to step S1607, and the voice synthesis data generated in step S1622 is notified to the smart speaker 102, and this process ends.
[0218] In step S1623, the CPU 402 performs a voice synthesis process to generate voice synthesis data to be notified to the smart speaker 102. For example, voice synthesis data consisting of a message such as "The preview image of the received fax has been displayed. Please select the fax to be printed" is generated. Then, the process proceeds to step S1607, and the voice synthesis data generated in step S1623 is notified to the smart speaker 102, and this process ends.
[0219] In step S1624, the CPU 402 performs a voice synthesis process to generate voice synthesis data to be notified to the smart speaker 102. For example, voice synthesis data consisting of a message such as "That command cannot be executed" is generated. Then, the process proceeds to step S1607, and the voice synthesis data generated in step S1624 is notified to the smart speaker 102, and this process ends.
[0220] <Flowchart of the process of step S1502 in FIG. 15 in the MFP> FIG. 17 is a flowchart showing details of the process of step S1503 in FIG. 15, that is, the process that is executed each time the data transmission / reception unit 601 of the device control program 600 receives job information in the MFP 101.
[0221] Here, in the process of FIG. 13, when it is confirmed in step S1311 that the smart speaker 102 is not attached to the MFP 101, the job list information response is notified to the cloud server 103 (step S1312), and this process ends. On the other hand, in the process of FIG. 17, after the process of step S1312, after executing the processes of steps S1701 and S1702 described below, this process ends.
[0222] Therefore, the description of the same process as that in FIG. 13 is omitted, and hereinafter, only the processes of steps S1701 to S1703 will be described.
[0223] In step S1701, the CPU 202 checks whether the smart speaker 102 is equipped with an LUI from the device configuration information (step S1501 in FIG. 15) transmitted together with the FAX job list acquisition command (job information) acquired in step S1301. If it is equipped (YES in step S1701), the process proceeds to step S1702, and if not, the process proceeds to step S1703.
[0224] In step S1702, when the CPU 202 receives the preview image acquisition command from the cloud server 103, it generates a preview image of the received FAX data that is the printing target in each FAX job held in the storage 205. Then, it notifies the cloud server 103 of a preview image response including the preview image.
[0225] <Flowchart of the process in the response processing state of the smart speaker equipped with an LUI> Figure 18 is a flowchart showing the details of the processing in the response processing state of the smart speaker 102 included in the LUI, that is, the processing when receiving response data such as voice synthesis data from the cloud server 103.
[0226] In step S1801, the CPU 302 checks whether the data received from the cloud server 103 is a job list. If it is a job list (YES in step S1801), the process proceeds to step S1802; otherwise, it proceeds to step S1804.
[0227] In step S1802, the CPU 302 stores the job list received from the cloud server 103 in the storage 305 and displays the job list on the LUI. Then, the process proceeds to step S1803, where the CPU 302 plays the voice synthesis data received from the cloud server 103 together with the job list through the speaker 310, executes the voice output process for voice notification to the user 107, and ends this process.
[0228] In step S1804, the CPU 302 checks whether the data received from the cloud server 103 is a preview image of the received FAX data that is the printing target of each FAX job in the job list queried by the user 107. If it is a preview image (YES in step S1804), the process proceeds to step S1805; otherwise, it proceeds to step S1803.
[0229] In step S1805, the CPU 302 stores the preview image received from the cloud server 103 in the storage 305 and displays the preview image on the LUI. Then, the process proceeds to step S1803, where the CPU 302 plays the voice synthesis data received from the cloud server 103 together with the preview image through the speaker 310, executes the voice output process for voice notification to the user 107, and ends this process.
[0230] As described above, the operation of the information processing system by voice control is changed according to whether the smart speaker 102 is attached to the MFP 101 or whether the LUI is attached to the smart speaker 102. This makes it possible to improve the usability of the entire information processing system.
[0231] Further, the present invention is not limited to the configurations described in the above-described first and second embodiments. For example, the present invention is also applicable to changing the operation of the information processing system according to the state of the smart speaker 102 and the commands included in the voice operations received by the smart speaker 102, as described in FIG. 22.
[0232] Note that the "vicinity" in FIG. 22 refers to the fact that, in step S1006 of FIG. 10, the attribute information indicating that the smart speaker 102 is attached to the MFP 101 was transmitted to the cloud server 103. Also, the "voice input only" in FIG. 22 refers to the fact that, in the same step of FIG. 10, the attribute information indicating that the smart speaker 102 is not attached to the MFP 101 was transmitted to the cloud server 103.
[0233] The "with LUI" in FIG. 22 refers to the fact that, in steps S1006 and S1501 of FIG. 15, the attribute information indicating that the smart speaker 102 is not attached to the MFP 101 and the device configuration information indicating that the smart speaker 102 has the LUI were transmitted.
[0234] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.
[0235] (Other Embodiments) The present invention can also be realized by supplying a program that implements one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and having one or more processors in a computer of the system or apparatus read and execute the program. It can also be executed by a circuit (e.g., ASIC) that implements one or more functions.
Explanation of Signs
[0236] 101 MFP 102 Smart Speaker 103 Cloud Server 104 Network 105 Gateway 106 Print Server 202, 302, 402 CPU 206, 306, 406 Network I / F 209 Operation Panel 308 Microphone 310 Speaker 600 Device Control Program 700 Voice Control Program 800 Voice Data Conversion Control Program 900 Print Data Control Program
Claims
1. A processing system having an image processing apparatus and an information processing apparatus and being communicable with an audio control apparatus, wherein the information processing apparatus has a receiving means for receiving first audio information generated based on attribute information set in the audio control apparatus and audio received from a user from the audio control apparatus; a generating means for generating a command based on the first audio information; and an output means for outputting response information to the audio control apparatus as a response to the command, and when the attribute information is first attribute information, the output means outputs first type response information as a response to the command; when the attribute information is second attribute information different from the first attribute information, the output means outputs second type response information different from the first type response information as a response to the command, the attribute information is information indicating whether the audio control apparatus is attached to the image processing apparatus, the first attribute information is information indicating that the audio control apparatus is attached to the image processing apparatus, and the second attribute information is information indicating that the audio control apparatus is not attached to the image processing apparatus, characterized by the processing system.
2. The output of the second type response information is a notification to the user using audio by the audio control apparatus without using display by the image processing apparatus, according to the processing system of claim 1.
3. The command is a command for executing a job using at least one of a FAX function, a printing function, and a scanning function, according to the processing system of claim 1 or 2.
4. The output of the response information of the first type is a notification to the user by voice by the voice control device and display by the image processing device, according to any one of claims 1 to 3, wherein the processing system is characterized in that.
Citation Information
Patent Citations
Voice operation system, voice operation method, and voice operation program
JP2020087347A
Image processing system, image forming device, voice input prohibition determination method and program
JP2020098229A
System, image formation device, method, and program
JP2020134903A
Controller, image formation system, and program
JP2020149602A
Image processing system, and voice response processing method and program
JP2021052220A