Device management apparatus, device management system, device management method and program
The device management system addresses the issue of multiple devices responding to voice commands by comparing voice information and device status to select the correct device for execution, improving operational accuracy and efficiency.
Patent Information
- Application Number
- JP2021150799
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-09-16
AI Technical Summary
When multiple voice-operable devices are installed in close proximity, there is a risk that a device other than the intended one may erroneously respond to voice commands, leading to operational inefficiencies.
A device management system that manages multiple voice-operable devices by comparing voice information across devices to identify the appropriate device for a specific operation instruction, considering factors like time of acquisition, voice characteristics, job content, device status, and user authentication to determine the correct device for execution.
Ensures accurate selection of the appropriate device for voice operations, reducing errors and enhancing operational efficiency by minimizing misinterpretation of voice commands among nearby devices.
Smart Images

Figure 0007735747000001 
Figure 0007735747000002 
Figure 0007735747000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a device management apparatus, a device management system, a device management method, and a program. [Background technology]
[0002] Conventionally, in addition to manual operation instructions from an operation unit, operation instructions are input by voice to devices such as image forming apparatuses. A voice-operable device outputs a voice response to the voice operation instructions and executes processing according to the voice operation instructions.
[0003] When operating a device such as an image forming apparatus using voice commands, the user and the device basically have a one-to-one relationship, but there are cases where the device picks up the voice of a person other than the actual operator. Therefore, a technology has been proposed to improve the operability of voice operations for image forming devices used in an environment where multiple people speak at the same time (see Patent Document 1). The image forming device described in Patent Document 1 stores in a storage unit settings based on job setting instructions received by voice in association with the identification information of the user who gave the setting instructions, and when an operation instruction for a job is received by voice, it extracts from the storage unit settings associated with the identification information of the same user as the user who gave the operation instruction, and executes the job based on the extracted settings. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-127104 Summary of the Invention [Problem to be solved by the invention]
[0005] However, when a plurality of voice-operable devices are installed in close proximity to one another, there is a possibility that a device other than the device to be operated may erroneously respond when the device is operated by voice.
[0006] The present invention has been made in consideration of the above-mentioned problems in the conventional technology, and an object of the present invention is to select an appropriate device from among a plurality of voice-operable devices as the operation target. [Means for solving the problem]
[0007] In order to solve the above problem, the invention described in claim 1 is a device management device that manages multiple devices that can be voice-operated by accepting voice input from a user regarding the content of a job and acquiring voice information, and is equipped with an extraction means that compares the voice information acquired from each of two or more devices among the multiple devices and extracts, from the two or more devices, a device that has acquired voice information related to the same operation instruction as a candidate for voice operation, a determination means that determines which of the candidate devices for voice operation should be used to execute the content of the job included in the same operation instruction, and an execution means that causes the device determined by the determination means to execute the content of the job included in the same operation instruction.
[0008] The invention described in claim 2 is a device management device described in claim 1, wherein the extraction means determines that the voice information relating to the same operation instruction has been acquired when it is determined that the time at which the voice information was acquired is the same between the two or more devices and that the content of the job indicated by the voice information is the same between the devices.
[0009] The invention described in claim 3 is a device management device described in claim 2, in which the extraction means determines whether the content of the jobs is the same by comparing the voice characteristics of the voice information acquired from each of the two or more devices.
[0010] The invention described in claim 4 is a device management device described in claim 2, wherein the extraction means determines whether the content of the jobs is the same by comparing character information recognized by voice from the voice information acquired from each of the two or more devices.
[0011] The invention described in claim 5 is a device management device described in claim 2, wherein the extraction means determines whether the job contents are the same by comparing the determination results determined by a job content determination algorithm based on character information obtained by voice recognition from voice information acquired from each of the two or more devices.
[0012] The invention described in claim 6 is a device management device described in any one of claims 1 to 5, wherein the determination means selects, as the cancellation target, a device among the candidates for voice operation other than the device determined to have the fastest response based on the sleep state or sleep level of the candidates for voice operation.
[0013] The invention described in claim 7 is a device management device described in any one of claims 1 to 6, wherein the determination means cancels devices among the candidate voice operation targets in which the power supply to the mechanism that acquires voice information and the power supply to the voice operation target mechanism are independent.
[0014] The invention described in claim 8 is a device management device described in any one of claims 1 to 7, wherein the determination means determines, as the operation target, from among the candidate voice operation targets, a device that is being manually operated by the same user as the user corresponding to the voice information related to the same operation instruction.
[0015] The invention described in claim 9 is a device management device described in any one of claims 1 to 8, wherein the determination means determines, among the candidates for voice operation targets, a device that is being manually operated by a user other than the user corresponding to the voice information related to the same operation instruction as the cancellation target.
[0016] The invention described in claim 10 is a device management device described in any one of claims 1 to 9, wherein the determination means determines, among the candidates for voice operation targets, a device that is being voice operated by a user other than the user corresponding to the voice information related to the same operation instruction as a target for cancellation.
[0017] The invention described in claim 11 is a device management device described in any one of claims 1 to 10, wherein the determination means compares the setting items that can be set for each of the candidates for voice operation with the voice instruction content related to the same operation instruction, and targets devices other than devices that can execute the voice instruction content for cancellation.
[0018] The invention described in claim 12 is a device management device described in any one of claims 1 to 11, wherein the determination means compares the installation status of optional items attached to each of the candidate voice operation targets with the voice instruction content related to the same operation instruction, and targets devices other than devices that can execute the voice instruction content for cancellation.
[0019] The invention described in claim 13 is a device management device described in any one of claims 1 to 12, wherein the determination means compares the software installed in each of the candidates for voice operation with the voice instruction content related to the same operation instruction, and targets devices other than devices that can execute the voice instruction content as cancellation targets.
[0020] The invention described in claim 14 is a device management device described in any one of claims 1 to 13, wherein the determination means compares the input sound volume of the voice information acquired for each of the voice operation target candidates, and selects the device with the loudest input sound as the operation target.
[0021] The invention described in claim 15 is a device management system comprising a plurality of voice-operable devices and a device management device that manages the plurality of devices, wherein each of the plurality of devices comprises a voice acquisition means that accepts voice input from a user regarding the contents of a job and acquires voice information, and the device management device comprises: an extraction means that compares the voice information acquired by each of two or more of the plurality of devices and extracts, from the two or more devices, a device that has acquired voice information related to the same operation instruction as a candidate for voice operation; a determination means that determines which of the candidate devices for voice operation should be made to execute the contents of the job included in the same operation instruction; and an execution means that makes the device determined by the determination means execute the contents of the job included in the same operation instruction.
[0022] The invention described in claim 16 is a device management method for managing a plurality of voice-operable devices, comprising: a voice acquisition step for accepting voice input from a user regarding the contents of a job at two or more of the plurality of devices and acquiring voice information; an extraction step for comparing the voice information acquired at each of the two or more devices and extracting, as a candidate for voice operation, a device from the two or more devices that has acquired voice information related to the same operation instruction; a determination step for determining which of the candidate voice operation devices should be made to execute the contents of the job included in the same operation instruction; and an execution step for making the device determined in the determination step execute the contents of the job included in the same operation instruction.
[0023] The invention described in claim 17 is a program for causing a computer of a device management apparatus that manages multiple voice-operable devices by accepting voice input from a user regarding the contents of a job and acquiring voice information to function as an extraction means that compares the voice information acquired from each of two or more of the multiple devices and extracts, from the two or more devices, a device that has acquired voice information related to the same operation instruction as a candidate for voice operation; a determination means that determines which of the candidate devices for voice operation should be used to execute the contents of the job included in the same operation instruction; and an execution means that causes the device determined by the determination means to execute the contents of the job included in the same operation instruction. [Effects of the Invention]
[0024] According to the present invention, it is possible to select an appropriate device from among a plurality of voice-operable devices as the operation target. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a diagram illustrating a system configuration of a device management system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing a functional configuration of the image forming apparatus. [Figure 3] FIG. 2 is a block diagram showing a functional configuration of a device management apparatus. [Figure 4] FIG. 10 is a diagram illustrating an example of the data configuration of a user authentication table. [Figure 5] FIG. 10 is a diagram illustrating an example of the data configuration of a determination condition setting table. [Figure 6] 10 is a ladder chart showing a first device management process executed in the device management system. [Figure 7] 10 is a ladder chart showing a first device management process executed in the device management system. [Figure 8] 10 is a flowchart illustrating an execution device determination process executed by the device management apparatus. [Figure 9] 10 is a ladder chart showing a second device management process executed in a device management system according to a second embodiment. [Figure 10] 10 is a ladder chart showing a second device management process executed in a device management system according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0026] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of a device management system according to the present invention will be described with reference to the drawings, but the scope of the invention is not limited to the illustrated examples.
[0027] [First embodiment] <Device management system configuration> FIG. 1 shows a system configuration of a device management system 100 according to a first embodiment of the present invention. 1, the device management system 100 includes a plurality of devices, namely, image forming apparatuses 10A, 10B, 10C, etc., smart speakers 20A, 20B, 20C, etc., and a device management apparatus 30. The image forming apparatuses 10A, 10B, 10C, etc. and the device management apparatus 30, and the smart speakers 20A, 20B, 20C, etc. and the device management apparatus 30 are connected via a communication network N so that data communication is possible.
[0028] The image forming apparatuses 10A, 10B, 10C, etc. are MFPs (Multi-Functional Peripherals) having a printer function, a copy function, a scanner function, etc. When there is no need to distinguish between the image forming apparatuses 10A, 10B, 10C, etc., they may be referred to as the image forming apparatus 10.
[0029] The smart speakers 20A, 20B, 20C, etc. are equipped with a microphone, a speaker, a communication unit, etc. The smart speakers 20A, 20B, 20C, etc. accept voice input via the microphone, convert the input voice into voice information (voice data), and transmit the voice information to the device management device 30 via the communication unit. The smart speakers 20A, 20B, 20C, etc. output voice from their speakers based on voice information for voice response received from the device management device 30. Note that when there is no need to particularly distinguish between the smart speakers 20A, 20B, 20C, etc., they may be referred to as smart speakers 20.
[0030] The smart speaker 20A receives an operation instruction for the image forming apparatus 10A and outputs a voice response to the operator of the image forming apparatus 10A. That is, the smart speaker 20A is associated with the image forming apparatus 10A. Similarly, the smart speakers 20B, 20C, etc. are associated with the image forming apparatuses 10B, 10C, etc., respectively.
[0031] Each of the multiple devices (image forming apparatuses 10A, 10B, 10C, etc.) is equipped with smart speakers 20A, 20B, 20C, etc. as voice acquisition means that accepts voice input from the user regarding the content of the job and acquires voice information. Note that "each of the multiple devices is equipped with voice acquisition means" includes the existence of voice acquisition means (smart speaker 20) associated with each of the multiple devices. In this embodiment, the mechanism for acquiring voice information (smart speaker 20) and the mechanism to be operated by voice (image forming device 10) are configured to have independent power supplies.
[0032] The device management apparatus 30 manages image forming apparatuses 10A, 10B, 10C, etc. as a plurality of voice-operable devices. The device management device 30 performs voice recognition processing on the voice information acquired from the smart speakers 20A, 20B, 20C, and so on to acquire the voice instruction content. Based on the voice instruction content, the device management device 30 transmits operation instructions to the image forming devices 10A, 10B, 10C, and so on, and transmits voice information to be responded to by the smart speakers 20A, 20B, 20C, and so on.
[0033] <Configuration of image forming device> FIG. 2 shows the functional configuration of the image forming apparatus 10. As shown in FIG. 2, the image forming apparatus 10 includes a control unit 11, a document reading unit 12, an image forming unit 13, a user authentication unit 14, an operation panel 15, a communication unit 16, and a storage unit 17.
[0034] The control unit 11 is configured with a CPU (Central Processing Unit) 111, a RAM (Random Access Memory) 112, a ROM (Read Only Memory) 113, etc. The CPU 111 reads various processing programs stored in the ROM 113, loads them into the RAM 112, and controls the operations of each unit of the image forming apparatus 10 in accordance with the loaded programs.
[0035] The document reading unit 12 optically scans a document transported from an ADF (Auto Document Feeder) onto a contact glass or a document placed on the contact glass, forms an image of the reflected light of light emitted from a light source and scanned onto the document on the light receiving surface of a CCD (Charge Coupled Device) sensor, reads the document image, A / D converts the read image, and outputs the obtained image data to the control unit 11.
[0036] Image forming unit 13 forms an image on paper based on image data generated by document reading unit 12, image data received by communication unit 16, etc. For example, image forming unit 13 performs electrophotographic image formation and is composed of a photosensitive drum, a charging unit that charges the photosensitive drum, an exposure unit that exposes and scans the surface of the photosensitive drum based on image data, a developing unit that attaches toner to the photosensitive drum, a transfer unit that transfers the toner image formed on the photosensitive drum to paper, and a fixing unit that fixes the toner image formed on paper, etc.
[0037] The user authentication unit 14 authenticates the user by comparing information such as a user ID and password entered by the user through the operation unit 152 with pre-registered reference data (such as a combination of a user ID and a password). The user authentication unit 14 may also read the user ID from an IC card carried by the user to identify the user. The user authentication unit 14 may also obtain the user's biometric information (fingerprint, iris, facial recognition data, etc.) and authenticate the user by comparing it with pre-registered reference data.
[0038] The operation panel 15 includes a display unit 151 that displays various information to the user, and an operation unit 152 that accepts operation inputs from the user. The display unit 151 is configured, for example, as a color liquid crystal display. The operation unit 152 is configured, for example, as a touch panel that is provided on the screen of the display unit 151 and that inputs information by touch operation, and push button keys that are arranged around the screen of the display unit 151.
[0039] The communication unit 16 is an interface that connects the image forming apparatus 10 to the communication network N. The communication unit 16 transmits and receives data to and from external devices such as the device management apparatus 30.
[0040] The storage unit 17 is a non-volatile storage device configured by a hard disk drive (HDD), a solid state drive (SSD), or the like.
[0041] The storage unit 17 stores the status of the device (image forming device 10), the user ID of the currently logged-in user, etc. Examples of the device status include the sleep state of the device, the sleep level, configurable items, the installation status of optional items (such as a finisher), the name of software installed in the device, and the usage status of consumables (such as the amount of toner remaining, the number of sheets of paper remaining in the paper feed tray).
[0042] The sleep state is information indicating whether the image forming apparatus 10 is in sleep mode or in operation. The sleep level is information indicating the degree (depth) of sleep when the sleep state is "sleeping", and is indicated by a number of stages (1 to 3, etc.).
[0043] The configurable items are setting items that can be set in the image forming apparatus 10, and are determined by usage permission / prohibition settings for the image forming apparatus 10. Examples of usage permission / prohibition settings include a setting to permit or prohibit sending of scanned data for each type of destination (e-mail, HDD, USB memory, manual setting, etc.), and a setting to permit or prohibit color / monochrome jobs for each user and for each copy operation, scan operation, and fax operation.
[0044] <Configuration of device management device> FIG. 3 shows the functional configuration of the device management apparatus 30. As shown in FIG. 3, the device management apparatus 30 includes a control unit 31, a clock unit 32, a communication unit 33, and a storage unit .
[0045] The control unit 31 is composed of a CPU 311, a RAM 312, a ROM 313, etc. The CPU 311 reads out various processing programs stored in the ROM 313, loads them into the RAM 312, and controls the operation of each unit of the device management apparatus 30 in accordance with the loaded programs.
[0046] The clock unit 32 has a clock circuit (RTC: Real Time Clock), and measures the current time using this clock circuit and outputs it to the control unit 31.
[0047] The communication unit 33 is an interface that connects the device management apparatus 30 to the communication network N. The communication unit 33 transmits and receives data to and from external devices such as the image forming apparatuses 10A, 10B, 10C, etc. and the smart speakers 20A, 20B, 20C, etc.
[0048] The storage unit 34 is a non-volatile storage device configured by an HDD, an SSD, or the like. The storage unit 34 stores a user authentication table 341, a judgment condition setting table 342, a device management table 343, an audio information management table 344, and the like.
[0049] The user authentication table 341 is a table for managing authentication data for users of the device management system 100. Fig. 4 shows an example of the data configuration of the user authentication table 341. In the user authentication table 341, a user ID and voice feature data are associated with each user. As the voice feature data, for example, a hash value obtained by processing voice information of each user's voice is used.
[0050] The judgment condition setting table 342 is a table that defines the judgment conditions for determining whether or not each device (image forming apparatus 10) is to be operated. FIG. 5 shows an example of the data configuration of the judgment condition setting table 342. In the judgment condition setting table 342, judgment priorities and judgment criteria are associated with each other. The judgment priorities are priorities defined for the judgment criteria. The judgment criteria are used to determine whether each device is to be operated or canceled. Here, the smaller the numerical value of the judgment priority associated with the judgment criterion, the higher the priority is.
[0051] The device management table 343 is a table for managing each device (image forming apparatus 10). For each device (image forming apparatus 10), the device management table 343 associates the device's (image forming apparatus 10's) identification information, device status, the user ID of the logged-in user, and identification information of the smart speaker 20 (e.g., smart speaker 20A if the device is image forming apparatus 10A). The device status includes the device's sleep state, sleep level, power status, configurable items, attachment status of optional items, name of software installed in the device, usage status of consumable items, etc. The power status is information indicating whether the image forming apparatus 10 is powered on or off.
[0052] When the image forming device 10 notifies the device management device 30 of the device's sleep state, sleep level, configurable items, installation status of optional items, name of software installed in the device, usage status of consumable items, etc., the CPU 311 of the device management device 30 updates the contents of the corresponding record in the device management table 343. Regarding the power state, when the image forming apparatus 10 notifies the device management apparatus 30 of the device state, the CPU 311 of the device management apparatus 30 sets the "power state" of the corresponding record in the device management table 343 to "on." When the image forming apparatus 10 does not periodically notify the device management apparatus 30 of the device state, the CPU 311 of the device management apparatus 30 sets the "power state" of the corresponding record in the device management table 343 to "off."
[0053] The voice information management table 344 is a table for managing voice information acquired by each device. The voice information management table 344 associates voice information with the identification information of the smart speaker 20 (the smart speaker 20 that transmitted the voice information) and the reception time.
[0054] Voice input from the user regarding the content of the job is accepted and voice information is obtained in two or more devices among the multiple devices (image forming apparatuses 10A, 10B, 10C, etc.) in the device management system 100. Specifically, voice input from the user regarding the content of the job is accepted and voice information is obtained in two or more smart speakers 20 among the smart speakers 20A, 20B, 20C, etc. corresponding to each of the multiple devices.
[0055] The CPU 311 (extraction means) compares the voice information acquired from each of two or more devices among the multiple devices, and extracts the device among the two or more devices that acquired voice information related to the same operation instruction (the image forming device 10 corresponding to the smart speaker 20 that captured the same voice input) as a candidate for voice operation.
[0056] Specifically, if the CPU 311 determines that the time at which the audio information was acquired between two or more devices is the same and that the content of the job indicated by the audio information between the devices is the same, the CPU 311 determines that the audio information related to the same operation instruction has been acquired.
[0057] The CPU 311 compares the voice characteristics of the voice information acquired from each of the two or more devices to determine whether the job contents are the same. Voice characteristics may include voiceprints and numerical values of frequency. The CPU 311 compares parameters indicating the voice characteristics obtained from each piece of voice information received from the smart speakers 20 corresponding to the two or more devices, and if it determines that the information belongs to the same user, it determines that the job contents are the same.
[0058] If the smart speaker 20 has a function for analyzing voice information and identifying the user from the voice characteristics, the user identification result may be transmitted to the device management device 30 in addition to the voice information. In this case, examples of the user identification result transmitted from the smart speaker 20 to the device management device 30 include a user name, a user ID, a hash value (a value obtained by processing the voice information), anonymized characteristic information such as a male in his twenties, etc.
[0059] Furthermore, the CPU 311 determines whether the job contents are the same by comparing the character information obtained by voice recognition from the voice information acquired from each of the two or more devices. The CPU 311 performs voice recognition processing on each of the voice information received from the smart speakers 20 corresponding to the two or more devices to convert it into text (character information), and determines the degree of match between the character information. The CPU 311 determines that the job contents are the same if the degree of match between the text in the converted text state is equal to or greater than a predetermined value (e.g., 80% or greater).
[0060] The CPU 311 may also determine whether the job contents are the same by comparing the results of a job content determination algorithm based on character information obtained by voice recognition from voice information acquired from each of two or more devices. The job content determination algorithm is a processing procedure for analyzing character information and determining the job contents. The CPU 311 performs voice recognition processing on each piece of voice information received from the smart speaker 20 corresponding to two or more devices to convert it into text (character information), and then converts the text information into a job command using the job content determination algorithm. The job command expresses the job contents, broken down into processing operations, means, file formats, transmission destinations, output destinations, etc. The CPU 311 then determines whether the job contents are the same by comparing the job commands corresponding to each device. For example, the CPU 311 analyzes text information such as "Send it to A as a PDF file by email" using a job content determination algorithm, and obtains job commands such as "Scan," "Send by email," "Destination: A," and "PDF file." The CPU 311 compares the job commands obtained by the analysis, and if the commands match, determines that the job contents are the same.
[0061] The method for determining whether the job contents indicated by the voice information are the same between devices may be a combination of the above methods as appropriate.
[0062] The CPU 311 (determination means) determines which of the voice operation target candidates the device is to execute the job content included in the same operation instruction.
[0063] Specifically, the CPU 311 selects, as the cancellation target, devices other than the device determined to have the fastest response based on the sleep state or sleep level of the voice operation target candidate. The "device determined to have the fastest response" does not have to be one, but may be multiple devices determined to have approximately the same response time. The CPU 311 references the device management table 343 in the storage unit 34 and acquires the sleep state or sleep level from the device state associated with the identification information of each device (image forming apparatus 10) as the voice operation target candidate. For example, when there are a device whose sleep state is "Starting" and a device whose sleep state is "Sleeping" as the voice operation target candidate, the CPU 311 selects the "Sleeping" device as the cancellation target. Furthermore, when there are multiple "Sleeping" devices, the CPU 311 selects the device whose sleep level indicates the deeper sleep depth as the cancellation target.
[0064] For devices among the voice operation target candidates in which the power supply to the mechanism that acquires voice information (smart speaker 20) and the voice operation target mechanism (image forming device 10) is independent, CPU 311 cancels devices in which the power to the voice operation target mechanism is turned off. CPU 311 refers to device management table 343 and acquires the power status from the device status associated with the identification information of each device (image forming device 10) that is a voice operation target candidate.
[0065] The CPU 311 compares the configurable setting items for each of the voice operation target candidates with the voice instruction content for the same operation instruction, and selects devices other than those that can execute the voice instruction content as cancellation targets. The CPU 311 references the device management table 343 and acquires the configurable items from the device status associated with the identification information of each device (image forming apparatus 10) that is a voice operation target candidate. The CPU 311 then selects devices that cannot execute the voice instruction content as cancellation targets based on the configurable items. For example, when a "color job" is set in the voice instruction content, the CPU 311 selects a monochrome-only device as cancellation target.
[0066] The CPU 311 compares the installation status of optional items attached to each of the voice operation target candidates with the voice instruction content related to the same operation instruction, and selects devices other than those that can execute the voice instruction content as cancellation targets. The CPU 311 references the device management table 343 and acquires the installation status of optional items from the device status associated with the identification information of each device (image forming apparatus 10) that is a voice operation target candidate. The CPU 311 then selects devices that cannot execute the voice instruction content as cancellation targets based on the installation status of optional items. For example, if the voice instruction content includes "staple," the CPU 311 selects devices that do not have a stapler attached as cancellation targets.
[0067] CPU 311 compares the software installed in each of the voice operation target candidates with the voice instruction content related to the same operation instruction, and selects devices other than those that can execute the voice instruction content as cancellation targets. CPU 311 references device management table 343 and acquires the name of software installed in the device from the device status associated with the identification information of each voice operation target candidate device (image forming apparatus 10). CPU 311 then selects devices that cannot execute the voice instruction content as cancellation targets based on the software name. For example, in response to the voice instruction content "Display the weather forecast," CPU 311 selects devices that do not have application software that displays the weather forecast installed as cancellation targets.
[0068] The CPU 311 compares the remaining amount of consumables for each of the candidate voice operation targets with the voice instruction content related to the same operation instruction, and selects devices other than the device determined to be more appropriate for consuming the consumables in accordance with the voice instruction content as the cancellation target. The CPU 311 references the device management table 343 and acquires the usage status of the consumables from the device status associated with the identification information of each candidate voice operation target device (image forming apparatus 10). For example, the CPU 311 may prioritize devices with larger remaining amounts of consumables predicted to be consumed in accordance with the voice instruction content, or may prioritize devices with smaller remaining amounts so that the consumables are used up more quickly.
[0069] The CPU 311 compares the volume of the input sound of the voice information acquired for each of the voice operation target candidates, and selects the device with the loudest input sound as the operation target.
[0070] The method for determining which of the voice operation target candidates devices should execute the job content included in the same operation instruction may be a combination of the above methods as appropriate.
[0071] For example, the CPU 311 narrows down the devices to be operated from among the voice operation target candidates using the determination condition setting table 342 (see FIG. 5) stored in the storage unit 34. First, when "configurable items" is set as the first judgment criterion, the configurable items of the candidate voice operation target are compared with the voice instruction content related to the same operation instruction, and devices that cannot execute the voice instruction content are targeted for cancellation.
[0072] Next, when "sleep state" is set as the second criterion, devices other than the device determined to have the fastest response are set as cancellation targets based on the sleep state or sleep level of the voice operation target candidates excluding the cancellation targets determined according to the first criterion. Here, devices that are powered off may also be set as cancellation targets.
[0073] Next, when "the installation status of optional items" is set as the third judgment criterion, the installation status of optional items of the candidates for voice operation, excluding those to be canceled as determined according to the first and second judgment criteria, is compared with the voice instruction content relating to the same operation instruction, and devices that cannot execute the voice instruction content are set as candidates for cancellation.
[0074] Next, when "software" is set as the fourth criterion, the software of the candidates for voice operation, excluding the cancellation targets determined according to the first to third criteria, is compared with the voice instruction content relating to the same operation instruction, and devices that cannot execute the voice instruction content are set as cancellation targets.
[0075] Next, when "voice volume" is set as the fifth criterion, the volume of the input sound of the voice information acquired for each of the voice operation target candidates, excluding the cancellation target determined according to the first to fourth criterion, is compared, and the device with the loudest input sound is set as the operation target. Note that it is also possible to leave as candidates the devices whose input sound volume is equal to or greater than a predetermined value, and set the other devices as cancellation targets.
[0076] CPU 311 sequentially performs judgment based on the judgment criteria in accordance with the priority set in judgment condition setting table 342, and if the operation target can be determined before all judgments are completed (if the device can be narrowed down to one), it ends the operation target narrowing down process, continues processing on the operation target device, and determines the cancellation target. On the other hand, if the operation target cannot be determined even after performing the narrowing down process based on all judgment criteria, CPU 311 determines one arbitrary device as the operation target from among the devices that were not determined as cancellation targets among the voice operation target candidates, and determines all others as cancellation targets.
[0077] The CPU 311 (execution means) causes the determined device to execute the job content included in the same operation instruction.
[0078] The CPU 311 performs a voice recognition process on the voice information received from the smart speakers 20A, 20B, 20C, . . . and converts the voice information into a character string (text). The CPU 311 extracts the voice characteristics of the voice information received from the smart speakers 20A, 20B, 20C, . . . and identifies the user who uttered the voice based on the extracted voice characteristics (speaker recognition).
[0079] <Device management system operation> Next, the operation of the device management system 100 will be described. 6 and 7 are ladder charts showing a first device management process executed in the device management system 100. The first device management process is a process for determining the device (image forming device 10) to be operated when voice information relating to the same operation instruction is acquired in duplicate in multiple image forming devices 10 (smart speakers 20). Here, an example will be described in which voice information relating to the same operation instruction is acquired in image forming devices 10A and 10B.
[0080] In the image forming apparatus 10A, the CPU 111 cooperates with a resident application program to periodically transmit the status of the device (image forming apparatus 10A) to the device management apparatus 30 via the communication unit 16 (step S1). For example, the device status information transmitted includes the sleep state (on / asleep) of the image forming device 10A, the sleep level, configurable items, the installation status of optional items, the name of software installed on the image forming device 10A, and the usage status of consumable items.
[0081] In the device management apparatus 30, the CPU 311 acquires the status of the device (image forming apparatus 10A) via the communication unit 33 and stores the acquired device status in the storage unit 34 (step S2). Specifically, the CPU 311 stores the status (device status) of the image forming apparatus 10A in the device management table 343 of the storage unit 34 in association with the identification information of the image forming apparatus 10A.
[0082] Similarly, in the image forming apparatus 10B, the CPU 111 cooperates with a resident application program to periodically transmit the status of the device (image forming apparatus 10B) to the device management apparatus 30 via the communication unit 16 (step S3).
[0083] In the device management apparatus 30, the CPU 311 acquires the status of the device (image forming apparatus 10B) via the communication unit 33 and stores the acquired device status in the storage unit 34 (step S4). Specifically, the CPU 311 stores the status (device status) of the image forming apparatus 10B in the device management table 343 of the storage unit 34 in association with the identification information of the image forming apparatus 10B.
[0084] The processes of steps S1 to S4 are repeated at the timing when each device transmits its own status.
[0085] Here, the user performs voice input near the image forming apparatus 10A and the image forming apparatus 10B (step S5). For example, the user utters a voice saying "Make a copy" in a location where the voice can reach the smart speakers 20A and 20B.
[0086] The smart speaker 20A corresponding to the image forming device 10A and the smart speaker 20B corresponding to the image forming device 10B each receive voice input via a microphone, convert the input voice into voice information, and transmit the voice information to the device management device 30 via a communication unit (steps S6 and S7). The smart speakers 20A and 20B add their own identification information to the voice information and transmit it to the device management device 30.
[0087] In the device management device 30, when the CPU 311 receives the voice information acquired by the smart speaker 20A via the communication unit 33, it stores the received voice information in the storage unit 34 (step S8). Specifically, in response to receiving the voice information and the identification information of the smart speaker 20A, the CPU 311 acquires the current time (reception time) from the clock unit 32, and stores the received voice information in the voice information management table 344 of the storage unit 34 in association with the identification information of the smart speaker 20A and the reception time.
[0088] Similarly, when the CPU 311 receives voice information acquired by the smart speaker 20B via the communication unit 33, it stores the received voice information in the storage unit 34 (step S9). Specifically, in response to receiving the voice information and the identification information of the smart speaker 20B, the CPU 311 acquires the current time (reception time) from the clock unit 32, and stores the received voice information in the voice information management table 344 of the storage unit 34 in association with the identification information and the reception time of the smart speaker 20B.
[0089] The CPU 311 of the device management apparatus 30 compares the voice information acquired by each of the smart speakers 20A and 20B, and extracts the device that acquired the voice information related to the same operation instruction as a candidate for voice operation (step S10). Specifically, when the CPU 311 determines that the time when the voice information was acquired between the devices (smart speakers 20A and 20B) is the same for two or more devices and the content of the job indicated by the voice information between the devices is the same, it determines that the voice information related to the same operation instruction has been acquired.
[0090] In this embodiment, the time when the device management device 30 receives the audio information is considered to be approximately the same as the "time when the audio information was acquired on the device," and the reception time associated with the audio information in the audio information management table 344 is used as the "time when the audio information was acquired on the device."
[0091] In addition, for each smart speaker 20 that has acquired voice information related to the same operation instruction, the CPU 311 acquires from the device management table 343 the identification information of each image forming device 10 that corresponds to the identification information of each smart speaker 20, and sets each image forming device 10 that corresponds to the acquired identification information as a candidate for voice operation.
[0092] In addition, when determining whether the job content indicated by the voice information between devices is the same, for example, CPU 311 obtains voice characteristics (voiceprint, numerical value of frequency, etc.) from each of the voice information acquired by smart speakers 20A and 20B, and compares the voice characteristics to determine whether the job content is the same. In addition, CPU 311 may perform voice recognition processing on each of the voice information acquired by smart speakers 20A and 20B, and determine whether the content of the jobs is the same by comparing the text information obtained by the voice recognition processing. In addition, CPU 311 may perform voice recognition processing on each of the voice information acquired by smart speakers 20A and 20B, convert the text information obtained by the voice recognition processing into a job command using a job content determination algorithm, and compare the job commands to determine whether the job content is the same.
[0093] Next, the CPU 311 of the device management apparatus 30 performs an execution device determination process (step S11). The execution device determination process is a process for determining which of the voice operation target candidates should be used to execute the contents of the job included in the same operation instruction. The execution device determination process will be described in detail later. Here, it is assumed that the image forming apparatus 10A is determined as the device to be operated, and the image forming apparatus 10B is determined as the device to be cancelled.
[0094] 7, the CPU 311 of the device management apparatus 30 transmits a job execution instruction to the device to be operated (image forming apparatus 10A) via the communication unit 33 (step S12). The CPU 311 of the device management apparatus 30 also transmits voice information related to the voice response to the smart speaker 20A corresponding to the image forming apparatus 10A via the communication unit 33. In the image forming device 10A, the CPU 111 executes the job (screen transition, copy operation, etc.) based on the job execution instruction (step S13). In addition, in the smart speaker 20A, a voice response based on the voice information is output from the speaker. For example, in response to the user's voice saying "Copy," the smart speaker 20A emits a voice saying "The copy operation will start. Do you want to change the screen settings?"
[0095] Furthermore, the CPU 311 of the device management apparatus 30 transmits a voice operation cancellation instruction to the device to be cancelled (image forming apparatus 10B) via the communication unit 33 (step S14). In image forming apparatus 10B, CPU 111 notifies the user of the cancellation of the voice operation based on the voice operation cancellation instruction (step S15). For example, CPU 111 of image forming apparatus 10B may display a message on display unit 151 indicating that the operation is not the target, or may emit a sound indicating the cancellation, such as a beep. Note that as long as only one device (image forming apparatus 10A) is able to respond to continue the operation, there is no problem with operability even if the other devices do not clearly respond, and therefore image forming apparatus 10B may not respond to the user. This completes the first device management process.
[0096] 8 is a flowchart showing the execution device determination process (step S11 of the first device management process) executed by the device management apparatus 30. The execution device determination process is a process for narrowing down the devices to be operated by excluding devices determined as cancellation targets from the voice operation target candidates.
[0097] First, the CPU 311 of the device management apparatus 30 sets i to 1 (step S21). Then, the CPU 311 refers to the judgment condition setting table 342 (see FIG. 5) in the storage unit 34, performs judgment on the voice operation target candidate using the i-th criterion (step S22), and determines the device with the lowest priority as the device to be canceled (step S23).
[0098] Next, the CPU 311 determines whether or not the number of candidates for voice operation has been narrowed down to one as a result of excluding the device to be canceled (step S24). If there are still two or more candidates for voice operation even after excluding the device to be canceled (step S24; NO), the CPU 311 determines whether or not the judgment has been completed for all criteria set in the judgment condition setting table 342 (step S25).
[0099] If there are any criteria remaining in the judgment condition setting table 342 for which judgment has not been completed (step S25; NO), the CPU 311 adds 1 to the value of i (step S26), returns to step S22, and repeats the process based on the judgment criterion corresponding to the next judgment priority.
[0100] In step S25, if the judgment is completed for all criteria set in the judgment condition setting table 342 (step S25; YES), the CPU 311 determines any one device from the narrowed down voice operation target candidates as the operation target (response target), and determines the other devices as cancellation targets (step S27).
[0101] In step S24, when the voice operation target candidates are narrowed down to one (step S24; YES), the CPU 311 determines the narrowed down voice operation target candidate as the operation target (response target) (step S28). After step S27 or step S28, the execution device determination process ends.
[0102] As described above, according to the first embodiment, when voice information relating to the same operation instruction is acquired from multiple devices installed in close proximity, the device management apparatus 30 can determine whether to operate each device or cancel it. Therefore, it is possible to select an appropriate device from among multiple voice-operable devices as the operation target.
[0103] Furthermore, if the device management device 30 determines that the time at which voice information was acquired is the same across multiple devices and that the job content indicated by the voice information is the same across devices, it determines that voice information relating to the same operation instruction has been acquired, thereby making it possible to appropriately determine whether the same utterance was made by the same user.
[0104] For example, if the voice characteristics of the voice information match, it can be determined that the job contents indicated by the voice information are the same. Furthermore, if the character information (character string) obtained by performing a voice recognition process on the voice information matches a predetermined percentage or more, it can be determined that the job contents indicated by the voice information are the same. Furthermore, if the character information recognized from the voice information matches the judgment result (job command, etc.) determined by the job content judgment algorithm, it can be determined that the job content indicated by the voice information is the same.
[0105] Furthermore, in the process of determining an operation target from among voice operation target candidates, a device that can respond more quickly can be determined as the operation target based on, for example, the sleep state or sleep level. Furthermore, devices that are powered off cannot respond, so by treating them as cancellation targets, the process of narrowing down the operation targets can be performed efficiently.
[0106] In addition, devices that cannot execute voice instructions based on the configurable items on the device, the status of optional items installed, installed software, etc. can be canceled, thereby efficiently narrowing down the objects to be operated.
[0107] Furthermore, by selecting the device with the loudest input sound of voice information as the operation target, it is possible to determine the device that is considered to be closest to the user who has input voice as the operation target.
[0108] In the first embodiment, the device management apparatus 30 combines multiple criteria in the process of determining the operation target from among the candidate voice operation targets, but the criteria to be used can be changed arbitrarily, or a single criterion can be used.
[0109] In the first embodiment, "audio volume" is set as one of the criteria in the judgment condition setting table 342 (see FIG. 5), but first, for each of the voice operation target candidates, the device status is compared to try to search for a device that is more appropriate as an operation target, and if there is no difference in the conditions of each device and the operation targets cannot be narrowed down, additional narrowing down based on the audio volume may be performed. Specifically, among the devices that were not determined as cancellation targets based on the comparison of the device status, the device with the loudest input sound is determined as the operation target, and the others are determined as cancellation targets.
[0110] [Second embodiment] Next, a second embodiment to which the present invention is applied will be described. The device management system in the second embodiment has the same configuration as the device management system 100 shown in the first embodiment, and therefore, the illustration and description of the configuration will be omitted and Figures 1 to 3 will be used. The configuration and processing characteristic of the second embodiment will be described below.
[0111] <Configuration of device management device> When a user logs in to a device (image forming apparatus 10), the device transmits to the device management apparatus 30 the user ID of the user currently logged in to the device (manually operating the device). The CPU 311 of the device management apparatus 30 associates the user ID received from the device with the identification information of the device and stores the information in the device management table 343 of the storage unit .
[0112] The CPU 311 (extraction means) compares the voice information acquired from two or more devices among the plurality of devices, and extracts, from the two or more devices, a device that has acquired voice information relating to the same operation instruction as a candidate for voice operation.
[0113] The CPU 311 obtains voice feature data from the voice information relating to the same operation instruction, and acquires a user ID corresponding to the voice feature data from the user authentication table 341 (see Figure 4) in the memory unit 34, thereby identifying the user corresponding to the voice information relating to the same operation instruction (the user who performed the voice input).
[0114] The CPU 311 (determination means) determines which of the voice operation target candidates the device is to execute the job content included in the same operation instruction.
[0115] Among the voice operation target candidates, CPU 311 selects as the operation target a device that is being manually operated by the same user as the user corresponding to the voice information related to the same operation instruction. When determining an operation target from the voice operation target candidates, CPU 311 refers to device management table 343 in storage unit 34, acquires a user ID (user ID of the logged-in user) associated with the identification information of each device, and if there is a device that is being manually operated by the same user as the user corresponding to the voice information related to the same operation instruction, selects that device as the operation target. Note that, when determining an operation target from the voice operation target candidates, CPU 311 may acquire the latest information about the user ID of the user manually operating each device from each device.
[0116] The CPU 311 (execution means) causes the determined device to execute the job content included in the same operation instruction.
[0117] <Device management system operation> Next, the operation of the device management system according to the second embodiment will be described. 9 and 10 are ladder charts showing a second device management process executed by the device management system of the second embodiment. In the second device management process, the device being manually operated by the user who has performed the voice input is set as the operation target.
[0118] First, the user manually performs an authentication operation on the image forming apparatus 10A (step S31). Specifically, the user inputs a user ID and a password through the operation unit 152 of the image forming apparatus 10A.
[0119] In image forming apparatus 10A, user authentication unit 14 performs user authentication (step S32). If the combination of the user ID and password entered by the user has been registered in advance, user authentication unit 14 determines that the user is a valid user and identifies the user who has logged in to image forming apparatus 10A. CPU 111 of image forming apparatus 10A stores the user ID of the identified user in storage unit 17 as the user ID of the currently logged-in user.
[0120] Next, the CPU 111 of the image forming apparatus 10A transmits the user ID of the user currently logged in to the image forming apparatus 10A to the device management apparatus 30 via the communication unit 16 (step S33). In the device management device 30, the CPU 311 acquires the user ID of the user currently logged in to the image forming device 10A via the communication unit 33, associates the acquired user ID with the identification information of the image forming device 10A, and stores the associated user ID in the device management table 343 of the memory unit 34 (step S34).
[0121] The processing of steps S35 to S40 is the same as the processing of steps S5 to S10 in the first device management processing (see FIGS. 6 and 7), and therefore a description thereof will be omitted.
[0122] 10, the CPU 311 of the device management apparatus 30 performs user authentication based on the voice information related to the same operation instruction (step S41). Specifically, the CPU 311 refers to the user authentication table 341 (see FIG. 4) in the storage unit 34, acquires a user ID corresponding to the voice feature data obtained by analyzing the voice information related to the same operation instruction, and identifies the user who made the voice input in step S35. CPU 311 of device management apparatus 30 determines, from among the candidates for voice operation targets, the device to which the user who made the voice input (the user identified in step S41) is logged in, as the operation target (step S42). That is, CPU 311 determines, as the operation target, image forming apparatus 10A to which the user who made the voice input has already logged in and is manually operating, CPU 311 accordingly determines image forming apparatus 10B as the cancellation target.
[0123] The processing of steps S43 to S46 is similar to the processing of steps S12 to S15 of the first device management processing (see FIGS. 6 and 7), and therefore a description thereof will be omitted. This completes the second device management process.
[0124] After that, when the user performs a logout operation from the operation unit 152 of the image forming device 10A, the CPU 111 of the image forming device 10A stores the fact that the user has logged out in the memory unit 17, and transmits logout information indicating that the user has logged out of the image forming device 10A to the device management device 30 via the communication unit 16. In the device management device 30, when the CPU 311 acquires logout information from the image forming device 10A via the communication unit 33, it changes the "user ID of the logged-in user" corresponding to the identification information of the image forming device 10A in the device management table 343 stored in the memory unit 34 to "not logged in."
[0125] As described above, according to the second embodiment, it is possible to select an appropriate device from among a plurality of voice-operable devices as the operation target. Specifically, if there is a device among the candidates for a voice operation target that is being manually operated by the same user as the user corresponding to the voice information relating to the same operation instruction, it can be set as the operation target as is.
[0126] In steps S31 and S32 of the second device management process, user authentication is performed based on information (user ID and password) manually entered by the user from the operation unit 152, but the user authentication unit 14 of the image forming device 10A may also identify the user based on an IC card carried by the user or the user's biometric information. Furthermore, each image forming device 10A may have a table similar to the user authentication table 341 (see FIG. 4) stored in the memory unit 34 of the device management device 30, and may obtain voice characteristics from voice information obtained by voice input to identify the user who is using the image forming device 10A.
[0127] [Variations] Next, a modification of the second embodiment will be described. When starting a voice operation, if there is no image forming device 10 to which the user making the voice input is logged in, but there is an image forming device 10 being used by another user, the image forming device 10 that is not being used by a user will be the object of operation. In other words, the device being used by another user will be the object of cancellation.
[0128] For example, assume that user X is logged in to image forming device 10A, and no user is logged in to image forming device 10B. If CPU 311 of device management device 30 identifies that the voice related to the voice information acquired when a voice operation is performed is from user Y, image forming device 10A, which is being used by another user (user X), is deemed unavailable for response (cancellation target), and image forming device 10B is determined to be the target for operation.
[0129] The CPU 311 of the device management apparatus 30 selects, as a cancellation target, a device that is being manually operated by a user other than the user corresponding to the voice information related to the same operation instruction, from among the candidate voice operation targets. When determining an operation target from among the candidate voice operation targets, the CPU 311 references the device management table 343 in the storage unit 34 to acquire a user ID associated with the identification information of each device, and selects a device that is being manually operated by a user other than the user corresponding to the voice information related to the same operation instruction as a cancellation target. Note that, when selecting an operation target from among the candidate voice operation targets, the CPU 311 may acquire the latest information about the user ID of the user manually operating each device from each device.
[0130] Furthermore, the CPU 311 of the device management apparatus 30 designates, among the candidate voice-operated devices, devices being voice-operated by a user other than the user corresponding to the voice information relating to the same operation instruction as a candidate for cancellation. Specifically, the CPU 311 obtains voice feature data from voice information acquired from each candidate voice-operated device (voice information different from the voice information relating to the same operation instruction and received within a few minutes before or after the time of the reception of the voice information relating to the same operation instruction), and acquires a user ID corresponding to the voice feature data from the user authentication table 341 (see FIG. 4 ) in the storage unit 34 to identify the user who is voice-operating each device. Then, if the user identified as being voice-operated by each device is a user other than the user corresponding to the voice information relating to the same operation instruction, the CPU 311 designates the device being voice-operated by the other user as a candidate for cancellation.
[0131] According to the modified example, among the candidates for voice operation, devices that are clearly being used (manually operated or voice operated) by someone other than the user who made the voice input can be canceled, thereby efficiently narrowing down the targets for operation.
[0132] The above-described embodiments and modifications are merely examples of the device management system according to the present invention, and the present invention is not limited to these. The detailed configuration and operation of each device constituting the system can be modified as appropriate without departing from the spirit of the present invention. For example, the characteristic processes in the above-described embodiments and modifications may be combined.
[0133] In the above-described embodiments and modifications, the device management apparatus 30 is provided independently of each device (image forming apparatus 10), but the voice processing function (function for determining the device to be operated from voice information) of the device management apparatus 30 may be installed in any of the devices. In other words, a serverless configuration may be achieved by managing the contents of voice instructions and the status of each device within any of the devices. Specifically, when there are image forming apparatuses 10A and 10B, image forming apparatus 10A is set as the parent device and image forming apparatus 10B is set as the child device, and image forming apparatus 10A is set as the connection destination of image forming apparatus 10B. When voice input is made to image forming apparatus 10B (the smart speaker 20B associated with image forming apparatus 10B), image forming apparatus 10B (the smart speaker 20B associated with image forming apparatus 10A) sends voice information to the voice processing function of image forming apparatus 10A to inquire whether simultaneous voice input is available. When voice input is received simultaneously from the parent device and the child device, the voice processing function installed in the image forming apparatus 10A determines which device is to be operated and notifies each of the parent device and the child device of the determination result as to whether or not the device is to be operated.
[0134] Furthermore, in the above-described embodiments and variations, the image forming device 10 is provided with a user authentication unit 14 that identifies the user who uses the device itself, but an external server (such as the device management device 30) may also be provided with a function to identify the user who uses the image forming device 10.
[0135] In addition, in each of the above embodiments and variant examples, the time when the device management device 30 received the voice information from the smart speaker 20 was treated as the time when the voice information was acquired in the image forming device 10, but the smart speaker 20 may add information about the time when the voice information was acquired to the voice information and then send the voice information to the device management device 30.
[0136] Furthermore, in each of the above embodiments and variant examples, the mechanism for acquiring voice information (smart speaker 20) and the mechanism to be operated by voice (image forming device 10) are configured independently, but the mechanism to be operated by voice (image forming device 10) may have a built-in mechanism for acquiring voice information (smart speaker 20).
[0137] Furthermore, in the above embodiments and variant examples, the case where the image forming apparatus 10 (MFP) is used as the multiple devices has been described, but the multiple devices managed by the device management apparatus 30 are not limited to this and may be other devices.
[0138] In the above description, an example has been disclosed in which a ROM is used as a computer-readable medium for storing a program for executing each process, but this is not limiting. Other computer-readable media may also be used, such as non-volatile memory such as flash memory, or portable recording media. Furthermore, a carrier wave may also be used as a medium for providing program data via a communication line. [Explanation of symbols]
[0139] 10A,10B,10C,... Image forming device 10 Image forming device 11 Control section 12 Document reading section 13 Image forming unit 14 User authentication section 15 Operation panel 16 Communications Department 17 Memory section 20A, 20B, 20C, Smart Speaker 20. Smart Speaker 30 Device management device 31 Control Unit 32 Timing section 33 Communications Department 34 Storage section 100 Device Management System N Communication Network
Claims
1. A device management apparatus that manages a plurality of voice-operable devices by receiving voice input from a user regarding job content and acquiring voice information, the device management apparatus comprising: an extraction means for comparing voice information acquired from two or more devices among the plurality of devices, and extracting, from the two or more devices, a device that has acquired voice information relating to the same operation instruction as a voice operation target candidate; a determining means for determining which of the candidate devices to be operated by the voice command should execute the job content included in the same operation instruction; an execution unit that causes the device determined by the determination unit to execute the job content included in the same operation instruction; A device management apparatus comprising:
2. The device management device according to claim 1, wherein the extraction means determines that the voice information relating to the same operation instruction has been acquired if it determines that the time at which the voice information was acquired is the same between the two or more devices and that the content of the job indicated by the voice information is the same between the devices.
3. The device management apparatus according to claim 2 , wherein the extraction unit determines whether the job contents are the same by comparing voice characteristics of the voice information acquired from each of the two or more devices.
4. The device management apparatus according to claim 2 , wherein the extraction unit determines whether the job contents are the same by comparing character information obtained by voice recognition from the voice information acquired from each of the two or more devices.
5. The device management device according to claim 2, wherein the extraction means determines whether the job contents are the same by comparing the determination results obtained by a job content determination algorithm based on character information obtained by voice recognition from the voice information acquired from each of the two or more devices.
6. The device management device according to any one of claims 1 to 5, wherein the determination unit selects, as the cancellation target, a device among the candidates for voice operation other than the device determined to have the fastest response based on the sleep state or sleep level of the candidates for voice operation.
7. The device management device described in any one of claims 1 to 6, wherein the determination means cancels devices among the candidate voice operation targets in which the power supply to the mechanism that acquires voice information and the power supply to the voice operation target mechanism are independent.
8. The device management device according to any one of claims 1 to 7, wherein the determination unit determines, as the operation target, a device that is being manually operated by the same user as the user corresponding to the voice information relating to the same operation instruction, from among the candidate voice operation targets.
9. The device management device according to any one of claims 1 to 8, wherein the determination unit determines, among the candidates for voice operation targets, a device that is being manually operated by a user other than the user corresponding to the voice information relating to the same operation instruction as the candidate for voice operation to be canceled.
10. The device management device according to any one of claims 1 to 9, wherein the determination unit determines, among the candidates for voice operation, a device that is being voice-operated by a user other than the user corresponding to the voice information relating to the same operation instruction as the candidate for voice operation.
11. The device management device according to any one of claims 1 to 10, wherein the determination means compares the setting items that can be set for each of the candidates for voice operation with the voice instruction content relating to the same operation instruction, and sets devices other than those that can execute the voice instruction content as cancellation targets.
12. The device management device described in any one of claims 1 to 11, wherein the determination means compares the installation status of optional items attached to each of the candidate voice operation targets with the voice instruction content related to the same operation instruction, and sets devices other than devices that can execute the voice instruction content as targets for cancellation.
13. The device management device according to any one of claims 1 to 12, wherein the determination means compares the software installed on each of the candidates for voice operation with the voice instruction content relating to the same operation instruction, and selects devices other than those capable of executing the voice instruction content as cancellation targets.
14. The device management apparatus according to claim 1 , wherein the determining unit compares the volume of the input sound of the voice information acquired for each of the voice operation target candidates, and selects the device with the loudest input sound as the operation target.
15. A device management system including a plurality of voice-operable devices and a device management apparatus that manages the plurality of devices, each of the plurality of devices includes a voice acquisition unit that receives voice input from a user regarding the content of a job and acquires voice information; the device management device, an extraction means for comparing voice information acquired from two or more devices among the plurality of devices, and extracting, from the two or more devices, a device that has acquired voice information relating to the same operation instruction as a voice operation target candidate; a determining means for determining which of the candidate devices to be operated by the voice command should execute the job content included in the same operation instruction; an execution unit that causes the device determined by the determination unit to execute the job content included in the same operation instruction; A device management system comprising:
16. A device management method for managing a plurality of voice-operable devices, comprising: a voice acquisition step of receiving voice input from a user regarding the content of a job and acquiring voice information in two or more devices among the plurality of devices; an extraction step of comparing the voice information acquired in each of the two or more devices and extracting, from the two or more devices, a device that has acquired voice information relating to the same operation instruction as a voice operation target candidate; a determining step of determining which of the voice operation target candidates should be caused to execute the job content included in the same operation instruction; an execution step of causing the device determined in the determination step to execute the content of the job included in the same operation instruction; A device management method including:
17. A computer of a device management apparatus that manages a plurality of voice-operable devices by receiving voice input from a user regarding the content of a job and acquiring voice information, an extraction means for comparing voice information acquired from two or more devices among the plurality of devices, and extracting, from the two or more devices, a device that has acquired voice information relating to the same operation instruction as a voice operation target candidate; a determining means for determining which of the candidate devices to be operated by the voice command should execute the job content included in the same operation instruction; an execution unit that causes the device determined by the determination unit to execute the job content included in the same operation instruction; A program to function as a
Citation Information
Patent Citations
Voice control system, control method, and program
JP2019095835A
Image forming apparatus, image forming system, and information processing method
JP2020127104A
Voice control system, control method, and non-transitory computer-readable storage medium storing program
US20190156824A1