Image forming apparatus
The image forming device uses a reduced number of identification units to synthesize sound data from multiple sensors, enhancing abnormality detection accuracy and reducing costs by generating composite images for efficient analysis.
Patent Information
- Application Number
- JP2024051361
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-10-09
AI Technical Summary
Existing image forming apparatuses require multiple hardware units for sound detection and abnormality identification, leading to increased hardware and processing costs and complexity.
An image forming device with a reduced number of identification units that utilize multiple sound detection units, generating composite images from sound data to identify abnormalities using machine learning and autoencoders, reducing hardware and processing requirements.
This approach reduces hardware and information transmission needs while improving abnormality detection accuracy and efficiency by synthesizing sound data from multiple sensors into a composite image for analysis.
Smart Images

Figure 2025150475000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image forming apparatus. [Background technology]
[0002] Patent Document 1 discloses an anomaly detection system that is operable to infer latent variables from input data to be subjected to anomaly detection based on an encoder of a VAE that has been trained in advance with training data including normal data, generate reconstructed data from the latent variables based on a decoder of the VAE that has been trained in advance, and determine whether the input data is normal or abnormal based on the input data and the reconstructed data. Patent Document 2 discloses a process that includes multiple acoustic sensors that detect the operating sounds of a device, compares previously collected operating sounds of the device with newly collected operating sounds, and determines the abnormal state and timing of the device. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6740247 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-184722 Summary of the Invention [Problem to be solved by the invention]
[0004] If an image forming apparatus is provided with a sound detection unit and a functional unit for identifying an abnormality in the image forming apparatus in correspondence with the sound detection unit, it is possible to identify an abnormality in the image forming apparatus. Here, if a functional unit for identifying an abnormality is provided in correspondence with each of the multiple sound detection units, the abnormality will be identified individually for each functional unit. An object of the present invention is to reduce the amount of hardware required to identify abnormalities compared to when a functional unit that identifies abnormalities is provided individually in correspondence with each of multiple sound detection units. [Means for solving the problem]
[0005] The invention described in claim 1 is an image forming device comprising an image forming unit that forms an image on a recording medium, a plurality of sound detection units that detect sounds from the image forming unit, and an identification unit that identifies abnormalities in the image forming unit based on information output from the plurality of sound detection units, wherein the number of identification units is less than the number of the plurality of sound detection units. The invention described in claim 2 is the image forming apparatus described in claim 1, further comprising a transmitting means for transmitting information indicating that an abnormality has occurred to an external device when an abnormality is identified by the identification unit. The invention described in claim 3 is the image forming device described in claim 1, wherein the identification unit performs machine learning using images generated based on information obtained by the sound detection unit and generated for each sound detection unit, calculates latent variables from features extracted from the images using the machine-learned anomaly detection model, generates an output image by restoring it using the latent variables, and identifies an abnormality in the image forming unit by comparing an input image with the output image. The invention described in claim 4 is the image forming device described in claim 1, further comprising a generation means for generating an image based on information obtained by the sound detection unit and generated for each sound detection unit, the image being generated based on a plurality of the images generated in accordance with the plurality of sound detection units provided, and generating a composite image which is the image obtained by the synthesis, and the identification unit identifies an abnormality based on the composite image generated by the generation means. The invention described in claim 5 is the image forming device described in claim 4, wherein each of the images generated for each sound detection unit has a time axis and a frequency axis and is an image that expresses sound intensity with pixel values, and the generation means generates the composite image in which each of the multiple images is arranged so that the time axis of each of the images is aligned in a specific direction. The invention described in claim 6 is the image forming device described in claim 5, wherein the generation means generates the composite image by synthesizing images based on multiple images, in which each of the multiple images is arranged so that the time axis of each of the multiple images is aligned with the one direction, and the multiple images are arranged side by side in a direction intersecting the one direction. The invention described in claim 7 is the image forming device described in claim 5, in which the generation means generates the composite image by synthesizing images based on multiple images, and generates the composite image in which the multiple images are arranged in a manner that is synchronized in time. The invention described in claim 8 is the image forming device described in claim 6, wherein the frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends and whose position in the direction in which the frequency axis extends is different from the first side, one of the first side and the second side is a low-frequency side side located on the low-frequency side and is arranged on the side of the image where low frequency component values are displayed, and the other side is a high-frequency side side located on the high-frequency side and is arranged on the side of the image where high frequency component values are displayed, and the generation means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, and generates a composite image in which the high-frequency side side of the one image is located on the other image side and the high-frequency side side of the other image is located on the first image side. The invention described in claim 9 is the image forming device described in claim 8, in which the generation means, when synthesizing images using at least the one image and the other image to generate the composite image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image. The invention described in claim 10 is an image forming apparatus described in claim 9, wherein the density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the plurality of pixels that are arranged in the same direction and that border the intermediate image, among the plurality of pixels that constitute the one image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the plurality of pixels that are arranged in the same direction, and the density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the plurality of pixels that are arranged in the same direction and that border the intermediate image among the plurality of pixels that constitute the other image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the plurality of pixels that are arranged in the same direction. The invention described in claim 11 is the image forming device described in claim 6, wherein the frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends and whose position in the direction in which the frequency axis extends is different from the first side, one of the first side and the second side is a low-frequency side side located on the low-frequency side and is arranged on the side of the image where low frequency component values are displayed, and the other side is a high-frequency side side located on the high-frequency side and is arranged on the side of the image where high frequency component values are displayed, and when the generation means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, it generates a composite image in which the low-frequency side side of the one image is located on the other image side and the low-frequency side side of the other image is located on the one image side. The invention described in claim 12 is the image forming device described in claim 11, in which the generation means, when synthesizing images using at least the one image and the other image to generate the composite image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image. The invention described in claim 13 is an image forming apparatus described in claim 12, wherein the density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the plurality of pixels that are arranged in the same direction and that are adjacent to the intermediate image, the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the plurality of pixels that are arranged in the same direction among the plurality of pixels that constitute the one image, the density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the plurality of pixels that are arranged in the same direction and that are adjacent to the intermediate image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the plurality of pixels that are arranged in the same direction among the plurality of pixels that constitute the other image. The invention described in claim 14 is an image forming apparatus comprising an image forming unit that forms an image on a recording medium, a plurality of sound detection units that detect sound from the image forming unit, a generation means that generates an image based on information obtained by the sound detection unit and for each sound detection unit, and that synthesizes images based on the images generated in plurality according to the plurality of sound detection units, to generate a composite image that is the image obtained by the synthesis, and an identification unit that identifies an abnormality in the image forming unit based on the composite image. [Effects of the Invention]
[0006] According to the invention of claim 1, it is possible to reduce the amount of hardware required to identify abnormalities compared to when a functional unit that identifies abnormalities is provided individually in correspondence with each of the multiple sound detection units. According to the invention of claim 2, the amount of information transmitted to the external device can be reduced compared to when information related to an abnormality is transmitted to the external device regardless of whether an abnormality has been identified or not. According to the invention of claim 3, even in a situation where the image forming apparatus basically produces only normal sounds and is unlikely to produce abnormal sounds, it becomes possible to identify abnormalities occurring in the image forming apparatus. According to the invention of claim 4, the processing load required to identify an abnormality can be reduced compared to when identifying an abnormality for each image obtained by each sound detection unit. According to the invention of claim 5, the accuracy of identifying abnormalities can be improved compared to when the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends. According to the invention of claim 6, the accuracy of identifying abnormalities can be improved compared to when the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends. According to the invention of claim 7, it is possible to improve the accuracy of identifying an abnormality compared to when a plurality of images are not arranged in a time-synchronized manner. According to the invention of claim 8, the accuracy of identifying abnormalities can be improved compared to when the low-frequency side of one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image. According to the invention of claim 9, it is possible to improve the accuracy of identifying an abnormality compared to when one image and another image are in direct contact with each other. According to the invention of claim 10, the accuracy of identifying an abnormality can be improved compared to when one image and another image are in direct contact with each other. According to the invention of claim 11, the accuracy of identifying abnormalities can be improved compared to when the low-frequency side of one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image. According to the invention of claim 12, the accuracy of identifying an abnormality can be improved compared to when one image and another image are in direct contact with each other. According to the invention of claim 13, it is possible to improve the accuracy of identifying an abnormality compared to when one image and another image are in direct contact with each other. According to the invention of claim 14, it is possible to reduce the amount of hardware required to identify an abnormality compared to when an abnormality is identified without generating a composite image. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 illustrates an example of a diagnostic system. [Figure 2] FIG. 1 is a diagram illustrating an image forming apparatus. [Figure 3] FIG. 2 illustrates an example of the hardware configuration of a specifier. [Figure 4] FIG. 10 is a diagram illustrating a process performed by a specifier. [Figure 5] 10A and 10B are diagrams showing sound corresponding images generated by a corresponding image generating unit; [Figure 6] 10A and 10B are diagrams illustrating a synthetic image generation process performed by a synthetic image generation unit. [Figure 7] 10A and 10B are diagrams illustrating synchronization between sound-corresponding images. [Figure 8] 10A and 10B are diagrams showing another example of processing for generating a composite image. [Figure 9] 10A and 10B are diagrams showing other processing examples. [Figure 10] FIG. 10 is an enlarged view of a portion indicated by the symbol X in FIG. 9. [Figure 11] 10A and 10B are diagrams showing another example of processing for generating a composite image. [Figure 12] 10A and 10B are diagrams showing other processing examples. [Figure 13] FIG. 13 is an enlarged view of a portion indicated by reference numeral XIII in FIG. 12. [Figure 14] FIG. 10 is a diagram showing another example of a composite image. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a diagram showing an example of a diagnostic system 1. As shown in FIG. The diagnostic system 1 of this embodiment is provided with a plurality of image forming apparatuses 100 and a server device 200 connected to each of the plurality of image forming apparatuses 100 via a communication line 190. In FIG. 1, one image forming apparatus 100 out of a plurality of image forming apparatuses 100 is shown.
[0009] In this embodiment, the server device 200 as an example of an information processing system acquires information about each of the image forming devices 100. The diagnostic system 1 further includes a user terminal 300. The user terminal 300 is connected to the server device 200. The user terminal 300 accepts operations from a user. An example of the user is a maintenance person of the image forming device 100. In this embodiment, the user terminal 300 that the maintenance person refers to is provided.
[0010] The user terminal 300 is provided with a display device 310. The user terminal 300 is realized by a computer. Examples of the form of the user terminal 300 include a PC (Personal Computer), a smartphone, and a tablet terminal. The image forming apparatus 100 is provided with an image forming section 100A that forms an image on a sheet of paper, which is an example of a recording medium. Furthermore, although not shown in FIG. 1, the image forming apparatus 100 is provided with a sound sensor and the like.
[0011] FIG. 2 is a diagram illustrating the image forming apparatus 100. As shown in FIG. In this embodiment, as described above, the image forming apparatus 100 is provided with the image forming unit 100A that forms an image on a sheet P, which is an example of a recording medium. The image forming unit 100A uses an electrophotographic method to form an image. The image forming section 100A, which is an example of an image forming means, is provided with an intermediate transfer belt 108, which is a rotating member, and a plurality of image forming units 107, which form images of different colors.
[0012] In this embodiment, the image formed by each of the plurality of image forming units 107 is first transferred onto the intermediate transfer belt 108 and then transferred onto the paper P. The multiple image forming units 107 form images of different colors on the intermediate transfer belt 108. Note that the intermediate transfer belt 108 is not essential. An image may be directly transferred onto the paper P from each of the multiple image forming units 107. Furthermore, multiple image forming units 107 are not essential. A configuration may be provided with only one image forming unit 107. In the case of a configuration in which only one image forming unit 107 is provided, the intermediate transfer belt 108 is omitted.
[0013] In this embodiment, the image forming units 107 include an image forming unit 107Y that forms a yellow image. Further, an image forming unit 107M that forms a magenta image is provided. Further, an image forming unit 107C that forms a cyan image is provided. Further, an image forming unit 107K that forms a black image is provided. The images formed by the image forming units 107 are transferred to an intermediate transfer belt 108, which is an example of a transfer member. Thereafter, this image is transferred at the transfer section T onto the paper P that is being conveyed to the transfer section T. In this way, the image is formed on the paper P.
[0014] Each of the image forming units 107 is provided with a photosensitive drum 101, which is an example of an image carrier. The photosensitive drum 101 rotates in a clockwise direction. Furthermore, each of the image forming units 107 is provided with a charging device 101C that charges the photosensitive drum 101. Each of the image forming units 107 is provided with an exposure device 102 that exposes the photosensitive drum 101 to light. Furthermore, each of the image forming units 107 is provided with a developing device 103. The developing device 103 develops the electrostatic latent image formed on the photosensitive drum 101 by exposure by the exposure device .
[0015] The developing device 103 is provided with a developing roll 103A disposed opposite the photosensitive drum 101. In this embodiment, the developer adhering to the outer peripheral surface of the developing roll 103A moves to the surface of the photosensitive drum 101. This causes development. When development is performed, an image made of, for example, toner is formed on the photosensitive drum 101. This image is then transferred to the outer peripheral surface of the intermediate transfer belt 108. This image on the intermediate transfer belt 108 is then transferred to paper P, and an image is formed on this paper P. The image forming unit 100A may form an image on the paper P using other methods, such as an inkjet method, instead of the electrophotographic method.
[0016] The image forming apparatus 100 is further provided with an image reading device 130 . The image reading device 130, which is an example of an image reading means, is a so-called scanner that reads an image formed on a sheet of paper (not shown), which is an example of a recording medium. The image reading device 130 includes a light source that emits light to be irradiated onto the paper, and a light receiving unit such as a CCD that receives the light reflected from the paper. In this embodiment, read image data of the paper is generated based on the reflected light received by the light receiving unit.
[0017] Furthermore, each image forming apparatus 100 is provided with an operation receiving unit 132 that receives operations from a user who uses the image forming apparatus 100. The operation reception unit 132 is configured by a so-called touch panel. The operation reception unit 132 displays information to the user and receives operations performed by the user. It should be noted that the display of information to the user and the reception of user operations are not limited to being performed by one operation reception unit 132, and the operation reception unit 132 and the information display unit may be provided separately.
[0018] Furthermore, each image forming apparatus 100 is provided with a sound sensor 120 as an example of a detector that detects the sound of image forming section 100A. This sound sensor 120 can also be considered a microphone. In this embodiment, a plurality of sound sensors 120 are provided in each image forming apparatus 100. In the example shown in Fig. 2, two sound sensors 120 are provided in the image forming apparatus 100. Furthermore, in this embodiment, a specifier 400 is provided as an example of a specifier that specifies an abnormality in the image forming unit 100A based on information output from the sound sensor 120.
[0019] In this embodiment, the number of identifiers 400 provided is less than the number of the plurality of sound sensors 120. Specifically, in this embodiment, one identifier 400 is provided. The identifier 400 identifies an abnormality in the image forming unit 100A based on information output from the multiple sound sensors 120. The identifier 400 identifies an abnormal sound generated in the image forming device 100, and identifies an abnormality in the image forming unit 100A based on this abnormal sound. Here, the "abnormal sound" refers to a sound generated due to a malfunction of the image forming device 100.
[0020] Furthermore, corresponding image generating units 410 are provided in correspondence with the respective sound sensors 120 . Each of the corresponding image generating units 410 converts the information output from the sound sensor 120 into an image and generates an image based on the information output from the sound sensor 120. In other words, each of the corresponding image generating units 410 generates an image representing the sound obtained by the sound sensor 120. Hereinafter, in this specification, the image generated by the corresponding image generating unit 410 will be referred to as a "sound corresponding image." The corresponding image generating units 410 are provided in correspondence with the respective sound sensors 120. Two corresponding image generating units 410 are provided.
[0021] In this embodiment, a corresponding image generating unit 410 is provided for each sound sensor 120, and a sound corresponding image is generated for each sound sensor 120. Furthermore, in this embodiment, a composite image generating unit 420 that generates a composite image is provided. The composite image generating unit 420, which is an example of a generating means, generates a composite image (details of which will be described later) by synthesizing images based on a plurality of generated sound corresponding images. In this embodiment, the composite image generating unit 420 generates a number of composite images that is smaller than the number of sound corresponding images.
[0022] In this embodiment, the corresponding image generating unit 410 and the composite image generating unit 420 are realized by, for example, a computer (not shown). More specifically, the corresponding image generating unit 410 and the composite image generating unit 420 are realized by a CPU (not shown) serving as an example of a processor executing a program relating to image processing stored in a ROM or the like.
[0023] Once the composite image is generated by the composite image generation unit 420, the composite image is input to the specification unit 400. In this embodiment, the generated composite image is analyzed by the identifier 400, which serves as an example of an identifying means. If an abnormal sound is generated in the image forming apparatus 100, the abnormal sound is identified by the identifier 400. In this embodiment, the identifier 400 identifies abnormal noises occurring in the image forming apparatus 100. In this case, the identifier 400 identifies an abnormality in the image forming unit 100A.
[0024] FIG. 3 is a diagram showing an example of the hardware configuration of the identifier 400. As shown in FIG. The specifying unit 400 is realized by a computer. The specifying unit 400 has an arithmetic processing unit 11 that executes digital arithmetic processing according to a program, and a secondary storage unit 12 that stores information. The secondary storage unit 12 is realized by an existing information storage device such as an HDD (Hard Disk Drive), a semiconductor memory, or a magnetic tape.
[0025] The arithmetic processing unit 11 is provided with a CPU 11a as an example of a processor. The arithmetic processing unit 11 also includes a RAM 11b used as a working memory for the CPU 11a, and a ROM 11c in which programs executed by the CPU 11a are stored. The arithmetic processing unit 11 is also provided with a non-volatile memory 11d that is rewritable and can retain data even if the power supply is interrupted, and an interface unit 11e that controls each part, such as a communication unit, connected to the arithmetic processing unit 11.
[0026] The nonvolatile memory 11d is configured with, for example, a battery-backed SRAM, a flash memory, etc. The secondary storage unit 12 stores files, etc., as well as programs executed by the arithmetic processing unit 11. In this embodiment, the CPU 11a reads a program stored in the ROM 11c or the secondary storage unit 12, and executes each process.
[0027] The program executed by the CPU 11a may be provided to the identifying unit 400 in a state where it is stored in a computer-readable recording medium such as a magnetic recording medium (such as a magnetic tape or a magnetic disk), an optical recording medium (such as an optical disk), a magneto-optical recording medium, or a semiconductor memory. The program executed by the CPU 11a may also be provided to the identifying unit 400 using a communication means such as the Internet.
[0028] In this specification, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Furthermore, the operations of the processors may not only be performed by a single processor, but may also be performed by multiple processors located at physically separate locations working together. The order of the operations of the processors is not limited to the order described in this embodiment, and may be changed.
[0029] In this embodiment, the identifier 400 (see FIG. 2) identifies allophones using an autoencoder, which uses the same data for both the input and output layers and performs unsupervised learning. The identifier 400 identifies abnormal sounds using the sound corresponding image generated based on the information obtained by the sound sensor 120 and an autoencoder. More specifically, the identifier 400 identifies an allophone using an autoencoder and a synthetic image generated by the synthetic image generator 420 based on two sound corresponding images.
[0030] FIG. 4 is a diagram illustrating the processing performed by the identifier 400. As shown in FIG. As described above, the identifier 400 identifies an allophone using an autoencoder. In this embodiment, as shown in FIG. 4, a composite image generated by a composite image generating unit 420 (see FIG. 2) is input to the specifier 400. The sounds that are normally generated by the image forming apparatus 100 are basically normal sounds. In this embodiment, the identifier 400 is usually input with a synthetic image generated based on a normal sound.
[0031] In this specifier 400 that uses an autoencoder, training is performed using synthetic images generated based on normal sounds. In other words, in this embodiment, training data is basically made up of synthetic images generated based on normal sounds, and the autoencoder is trained using the training data. In the specifier 400 of this embodiment, learning is basically performed so that the input composite image and the output composite image match. In this case, there will be no difference between the composite image before being input to the identifier 400 and the composite image output from the identifier 400. Hereinafter, in this specification, the composite image before being input to the identifier 400 may be referred to as an "input image," and the composite image output from the identifier 400 may be referred to as an "output image."
[0032] In this embodiment, when identifying an abnormal sound, the identifier 400 generates a difference image based on an input image and an output image. More specifically, the identifier 400 performs a process for each pixel, for example, subtracting the pixel values of the pixels that make up the output image from the pixel values of the pixels that make up the input image, to generate a difference image that represents the difference between the input image and the output image. In this embodiment, the input image and the output image are each an image in which sound intensity is expressed by pixel values. In this embodiment, the specifier 400 performs a process for each pixel, for example, subtracting the pixel value of a pixel that constitutes the output image from the pixel value of a pixel that constitutes the input image, to generate a difference image that represents the difference between the input image and the output image. In this case, if the input image input to the identifier 400 is an input image generated based on a normal sound, the differential image will not have image distortion or the like.
[0033] In contrast, when a synthetic image generated based on an abnormal sound that rarely occurs is input to the identifier 400, image distortion caused by this abnormal sound occurs in the differential image. In other words, when a synthetic image generated based on an abnormal sound that rarely occurs is input to the identifier 400, an image that does not appear in the case of normal sounds appears in the differential image. In this case, the identifier 400 identifies that an abnormal sound has occurred in the image forming apparatus 100. In this case, the identifier 400 outputs information indicating that an abnormal sound has occurred. The identifier 400 identifies the abnormal noise occurring in the image forming apparatus 100 based on the image that appears in the differential image. In this embodiment, the identifier 400 performs machine learning using images generated based on information obtained by the sound sensors 120 and generated for each sound sensor 120. The identifier 400 then uses the machine-learned anomaly detection model to calculate latent variables from feature amounts extracted from the images, and generates an output image by restoring the image using these latent variables. The identifier 400 then compares the input image with the output image to identify an abnormal sound occurring in the image forming apparatus 100. In other words, the identifier 400 compares the input image with the output image to identify an abnormality in the image forming unit 100A.
[0034] When the identifier 400 outputs information indicating that an abnormal sound has occurred, the information indicating that an abnormal sound has occurred is transmitted to the server device 200 (see FIG. 1), which is an example of an external device. Although not described above, in this embodiment, as shown in FIG. 2, a transmitting unit 430 is provided as an example of a transmitting means for transmitting information indicating that an abnormal noise has occurred to the server device 200. When the identifier 400 identifies an abnormal sound, the transmitter 430 transmits to the server device 200 information indicating that an abnormal sound has occurred.
[0035] The transmitting unit 430 is composed of a computer (not shown) and a known transmitting device (not shown) for transmitting information. This computer has a CPU (not shown) as an example of a processor. In the transmitting unit 430, the transmitting device operates in response to instructions from the CPU. When the identifier 400 identifies an abnormal sound, the transmitter 430 transmits to the server device 200 information indicating that an abnormal sound has occurred. In this embodiment, when the identifier 400 outputs information indicating that an abnormal sound has occurred, the identifier 400 also outputs information indicating that an abnormality in the image forming unit 100A has been identified. In this embodiment, the information indicating that an abnormality in the image forming unit 100A has been identified is also transmitted to the server device 200.
[0036] In this embodiment, the differential images sequentially generated by the identifier 400 are not transmitted to the server device 200. Furthermore, information indicating that no abnormal sound is occurring is also not transmitted to the server device 200. In this embodiment, only when the identifier 400 identifies an abnormal sound, information indicating that an abnormal sound has occurred is transmitted to the server device 200. At this time, in addition to the information indicating that an abnormal sound has occurred, other information may also be transmitted to the server device 200. For example, the content of the analysis by the identifier 400 may also be transmitted to the server device 200. In this embodiment, information indicating that an abnormality in the image forming unit 100A has been identified is also transmitted to the server device 200 only when the identifier 400 identifies an abnormal sound.
[0037] FIG. 5 is a diagram showing a sound corresponding image generated by the corresponding image generating unit 410. As shown in FIG. The corresponding image generation unit 410 (see FIG. 2) performs STFT (short-time Fourier transform) processing on the information obtained from the sound sensor 120. In other words, the corresponding image generation unit 410 performs short-time Fourier transform processing on the information obtained from the sound sensor 120. As a result, a sound corresponding image 81 shown in FIG. 5 is generated.
[0038] The sound corresponding image 81 shown in FIG. 5 has two axes, a horizontal axis 81A and a vertical axis 81B, which are orthogonal to each other. In the sound corresponding image 81, the horizontal axis 81A is the time axis, and the vertical axis 81B is the axis corresponding to the magnitude of the frequency component value. The vertical axis 81B can also be said to be the frequency axis. In this embodiment, the sound corresponding image 81 has a time axis and a frequency axis, and is an image in which the intensity of the sound is expressed by pixel values. In the sound corresponding image 81, frequency component values corresponding to high frequencies are displayed on the side farther from the time axis, and frequency component values corresponding to low frequencies are displayed on the side closer to the time axis. A white portion in the sound corresponding image 81 indicates that a sound is being generated, and a black portion in the sound corresponding image 81 indicates that no sound is being generated.
[0039] The vertical axis 81B of the sound corresponding image 81 corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound sensor 120. Each sound corresponding image 81 has a first side 81C along the direction in which horizontal axis 81A, which is the time axis, extends, and a second side 81D along the direction in which horizontal axis 81A, which is the time axis, extends. In the direction in which vertical axis 81B extends, the position of first side 81C and the position of second side 81D are different.
[0040] In this embodiment, one of the first side 81C and the second side 81D, that is, the first side 81C, is a low-frequency side side 81X located on the low-frequency side. This low-frequency side side 81X is disposed on the side of the sound corresponding image 81 where low frequency component values are displayed. The other of the first side 81C and the second side 81D, that is, the second side 81D, is a high-frequency side side 81Y located on the high-frequency side. This high-frequency side side 81Y is arranged on the side of the sound corresponding image 81 where high frequency component values are displayed.
[0041] In this embodiment, as described above, a plurality of sound sensors 120 (see FIG. 2) are provided. In this case, a specifier 400 may be provided for each sound sensor 120. In this case, the number of specifiers 400 increases in accordance with the number of sound sensors 120, which tends to increase costs. Furthermore, if a specifier 400 is provided for each sound sensor 120, the processing load on the entire image forming apparatus 100 will increase.
[0042] In contrast to this, in this embodiment, the number of installed identifiers 400 is smaller than the number of installed sound sensors 120. In this case, an increase in costs due to an increase in the number of identifiers 400 can be suppressed. Furthermore, in this case, the processing load on the entire image forming apparatus 100 is reduced. If the number of installed specifiers 400 is made smaller than the number of installed sound sensors 120, the implementation scale of the circuitry implemented in the image forming apparatus 100 is reduced.
[0043] In this embodiment, the number of installed specifiers 400 is thus smaller than the number of installed sound sensors 120. In this embodiment, in order to reduce the number of installed identifiers 400, a smaller number of synthetic images than the total number of sound sensors 120 are generated based on the sound corresponding images 81 generated for each sound sensor 120. This composite image is then input to the identifier 400. The identifier 400 then analyzes this composite image to identify abnormal sounds.
[0044] The composite image generation unit 420 (see Figure 2), which is an example of a generation means, synthesizes images based on multiple sound corresponding images 81 generated by multiple corresponding image generation units 410, and generates a composite image, which is the image obtained by this synthesis. In this embodiment, the sound corresponding image 81 is generated for each sound sensor 120 based on the information obtained by the sound sensor 120, as described above.
[0045] The composite image generation unit 420 synthesizes an image based on the sound corresponding image 81 generated by the corresponding image generation unit 410, which is a plurality of sound corresponding images 81 generated corresponding to each of the plurality of sound sensors 120 provided. As a result, a composite image, which is an image obtained by combining the sound corresponding images 81, is generated. In this embodiment, a composite image is obtained in which the number of images is smaller than the total number of sound sensors 120. Then, in this embodiment, the identifier 400 identifies abnormal sounds using an autoencoder based on this generated composite image.
[0046] In this embodiment, a maintenance person for the image forming apparatus 100 (see FIG. 1) accesses the server device 200 via the user terminal 300 and refers to information about abnormal noise stored in the server device 200. In this embodiment, when an abnormal sound is identified in the image forming apparatus 100, information indicating that the abnormal sound has occurred in the image forming apparatus 100 is stored in the server apparatus 200. If information indicating that an abnormal noise has occurred in the image forming apparatus 100 is stored in the server apparatus 200, the maintenance person for the image forming apparatus 100 takes measures for the image forming apparatus 100, such as replacing parts.
[0047] 6A and 6B are diagrams illustrating the synthetic image generation process performed by the synthetic image generation unit 420. FIG. FIG. 6(A) shows each of the sound corresponding images 81 before synthesis, and FIG. 6(B) shows the synthesized image 83. As shown in Figure 6(B), the composite image generation unit 420 generates a composite image 83 in which each of the multiple sound corresponding images 81 is arranged so that the horizontal axis 81A (not shown in Figure 6(B)) of each of the sound corresponding images 81 is aligned along a specific direction 6A.
[0048] In this embodiment, each of the sound corresponding images 81 generated for each sound sensor 120 has a horizontal axis 81A which is the time axis, and a vertical axis 81B which corresponds to the magnitude of the frequency component value obtained at each specific time, as shown in FIG. When generating the composite image 83, the composite image generating unit 420 arranges each of the sound corresponding images 81 so that the horizontal axis 81A is aligned with a specific direction 6A, as shown in FIG. 6(B). In the example shown in FIG. 6, this "one direction 6A" is the direction from the left side to the right side in the figure.
[0049] 6(B), the plurality of sound corresponding images 81 are arranged in a cross direction 6B that crosses the one direction 6A. More specifically, the plurality of sound corresponding images 81 are arranged in a direction perpendicular to the one direction 6A. The composite image generation unit 420 generates a composite image 83 in which each of the sound corresponding images 81 is arranged so that the horizontal axis 81A of each of the sound corresponding images 81 is aligned along one direction 6A, and a plurality of these sound corresponding images 81 are arranged side by side in an intersecting direction 6B that intersects with this one direction 6A.
[0050] Furthermore, when generating the composite image 83, the composite image generating unit 420 makes the plurality of sound corresponding images 81 contact each other as shown in FIG. 6(B). The arrangement of the sound corresponding images 81 is not limited to a configuration in which the multiple sound corresponding images 81 are in contact with each other. For example, the multiple sound corresponding images 81 may be configured to partially overlap each other. Also, the multiple sound corresponding images 81 may be configured to have gaps between them.
[0051] Thereafter, in this embodiment, the generated composite image 83 shown in Fig. 6(B) is input to the identifier 400 (see Fig. 2). As a result, the identifier 400 performs processing to identify the allophone. Specifically, the identifier 400 uses the synthesized image 83 and an autoencoder to perform processing to identify allophones. More specifically, the identifier 400 performs processing to identify an abnormal sound based on the difference image obtained based on the composite image 83 and the output image output via the autoencoder.
[0052] 7A and 7B are diagrams illustrating synchronization between sound corresponding images 81. FIG. 7(A) and (B) show a simplified representation of each of the sound corresponding images 81. Also, in these Figures 7(A) and (B), each of the sound corresponding images 81 includes a vertical streak image 88 caused by the impulse sound generated in the image forming apparatus 100. When generating the composite image 83, the composite image generating unit 420 generates the composite image 83 in which a plurality of sound corresponding images 81 are arranged in a time-synchronized manner, as shown in FIG. 7(B).
[0053] FIG. 7A illustrates an example in which a composite image 83 is generated in a manner in which two sound corresponding images 81 are not synchronized in time. In contrast to this, FIG. 7B illustrates an example in which a composite image 83 is generated in such a manner that two sound corresponding images 81 are synchronized in time. In this embodiment, an impulse sound may be generated as a normal sound in the image forming apparatus 100. Each time a sheet of paper P is transported, the sheet of paper P may hit a member on the transport path, which may generate an impulse sound as a normal sound. In FIGS. 7A and 7B, each of the sound corresponding images 81 includes a vertical streak image 88 caused by this impulse sound.
[0054] A case will be described where two sound corresponding images 81 are arranged in a time-synchronized manner. In this case, as shown in Figure 7(B), when comparing positions in the time axis direction, the position of the image 88 corresponding to the impulse sound that appears in one sound corresponding image 81 matches the position of the image 88 corresponding to the impulse sound that appears in the other sound corresponding image 81.
[0055] Here, as shown in FIG. 7(A), it is assumed that two sound corresponding images 81 are arranged in a time-unsynchronized manner. In this case, even though the above impulse sounds are generated at the same timing, the position of the image 88 corresponding to the impulse sound that appears on one sound corresponding image 81 is shifted from the position of the image 88 corresponding to the impulse sound that appears on the other sound corresponding image 81. In this case, there is a risk that the accuracy of identifying abnormal noises occurring in image forming apparatus 100 may decrease.
[0056] If the sound corresponding images 81 are arranged in a manner that is not synchronized in time, there is a risk that the accuracy of identifying abnormal sounds will decrease due to interactions between the sound corresponding images 81. If the position of the image 88 corresponding to the impulse sound is shifted, there is a risk that learning by the autoencoder will not be performed correctly due to interactions between the sound corresponding images 81. In this case, there is a risk that the accuracy of identification of abnormal sounds by the identifier 400 will decrease. If the autoencoder is no longer able to learn correctly, sounds that are actually normal may be identified as abnormal sounds, and the accuracy of identifying abnormal sounds may decrease. In contrast to this, in this embodiment, the sound corresponding images 81 are arranged in a time-synchronized manner when generating the composite image 83. In this case, the problem of a decrease in the accuracy of identifying abnormal sounds is less likely to occur.
[0057] Here, if the two sound corresponding images 81 are not synchronized in time, for example, one of the sound corresponding images 81 is offset in the time axis direction, thereby synchronizing the two sound corresponding images 81 in time. More specifically, for example, due to differences in processing speed in the corresponding image generating unit 410 (see FIG. 2) that generates the sound corresponding image 81, it is conceivable that the two sound corresponding images 81 may not be synchronized in time. In this case, for example, the sound corresponding image 81 output from the faster processing corresponding image generating unit 410 is delayed, thereby synchronizing the two sound corresponding images 81 in time.
[0058] [Another example of synthesis processing] 8A and 8B are diagrams showing another example of processing for generating a composite image 83. In FIG. FIG. 8(A) shows the state before the composite image 83 is generated, and FIG. 8(B) shows the state after the composite image 83 is generated. 8, when generating the composite image 83, the composite image generation unit 420 positions the high-frequency side 81Y of one sound corresponding image 81E on the side of the other sound corresponding image 81F, as shown in Fig. 8(B). Also, the composite image generation unit 420 positions the high-frequency side 81Y of the other sound corresponding image 81F on the side of the one sound corresponding image 81E.
[0059] More specifically, in this example, the composite image generation unit 420 inverts one sound corresponding image 81E located at the top of Figure 8(A) upside down, and then generates a composite image 83 based on this inverted one sound corresponding image 81E and another sound corresponding image 81F located at the bottom of Figure 8(A). In this example, a composite image 83 is generated in such a manner that the high-frequency side 81Y of one sound corresponding image 81E is positioned on the side of the other sound corresponding image 81F, and the high-frequency side 81Y of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E.
[0060] In this case, in the composite image 83, at a location where the one sound corresponding image 81E and the other sound corresponding image 81F contact each other, a difference in density between the one sound corresponding image 81E and the other sound corresponding image 81F is unlikely to occur. 6(B), it is assumed that the low-frequency side 81X of one sound corresponding image 81E is located on the side of the other sound corresponding image 81F, and the high-frequency side 81Y of the other sound corresponding image 81F is located on the side of the one sound corresponding image 81E. In this case, a difference in density between the one sound corresponding image 81E and the other sound corresponding image 81F is likely to occur at the point where the one sound corresponding image 81E and the other sound corresponding image 81F contact each other. In contrast, this difference is less likely to occur when the processing shown in Fig. 8 is performed, and in this case, the decrease in accuracy in identifying abnormal noise is suppressed.
[0061] Assume that there is a large difference in density between one sound corresponding image 81E and another sound corresponding image 81F at a location where the two images meet. In this case, as in the case of the above-described impulse sound, there is a risk that the accuracy of identifying the allergic sound will decrease due to interactions between the sound corresponding images 81. More specifically, in this case, there is a risk that learning by the autoencoder will not be performed correctly, and accordingly, there is a risk that the accuracy of identifying the allergic sound will decrease. In contrast to this, as shown in FIG. 8(B), when the high-frequency side 81Y of one sound-corresponding image 81E is positioned on the side of another sound-corresponding image 81F, and the high-frequency side 81Y of the other sound-corresponding image 81F is positioned on the side of the one sound-corresponding image 81E, the accuracy of identifying abnormal sounds is less likely to decrease.
[0062] 9(A) and (B) are diagrams showing another processing example. 9A shows each of the sound corresponding images 81 etc. before synthesis, and FIG. 9B shows the synthesized image 83. In this processing example, similarly to the above, when generating a composite image 83, the high frequency side 81Y of one sound corresponding image 81E is positioned on the side of another sound corresponding image 81F as shown in FIG. 9(B). Moreover, the high frequency side 81Y of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E.
[0063] Furthermore, in this processing example, the composite image generating unit 420 places an intermediate image 78 between one sound corresponding image 81E and another sound corresponding image 81F, as shown in FIGS. 9(A) and 9(B). This intermediate image 78 is an image that is configured by an image other than the one sound corresponding image 81E and the other sound corresponding image 81F. In this processing example, the intermediate image 78 is arranged to further reduce the density difference at the boundary between one sound corresponding image 81 E and another sound corresponding image 81 F. Here, the density of the intermediate image 78 is uniform.
[0064] FIG. 10 is an enlarged view of the portion indicated by the symbol X in FIG. In FIG. 10, an intermediate image 78, a part of one sound corresponding image 81E, and a part of another sound corresponding image 81F are displayed. Here, consider the plurality of pixels 208 that make up one sound corresponding image 81E. These plurality of pixels 208 are the plurality of pixels that are aligned in the above-mentioned one direction 6A (see FIG. 6(B)), and are the plurality of pixels that contact the intermediate image 78. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of pixel 208A, which has the largest density value among the plurality of pixels 208. Furthermore, in this embodiment, the density value of the intermediate image 78 is larger than the density value of pixel 208B, which has the smallest density value among the plurality of pixels 208.
[0065] The plurality of pixels 208 includes pixels with large density values and pixels with small density values. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of pixel 208A, which has the largest density value among the plurality of pixels 208. Also, the density value of the intermediate image 78 is larger than the density value of pixel 208B, which has the smallest density value among the plurality of pixels 208. In this case, the shading at the boundary between one sound corresponding image 81E and another sound corresponding image 81F is smaller than when one sound corresponding image 81E and another sound corresponding image 81F are in direct contact with each other.
[0066] When one sound corresponding image 81E and another sound corresponding image 81F are in direct contact with each other, a pixel with a large density value included in one sound corresponding image 81E and a pixel with a small density value included in the other sound corresponding image 81F will be adjacent to each other. In this case, the difference in shading at the boundary between one sound corresponding image 81E and another sound corresponding image 81F becomes large, which may result in a decrease in the accuracy of identifying the abnormal sound. In contrast to this, providing the intermediate image 78 prevents the shading from becoming too large, and in this case, the accuracy of identifying abnormal noises is less likely to decrease.
[0067] In this embodiment, the other sound corresponding image 81F also has the same configuration. Here, consider a plurality of pixels 209 that constitute another sound corresponding image 81F. Specifically, consider a plurality of pixels 209 that are aligned in one direction 6A (see FIG. 6(B)) and that contact the intermediate image 78. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of pixel 209A, which has the largest density value among the plurality of pixels 209. Also, the density value of the intermediate image 78 is larger than the density value of pixel 209B, which has the smallest density value among the plurality of pixels 209.
[0068] The density of the intermediate image 78 is set in advance based on, for example, the density of one previously obtained sound corresponding image 81E and the density of another previously obtained sound corresponding image 81F. Specifically, the density of the intermediate image 78 is set based on, for example, a first average value which is the average value of the density of each of the above-mentioned multiple pixels 208 of one sound corresponding image 81E, and a second average value which is the average value of the density of each of the above-mentioned multiple pixels 209 of another sound corresponding image 81F. More specifically, for example, the average value of the first average value and the second average value is set as the density of the intermediate image 78.
[0069] Alternatively, for example, the density of the intermediate image 78 may be set based on a first average value, which is the average value of the density of each of the pixels that make up one sound-corresponding image 81E and that are included in a specific area indicated by reference symbol 10A in Figure 10, and a second average value, which is the average value of the density of each of the pixels that make up another sound-corresponding image 81F and that are included in a specific area indicated by reference symbol 10B. More specifically, in this case, the average value of the first average value and the second average value may be further calculated, and the calculated average value may be set as the density of the intermediate image 78.
[0070] 11(A) and (B) are diagrams showing another example of processing for generating a composite image 83. In FIG. FIG. 11(A) shows the state before the composite image 83 is generated, and FIG. 11(B) shows the state after the composite image 83 is generated. 11, when generating a composite image 83, the composite image generation unit 420 positions the low-frequency side edge 81X of one sound corresponding image 81E on the side of another sound corresponding image 81F, as shown in Fig. 11(B). Also, the composite image generation unit 420 positions the low-frequency side edge 81X of the other sound corresponding image 81F on the side of the one sound corresponding image 81E, as shown in Fig. 11(B). More specifically, the composite image generation unit 420 inverts the other sound corresponding image 81F located at the bottom of Fig. 11(A) upside down. Then, the composite image generation unit 420 generates the composite image 83 based on the other sound corresponding image 81F after the inversion and the one sound corresponding image 81E located at the top of the figure.
[0071] In this example, a composite image 83 is generated in such a way that the low-frequency side edge 81X of one sound corresponding image 81E is positioned on the side of the other sound corresponding image 81F, and the low-frequency side edge 81X of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E. In this case as well, at the location where one sound corresponding image 81E and another sound corresponding image 81F contact each other, a difference in density between one sound corresponding image 81E and another sound corresponding image 81F is unlikely to occur.
[0072] 12(A) and (B) are diagrams showing another processing example. FIG. 12(A) shows each of the sound corresponding images 81 etc. before synthesis, and FIG. 12(B) shows a synthesized image 83 obtained by synthesis. In this processing example, when generating a composite image 83, as shown in Fig. 12(B), the low-frequency side edge 81X of one sound corresponding image 81E is positioned on the side of another sound corresponding image 81F. Also, the low-frequency side edge 81X of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E.
[0073] Furthermore, in this processing example, when generating the composite image 83, the composite image generating unit 420 places the intermediate image 78 between one sound corresponding image 81E and another sound corresponding image 81F, in the same manner as described above. In this processing example as well, the shading at the boundary between one sound corresponding image 81E and another sound corresponding image 81F is reduced by this intermediate image 78. Here, the density of the intermediate image 78 is uniform, as in the above case.
[0074] FIG. 13 is an enlarged view of the portion indicated by reference numeral XIII in FIG. FIG. 13 also displays the intermediate image 78, a portion of one sound corresponding image 81E, and a portion of another sound corresponding image 81F. Here again, consider the plurality of pixels 212 that make up one sound corresponding image 81E. As in the above, the plurality of pixels 212 are a plurality of pixels aligned in the one direction 6A (see FIG. 6(B)) and are a plurality of pixels that contact the intermediate image 78. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of pixel 212A, which has the largest density value among the plurality of pixels 212. Furthermore, in this embodiment, the density value of the intermediate image 78 is larger than the density value of pixel 212B, which has the smallest density value among the plurality of pixels 212.
[0075] As in the above, the plurality of pixels 212 includes pixels with large density values and pixels with small density values. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of the pixel 212A, which has the largest density value among the plurality of pixels. In this embodiment, the density value of the intermediate image 78 is greater than the density value of the pixel 212B, which has the smallest density value among the plurality of pixels. In this case, similarly to the above, the shading at the boundary between one sound corresponding image 81E and another sound corresponding image 81F is smaller than when one sound corresponding image 81E and another sound corresponding image 81F are in direct contact with each other.
[0076] The same is true for the other sound-enabled images 81F. Here, consider a plurality of pixels 213 that constitute another sound corresponding image 81F. Specifically, consider a plurality of pixels 213 that are aligned in one direction 6A (see FIG. 6(B)) and that contact the intermediate image 78. In this embodiment, the density value of the intermediate image 78 is smaller than the density value of pixel 213A, which has the largest density value among the plurality of pixels 213. In this embodiment, the density value of the intermediate image 78 is also larger than the density value of pixel 213B, which has the smallest density value among the plurality of pixels 213.
[0077] As described above, the density of the intermediate image 78 is set in advance based on, for example, the density of one previously obtained sound corresponding image 81E and the density of another previously obtained sound corresponding image 81F. Specifically, the density of the intermediate image 78 is set based on, for example, a first average value which is the average value of the density of each of the above-mentioned multiple pixels 212 of one sound-corresponding image 81E, and a second average value which is the average value of the density of each of the above-mentioned multiple pixels 213 of another sound-corresponding image 81F. More specifically, for example, the average value of the first average value and the second average value is set as the density of the intermediate image 78.
[0078] Alternatively, for example, the density of the intermediate image 78 may be set based on a first average value, which is the average value of the density of each of the pixels that make up one sound-corresponding image 81E and that are included in a specific area indicated by the symbol 13A, and a second average value, which is the average value of the density of each of the pixels that make up another sound-corresponding image 81F and that are included in a specific area indicated by the symbol 13B. More specifically, in this case, the average value of the first average value and the second average value may be further calculated, and the calculated average value may be set as the density of the intermediate image 78.
[0079] Although not shown, in the composite image 83 shown in FIG. 6, an intermediate image 78 may also be placed between one sound corresponding image 81E and another sound corresponding image 81F that make up this composite image 83. In the example shown in FIG. 6, when the intermediate image 78 is arranged, it is preferable that the density of the intermediate image 78 gradually increases from one sound corresponding image 81E toward the other sound corresponding image 81F. In the example shown in FIG. 6, the low-frequency side 81X of one sound corresponding image 81E is located on another sound corresponding image 81F, and the high-frequency side 81Y of the other sound corresponding image 81F is located on the one sound corresponding image 81E side. In this case, it is preferable that the density of the intermediate image 78 gradually increases from the one sound corresponding image 81E side toward the other sound corresponding image 81F side.
[0080] In the above description, the composite image 83 is generated based on the two sound corresponding images 81, but the present invention is not limited to this. For example, as shown in FIG. 14 (a diagram showing another example of a composite image), a composite image 83 may be generated based on three or more sound corresponding images 81. In the example shown in FIG. 14, one composite image 83 is formed based on four sound corresponding images 81 generated in response to four provided sound sensors 120.
[0081] In this example shown in Figure 14, as described above, between two adjacent sound-corresponding images 81 in the cross direction 6B, which is the direction that crosses the one direction 6A, the low-frequency side edges 81X are in contact with each other, or the high-frequency side edges 81Y are in contact with each other. Additionally, in the example shown in FIG. 14, an intermediate image 78 may be placed between two adjacent sound corresponding images 81, similar to the forms shown in FIGS.
[0082] (Addendum) (((1))) an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect sounds from the image forming unit; an identifying unit that identifies an abnormality in the image forming unit based on information output from the plurality of sound detecting units; Equipped with The image forming apparatus is provided with a smaller number of identifying units than the number of the plurality of sound detecting units. (((2))) The image forming apparatus according to (((1))) further comprises a transmitting unit that, when an abnormality is identified by the identifying unit, transmits information indicating that an abnormality has occurred to an external device. (((3))) The identification unit performs machine learning using images generated based on information obtained by the sound detection unit and generated for each sound detection unit, calculates latent variables from features extracted from the images using a machine-learned anomaly detection model, generates an output image by restoring the latent variables, and identifies an abnormality in the image forming unit by comparing an input image with the output image. (((1))) or (((2))) (((4))) The image processing device further includes a generating unit that generates an image based on information obtained by the sound detection unit and for each of the sound detection units, and generates a composite image by synthesizing images based on a plurality of the images generated in accordance with the plurality of sound detection units, and generating a composite image that is an image obtained by the synthesis, The identification unit identifies an abnormality based on the composite image generated by the generation means. The image forming apparatus according to any one of (((1))) to (((3))). (((5))) Each of the images generated for each of the sound detection units has a time axis and a frequency axis, and is an image in which sound intensity is expressed by pixel values, the generating means generates the composite image in which each of the plurality of images is arranged such that the time axis of each of the images is aligned along a specific direction. The image forming apparatus according to (((4))). (((6))) the generation means generates the composite image by synthesizing images based on the plurality of images, in which each of the plurality of images is arranged so that the time axis of each of the plurality of images is aligned with the one direction, and the plurality of images are arranged side by side in a direction intersecting the one direction; The image forming apparatus according to (((5))). (((7))) the generating means generates the composite image in which the plurality of images are arranged in a time-synchronized manner when synthesizing images based on the plurality of images to generate the composite image; The image forming apparatus according to (((5))). (((8))) the frequency axis of the image corresponds to the magnitude of a frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends, the second side being positioned differently from the first side in the direction in which the frequency axis extends; one of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where low frequency component values of the image are displayed, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where high frequency component values of the image are displayed; the generation means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, such that the high-frequency side edge of the one image is located on the side of the other image, and the high-frequency side edge of the other image is located on the side of the one image. The image forming apparatus according to (((6))). (((9))) the generating means, when generating the composite image by synthesizing images using at least the one image and the other image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image; The image forming apparatus according to (((8))). (((10))) a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the one image and are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction, a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the other image and that are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction; The image forming apparatus according to (((9))). (((11))) the frequency axis of the image corresponds to the magnitude of a frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends, the second side being positioned differently from the first side in the direction in which the frequency axis extends; one of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where low frequency component values of the image are displayed, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where high frequency component values of the image are displayed; the generating means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, in which the low-frequency side edge of the one image is located on the other image side and the low-frequency side edge of the other image is located on the one image side; The image forming apparatus according to (((6))). (((12))) the generating means, when generating the composite image by synthesizing images using at least the one image and the other image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image; The image forming apparatus according to (((11))). (((13))) a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the one image and are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction, a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the other image and that are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction; The image forming apparatus according to (((12))). (((14))) an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect sounds from the image forming unit; a generating means for generating a composite image by synthesizing images based on the images generated for each sound detection unit based on information obtained by the sound detection unit, the images being generated in plurality according to the plurality of sound detection units provided; and an identifying unit that identifies an abnormality in the image forming unit based on the composite image; An image forming apparatus comprising:
[0083] According to the image forming device of (((1))), it is possible to reduce the amount of hardware required to identify abnormalities compared to when a functional unit that identifies abnormalities individually is provided in correspondence with each of the multiple sound detection units. According to the image forming apparatus of (((2))), the amount of information sent to the external device can be reduced compared to when information about the abnormality is sent to the external device regardless of whether the abnormality is identified or not. According to the image forming apparatus of (((3))), even in a situation where only normal sounds are basically generated in the image forming apparatus and abnormal sounds are unlikely to occur, it becomes possible to identify abnormalities occurring in the image forming apparatus. According to the image forming apparatus of (((4))), the processing load required to identify an abnormality can be reduced compared to when identifying an abnormality for each image obtained by each sound detection unit. According to the image forming device of (((5))), the accuracy of identifying abnormalities can be improved compared to when the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends. According to the image forming device of (((6))), the accuracy of identifying abnormalities can be improved compared to when the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of the other images extends. According to the image forming apparatus of (((7))), it is possible to improve the accuracy of identifying abnormalities compared to when a plurality of images are not arranged in a time-synchronized manner. According to the image forming device of (((8))), the accuracy of identifying abnormalities can be improved compared to when the low-frequency side of one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image. According to the image forming apparatus of (((9))), it is possible to improve the accuracy of identifying abnormalities compared to when one image and another image are in direct contact with each other. According to the image forming apparatus of (((10))), it is possible to improve the accuracy of identifying abnormalities compared to when one image and another image are in direct contact with each other. According to the image forming device of (((11))), the accuracy of identifying abnormalities can be improved compared to when the low-frequency side of one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image. According to the image forming apparatus of (((12))), it is possible to improve the accuracy of identifying abnormalities compared to when one image and another image are in direct contact with each other. According to the image forming apparatus of (((13))), it is possible to improve the accuracy of identifying abnormalities compared to when one image and another image are in direct contact with each other. According to the image forming apparatus of (((14))), it is possible to reduce the amount of hardware required to identify an abnormality compared to when an abnormality is identified without generating a composite image. [Explanation of symbols]
[0084] 78...intermediate image, 81A...horizontal axis, 81B...vertical axis, 81E...one sound corresponding image, 81F...other sound corresponding image, 81X...low frequency side, 81Y...high frequency side, 83...composite image, 100...image forming device, 100A...image forming unit, 120...sound sensor, 200...server device, 400...specifier, 420...composite image generating unit, 430...transmitting unit
Claims
1. an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect sounds from the image forming unit; an identifying unit that identifies an abnormality in the image forming unit based on information output from the plurality of sound detecting units; Equipped with The image forming apparatus is provided with a smaller number of identifying units than the number of the plurality of sound detecting units.
2. The image forming apparatus according to claim 1 , further comprising a transmitting unit that, when the identifying unit identifies an abnormality, transmits information indicating that an abnormality has occurred to an external device.
3. 2. The image forming apparatus according to claim 1, wherein the identification unit performs machine learning using images generated based on information obtained by the sound detection unit and generated for each sound detection unit, calculates latent variables from features extracted from the images using a machine-learned anomaly detection model, generates an output image by restoring the output image using the latent variables, and identifies an abnormality in the image forming unit by comparing an input image with the output image.
4. The image processing device further includes a generating unit that generates an image based on information obtained by the sound detection unit and for each of the sound detection units, and generates a composite image by synthesizing images based on a plurality of the images generated in accordance with the plurality of sound detection units, and generating a composite image that is an image obtained by the synthesis, The identification unit identifies an abnormality based on the composite image generated by the generation means. The image forming apparatus according to claim 1 .
5. Each of the images generated for each of the sound detection units has a time axis and a frequency axis, and is an image in which sound intensity is expressed by pixel values, the generating means generates the composite image in which each of the plurality of images is arranged such that the time axis of each of the images is aligned along a specific direction. The image forming apparatus according to claim 4 .
6. the generation means generates the composite image by synthesizing images based on the plurality of images, in which each of the plurality of images is arranged so that the time axis of each of the plurality of images is aligned with the one direction, and the plurality of images are arranged side by side in a direction intersecting the one direction; The image forming apparatus according to claim 5 .
7. the generation means generates the composite image by synthesizing images based on the plurality of images, in which the plurality of images are arranged in a time-synchronized manner; The image forming apparatus according to claim 5 .
8. the frequency axis of the image corresponds to the magnitude of a frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends, the second side being positioned differently from the first side in the direction in which the frequency axis extends; one of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where low frequency component values of the image are displayed, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where high frequency component values of the image are displayed, the generation means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, such that the high-frequency side edge of the one image is located on the side of the other image, and the high-frequency side edge of the other image is located on the side of the one image. The image forming apparatus according to claim 6 .
9. the generating means, when generating the composite image by synthesizing images using at least the one image and the other image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image; The image forming apparatus according to claim 8 .
10. a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the one image and are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction, a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the other image and that are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction; The image forming apparatus according to claim 9 .
11. the frequency axis of the image corresponds to the magnitude of a frequency component value obtained by analyzing the information obtained by the sound detection unit, each of the images has a first side along the direction in which the time axis extends and a second side along the direction in which the time axis extends, the second side being positioned differently from the first side in the direction in which the frequency axis extends; one of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where low frequency component values of the image are displayed, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where high frequency component values of the image are displayed, the generating means generates the composite image by synthesizing images using at least one image and another image included in the plurality of images, in which the low-frequency side edge of the one image is located on the other image side and the low-frequency side edge of the other image is located on the one image side; The image forming apparatus according to claim 6 .
12. the generating means, when generating the composite image by synthesizing images using at least the one image and the other image, places an intermediate image, which is an image other than the one image and the other image, between the one image and the other image; The image forming apparatus according to claim 11.
13. a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the one image and are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction, a density value of the intermediate image is smaller than a density value of a pixel having a largest density value among a plurality of pixels that are arranged in the one direction among a plurality of pixels that constitute the other image and that are in contact with the intermediate image, and a density value of the intermediate image is larger than a density value of a pixel having a smallest density value among a plurality of pixels that are arranged in the one direction; The image forming apparatus according to claim 12.
14. an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect sounds from the image forming unit; a generating means for generating a composite image by synthesizing images based on the images generated for each sound detection unit based on information obtained by the sound detection unit, the images being generated in plurality according to the plurality of sound detection units provided; and an identifying unit that identifies an abnormality in the image forming unit based on the composite image; An image forming apparatus comprising:
Citation Information
Patent Citations
Image forming apparatus with self-checking function
JP2006184722A
Anomaly detection system, anomaly detection method, anomaly detection program, and trained model generation method
JP6740247B2