Image forming apparatus

By using a small number of determination units in an image forming apparatus to generate a synthetic image for abnormality detection, the problems of increased hardware and processing load caused by multiple sound detection units are solved, and efficient abnormality detection is achieved.

CN120722701APending Publication Date: 2025-09-30FUJIFILM BUSINESS INNOVATION CORP
1 Cites 0 Cited by

Patent Information

Application Number
CN202411202073.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2024-08-29
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the prior art, abnormality detection of an image forming apparatus requires multiple sound detection units to be independently provided with corresponding functional units for determining abnormalities, resulting in increased hardware requirements and excessive processing load.

Method used

A small number of determination units are used to generate a synthetic image through multiple sound detection units, and anomaly detection is performed using an autoencoder, reducing the amount of hardware and improving detection accuracy through synthetic images.

Benefits of technology

It reduces the hardware requirements for determining anomalies, reduces the amount of information sent, and improves the accuracy and processing efficiency of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120722701A_ABST
    Figure CN120722701A_ABST
Patent Text Reader

Abstract

An image forming apparatus includes: an image forming section that forms an image on a recording medium; a plurality of sound detection units that detect sounds from the image forming unit; and determining sections that determine an abnormality of the image forming section on the basis of information output from the plurality of sound detection sections, the number of the determining sections being set to be less than the number of the plurality of sound detection sections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image forming apparatus. Background Art

[0002] Patent document 1 discloses an anomaly detection system that is capable of operating by performing the following processing: inferring latent variables from input data of an anomaly detection object based on an encoder of a VAE pre-learned using training data containing normal data; generating restored data from the latent variables based on a decoder of the pre-learned VAE; and determining whether the input data is normal or abnormal based on the input data and the restored data.

[0003] Patent Document 2 discloses a process in which a plurality of sound sensors for detecting the operating sounds of a device are provided, and the previously collected operating sounds of the device are compared with newly collected operating sounds to determine the abnormal state and timing of the device.

[0004] Patent Document 1: Japanese Patent No. 6740247

[0005] Patent Document 2: Japanese Patent Application Laid-Open No. 2006-184722 Summary of the Invention

[0006] If a sound detection unit is provided in the image forming apparatus and a functional unit for identifying an abnormality in the image forming apparatus is provided in a manner corresponding to the sound detection unit, an abnormality in the image forming apparatus can be identified.

[0007] Here, if the functional unit for identifying abnormality is provided so as to correspond to each of the plurality of sound detection units, abnormality is identified individually for each of the functional units.

[0008] An object of the present invention is to reduce the hardware required for abnormality identification, compared with a case where a functional unit for identifying abnormality is separately provided corresponding to each of a plurality of sound detection units.

[0009] The invention described in Option 1 is an image forming device, which comprises: an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect the sound of the image forming unit; and a determination unit that determines an abnormality of the image forming unit based on information output from the plurality of sound detection units, and the number of the determination units is set to be less than the number of the plurality of sound detection units.

[0010] The invention according to claim 2 is the image forming apparatus according to claim 1, further comprising a transmitting unit configured to transmit information indicating that an abnormality has occurred to an external device when the abnormality has been identified by the identifying unit.

[0011] The invention described in Option 3 is an image forming device described in Option 1 or Option 2, wherein the determination unit uses machine learning for images generated based on information obtained by the sound detection unit and for each image generated by the sound detection unit, uses an abnormality detection model after machine learning, calculates latent variables based on feature quantities extracted from the image, and uses the latent variables to perform recovery to generate an output image, and determines the abnormality of the image forming unit by comparing the input image and the output image.

[0012] The invention described in Option 4 is an image forming device described in any one of Options 1 to 3, which further includes a generating component, which synthesizes an image based on multiple images to generate an image obtained by the synthesis, namely a composite image, wherein the image is generated based on the information obtained by the sound detection unit and is generated for each sound detection unit, and is a plurality of such images generated corresponding to the plurality of sound detection units, and the determination unit determines the abnormality based on the composite image generated by the generating component.

[0013] The invention described in Option 5 is the image forming device described in Option 4, wherein each of the images generated by the sound detection unit is an image having a time axis and a frequency axis and expressing the intensity of the sound using pixel values, and the generating component generates the composite image in which the images are respectively arranged in the form of a plurality of images whose respective time axes are along a specific direction.

[0014] The invention described in Option 6 is the image forming device described in Option 5, wherein, when synthesizing an image based on a plurality of the images to generate the synthesized image, the generating component generates the synthesized image as follows: the images are respectively arranged in the form of a plurality of the images with their respective time axes along the one direction and the plurality of the images are arranged in the form of being arranged in a direction intersecting the one direction.

[0015] The invention described in claim 7 is the image forming apparatus described in claim 5, wherein, when the composite image is generated by synthesizing the images based on the plurality of images, the generating means generates the composite image in which the plurality of images are arranged in a temporally synchronized manner.

[0016] The invention described in Scheme 8 is an image forming device described in Scheme 6, wherein the frequency axis of the image corresponds to the size of the frequency component value obtained by analyzing the information obtained by the sound detection unit, and each of the images has a first side along the direction extending along the time axis, and a second side along the direction extending along the time axis and whose position in the direction extending along the frequency axis is different from that of the first side, one of the first side and the second side is a low-frequency side located on the low-frequency side and is arranged on the side of the image displaying the low-frequency component value, and the other side is a high-frequency side located on the high-frequency side and is arranged on the side of the image displaying the high-frequency component value, and when the composite image is generated by synthesizing the image using at least one image and another image included in a plurality of the images, the generating component generates a composite image in which the high-frequency side of the one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image.

[0017] The invention described in Scheme 9 is an image forming device described in Scheme 8, wherein, when the composite image is generated by synthesizing the image using at least the one image and the other image, the generating component configures an image other than the one image and the other image, i.e., an intermediate image, between the one image and the other image.

[0018] The invention described in Option 10 is the image forming device described in Option 9, wherein the concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the one image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction, and the concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction.

[0019] The invention described in Scheme 11 is an image forming device described in Scheme 6, wherein the frequency axis of the image corresponds to the size of the frequency component value obtained by analyzing the information obtained by the sound detection unit, and each of the images has a first side along the direction extending along the time axis, and a second side along the direction extending along the time axis and whose position in the direction extending along the frequency axis is different from that of the first side, one of the first side and the second side is a low-frequency side located on the low-frequency side and is arranged on the side of the image displaying the low-frequency component value, and the other side is a high-frequency side located on the high-frequency side and is arranged on the side of the image displaying the high-frequency component value, and when the composite image is generated by synthesizing the image using at least one image and another image included in a plurality of the images, the generating component generates a composite image in which the low-frequency side of the one image is located on the side of the other image and the low-frequency side of the other image is located on the side of the one image.

[0020] The invention described in Scheme 12 is an image forming device described in Scheme 11, wherein, when the composite image is generated by synthesizing the image using at least the one image and the other image, the generating component configures an image other than the one image and the other image, i.e., an intermediate image, between the one image and the other image.

[0021] The invention described in Option 13 is the image forming device described in Option 12, wherein the concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the one image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction, and the concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction.

[0022] The invention described in Scheme 14 is an image forming device, which comprises: an image forming unit that forms an image on a recording medium; a plurality of sound detection units that detect the sound of the image forming unit; a generating component that synthesizes images based on the images to generate an image obtained by the synthesis, namely a composite image, wherein the image is generated based on the information obtained by the sound detection unit and is generated for each of the sound detection units, and a plurality of such images are generated corresponding to the plurality of sound detection units provided; and a determining unit that determines the abnormality of the image forming unit based on the composite image.

[0023] Effects of the Invention

[0024] According to the first aspect of the present invention, compared with a case where a functional unit for identifying abnormality is separately provided corresponding to each of a plurality of sound detection units, hardware required for identifying abnormality can be reduced.

[0025] According to the second aspect of the present invention, the amount of information transmitted to the external device can be reduced compared to a case where information related to an abnormality is transmitted to the external device regardless of whether an abnormality is confirmed or not.

[0026] According to the third aspect of the present invention, even when only normal sounds are basically generated in the image forming apparatus and abnormal noise is unlikely to be generated, an abnormality occurring in the image forming apparatus can be identified.

[0027] According to the fourth aspect of the present invention, the load of processing required for identifying an abnormality can be reduced compared to a case where an abnormality is identified for each image obtained by each sound detection unit.

[0028] According to the fifth aspect of the present invention, the accuracy of identifying an abnormality can be improved compared to a case where the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends.

[0029] According to the sixth aspect of the present invention, the accuracy of identifying an abnormality can be improved compared to a case where the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends.

[0030] According to the seventh aspect of the present invention, the accuracy of identifying abnormalities can be improved compared to a case where a plurality of images are not arranged in a temporally synchronized manner.

[0031] According to the eighth aspect of the present invention, the accuracy of abnormality identification can be improved compared to a case where the low-frequency side of one image is located on the other image side and the high-frequency side of the other image is located on the one image side.

[0032] According to the ninth aspect of the present invention, the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0033] According to the tenth aspect of the present invention, the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0034] According to the eleventh aspect of the present invention, the accuracy of abnormality identification can be improved compared to a case where the low-frequency side of one image is located on the other image side and the high-frequency side of the other image is located on the one image side.

[0035] According to the twelfth aspect of the present invention, the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0036] According to the thirteenth aspect of the present invention, the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0037] According to the fourteenth aspect of the present invention, the hardware required for determining an abnormality can be reduced compared to a case where an abnormality is determined without generating a synthetic image. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Embodiments of the present invention will be described in detail with reference to the following drawings.

[0039] Figure 1 is a diagram showing an example of a diagnostic system;

[0040] Figure 2 is a diagram illustrating an image forming apparatus;

[0041] Figure 3 is a diagram showing an example of the hardware configuration of a determiner;

[0042] Figure 4 is a diagram illustrating processing performed by a determiner;

[0043] Figure 5 is a diagram showing a sound corresponding image generated by a corresponding image generating unit;

[0044] exist Figure 6 middle, Figure 6 (A) Figure 6 (B) is a diagram illustrating a process of generating a synthetic image by a synthetic image generating unit;

[0045] exist Figure 7 middle, Figure 7 (A) Figure 7 (B) is a diagram illustrating synchronization between sound and image;

[0046] exist Figure 8 middle, Figure 8 (A) Figure 8 (B) is a diagram showing another processing example for generating a composite image;

[0047] exist Figure 9 middle, Figure 9 (A) Figure 9 (B) is a diagram showing another processing example;

[0048] Figure 10 is Figure 9 An enlarged view of the portion indicated by symbol X;

[0049] exist Figure 11 middle, Figure 11 (A) Figure 11 (B) is a diagram showing another processing example of generating a composite image;

[0050] exist Figure 12 middle, Figure 12 (A) Figure 12 (B) is a diagram showing another processing example;

[0051] Figure 13 is Figure 12 An enlarged view of the portion indicated by symbol XIII;

[0052] Figure 14 This is a diagram showing another example of a synthesized image.

[0053] Explanation of symbols

[0054] 78-middle image, 81A-horizontal axis, 81B-vertical axis, 81E-one sound corresponding image, 81F-another sound corresponding image, 81X-low-frequency side, 81Y-high-frequency side, 83-synthetic image, 100-image forming device, 100A-image forming unit, 120-sound sensor, 200-server device, 400-determinator, 420-synthetic image generating unit, 430-sending unit. DETAILED DESCRIPTION

[0055] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0056] Figure 1 1 is a diagram showing an example of the diagnostic system 1 .

[0057] The diagnostic system 1 of the present embodiment includes a plurality of image forming apparatuses 100 and a server device 200 connected to each of the plurality of image forming apparatuses 100 via a communication line 190 .

[0058] In addition, Figure 1 1 shows one image forming apparatus 100 among the plurality of image forming apparatuses 100 .

[0059] In the present embodiment, the server device 200 , which is an example of an information processing system, acquires information about each image forming apparatus 100 .

[0060] The diagnostic system 1 is further provided with a user terminal 300. The user terminal 300 is connected to the server device 200. The user terminal 300 receives operations from a user. An example of a user is a person who maintains the image forming device 100. In this embodiment, the user terminal 300 is provided for reference by the person who maintains the image forming device 100.

[0061] The user terminal 300 is provided with a display device 310. The user terminal 300 is implemented by a computer. Examples of the user terminal 300 include a PC (Personal Computer), a smartphone, and a tablet terminal.

[0062] The image forming apparatus 100 is provided with an image forming unit 100A that forms an image on a sheet as an example of a recording medium.

[0063] Moreover, in Figure 1 Although not shown in the figure, the image forming apparatus 100 is provided with an acoustic sensor and the like.

[0064] Figure 2 It is a diagram for explaining the image forming apparatus 100 .

[0065] In this embodiment, as described above, the image forming apparatus 100 is provided with the image forming unit 100A that forms an image on a sheet of paper P, which is an example of a recording medium. The image forming unit 100A forms an image using electrophotography.

[0066] The image forming portion 100A, which is an example of an image forming member, is provided with an intermediate transfer belt 108 as a member that circulates and a plurality of image forming units 107 that form images of different colors.

[0067] In the present embodiment, the images formed by each of the plurality of image forming units 107 are once transferred onto the intermediate transfer belt 108 and then transferred onto the paper P.

[0068] The plurality of image forming units 107 form images of different colors on the intermediate transfer belt 108. The intermediate transfer belt 108 is not essential. An image may be directly transferred from each of the plurality of image forming units 107 to the paper P.

[0069] Furthermore, a plurality of image forming units 107 is not essential, and a configuration may be provided with only one image forming unit 107. In the case of a configuration with only one image forming unit 107, the intermediate transfer belt 108 is omitted.

[0070] In this embodiment, the image forming units 107 include an image forming unit 107Y that forms a yellow image, an image forming unit 107M that forms a magenta image, an image forming unit 107C that forms a cyan image, and an image forming unit 107K that forms a black image.

[0071] The images formed by the image forming units 107 are transferred onto an intermediate transfer belt 108 , which is an example of a transfer member.

[0072] Then, the image is transferred to the paper P conveyed to the transfer section T at the transfer section T. Thus, the image is formed on the paper P.

[0073] A photosensitive drum 101 as an example of an image holding member is provided in each of the image forming units 107. The photosensitive drum 101 rotates in the clockwise direction.

[0074] Furthermore, each of the image forming units 107 is provided with a charging device 101C for charging the photosensitive drum 101. Each of the image forming units 107 is provided with an exposure device 102 for exposing the photosensitive drum 101.

[0075] Furthermore, each of the image forming units 107 is provided with a developing device 103 . The developing device 103 develops the electrostatic latent image formed on the photosensitive drum 101 by the exposure by the exposure device 102 .

[0076] The developing device 103 includes a developing roller 103A disposed at a position facing the photosensitive drum 101. In this embodiment, the developer adhering to the outer peripheral surface of the developing roller 103A moves toward the surface of the photosensitive drum 101. This causes development to be performed.

[0077] When the development is performed, an image composed of, for example, toner is formed on the photosensitive drum 101. The image is then transferred to the outer peripheral surface of the intermediate transfer belt 108. The image on the intermediate transfer belt 108 is then transferred to the paper P, forming an image on the paper P.

[0078] Furthermore, the formation of an image on the paper P by the image forming unit 100A is not limited to the electrophotographic method, and may be performed using other methods such as an inkjet method.

[0079] The image forming apparatus 100 is further provided with an image reading device 130 .

[0080] The image reading device 130 , which is an example of image reading means, is a so-called scanner that reads an image formed on a paper (not shown) which is an example of a recording medium.

[0081] The image reading device 130 includes a light source that emits light to illuminate the paper and a light receiving unit such as a CCD that receives light reflected from the paper. In this embodiment, the image data of the paper is generated based on the reflected light received by the light receiving unit.

[0082] Furthermore, each image forming apparatus 100 is provided with an operation receiving unit 132 for receiving an operation from a user using the image forming apparatus 100 .

[0083] The operation receiving unit 132 is formed of a so-called touch panel. The operation receiving unit 132 displays information to the user and receives operations performed by the user.

[0084] Furthermore, the display of information to the user and the reception of the user's operation are not limited to being performed by a single operation reception unit 132 , and the operation reception unit 132 and the information display unit may be provided separately.

[0085] Furthermore, each of the image forming apparatuses 100 is provided with a sound sensor 120 as an example of a detection unit for detecting sound from the image forming unit 100A. The sound sensor 120 may also be called a microphone.

[0086] In this embodiment, a plurality of sound sensors 120 are provided in each of the image forming apparatuses 100. Figure 2 In the illustrated example, two acoustic sensors 120 are provided in the image forming apparatus 100 .

[0087] Furthermore, in the present embodiment, a determiner 400 is provided as an example of a determining unit that determines an abnormality in the image forming unit 100A based on information output from the acoustic sensor 120 .

[0088] In the present embodiment, the number of determiners 400 is provided to be smaller than the number of the plurality of sound sensors 120. Specifically, in the present embodiment, one determiner 400 is provided.

[0089] Determiner 400 determines an abnormality in image forming unit 100A based on information output from multiple sound sensors 120. Determiner 400 determines abnormal noise generated in image forming apparatus 100 and determines an abnormality in image forming unit 100A based on the abnormal noise. Here, "abnormal noise" refers to a sound generated due to a malfunction in image forming apparatus 100.

[0090] Furthermore, a corresponding image generating unit 410 is provided so as to correspond to each of the acoustic sensors 120 .

[0091] Each corresponding image generating unit 410 images the information output from the acoustic sensor 120 and generates an image based on the information output from the acoustic sensor 120. In other words, each corresponding image generating unit 410 generates an image representing the sound acquired by the acoustic sensor 120.

[0092] Hereinafter, in this specification, the image generated by the corresponding image generating unit 410 is referred to as a “sound corresponding image”.

[0093] The corresponding image generating units 410 are provided so as to correspond to the respective acoustic sensors 120. Two corresponding image generating units 410 are provided.

[0094] In this embodiment, the correspondence image generating unit 410 is provided for each sound sensor 120 , and a sound correspondence image is generated for each sound sensor 120 .

[0095] Furthermore, in this embodiment, a synthetic image generating unit 420 for generating a synthetic image is provided.

[0096] The composite image generating unit 420 , which is an example of generating means, generates a composite image by synthesizing images based on a plurality of generated sound-corresponding images (details will be described later).

[0097] In the present embodiment, the synthetic image generating unit 420 generates synthetic images whose number is smaller than the number of audio-corresponding images.

[0098] In the present embodiment, the corresponding image generation unit 410 and the synthesized image generation unit 420 are implemented by, for example, a computer (not shown).

[0099] More specifically, the corresponding image generating unit 410 and the synthesized image generating unit 420 are implemented by a CPU (not shown) as an example of a processor executing a program related to image processing stored in a ROM or the like.

[0100] When the composite image is generated by the composite image generating unit 420 , the composite image is input to the determiner 400 .

[0101] Then, in this embodiment, the generated composite image is analyzed by the determiner 400 as an example of determining means. Then, when abnormal noise occurs in the image forming apparatus 100, the determiner 400 determines the abnormal noise.

[0102] In the present embodiment, abnormal noise generated in the image forming apparatus 100 is determined by the determiner 400. Then, at this time, the determiner 400 determines that the image forming portion 100A is abnormal.

[0103] Figure 3 4 is a diagram showing a hardware configuration example of the determiner 400 .

[0104] The determiner 400 is implemented by a computer and includes a calculation processing unit 11 that performs digital calculation processing according to a program and a secondary storage unit 12 that stores information.

[0105] The secondary storage unit 12 is realized by, for example, an existing information storage device such as an HDD (Hard Disk Drive), a semiconductor memory, or a magnetic tape.

[0106] The calculation processing unit 11 includes a CPU 11 a as an example of a processor.

[0107] Furthermore, the calculation processing unit 11 is provided with a RAM 11 b used as a working memory or the like of the CPU 11 a , and a ROM 11 c storing programs and the like executed by the CPU 11 a .

[0108] The arithmetic processing unit 11 is provided with a nonvolatile memory 11d and an interface unit 11e for controlling various units connected to the arithmetic processing unit 11, such as a communication unit. The nonvolatile memory 11d is configured to be rewritable and can retain data even when power is interrupted.

[0109] The nonvolatile memory 11d is composed of, for example, a battery-backed SRAM or flash memory, etc. The secondary storage unit 12 stores not only files, but also programs executed by the arithmetic processing unit 11 .

[0110] In the present embodiment, each process is executed by the CPU 11 a reading a program stored in the ROM 11 c or the secondary storage unit 12 .

[0111] The program executed by the CPU 11a can be provided to the determiner 400 in a state stored in a computer-readable recording medium such as a magnetic recording medium (a magnetic tape, a magnetic disk, etc.), an optical recording medium (an optical disk, etc.), a magneto-optical recording medium, or a semiconductor memory. Furthermore, the program executed by the CPU 11a can be provided to the determiner 400 using a communication method such as the Internet.

[0112] In this specification, the processor refers to a processor in a broad sense, and includes general-purpose processors (such as CPU: Central Processing Unit, etc.), special-purpose processors (such as GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic devices, etc.).

[0113] Furthermore, the operations of the processors may be performed not only by a single processor but also by a plurality of processors located in physically separate locations in collaboration. Furthermore, the order of the operations of the processors is not limited to that described in this embodiment and may be changed.

[0114] In this embodiment, the determiner 400 (refer to Figure 2 ) uses an autoencoder to identify abnormal noise. The autoencoder performs unsupervised learning using the same data in both the input and output layers.

[0115] The determiner 400 determines abnormal noise using the above-mentioned sound corresponding image generated from the information obtained by the sound sensor 120 and the autoencoder.

[0116] More specifically, the determiner 400 determines abnormal noise using the synthesized image generated by the synthesized image generation unit 420 based on the two sound corresponding images and the autoencoder.

[0117] Figure 4 4 is a diagram illustrating the processing performed by the determiner 400 .

[0118] As described above, the determiner 400 determines abnormal noise using an autoencoder.

[0119] In this embodiment, if Figure 4 As shown, the composite image generating unit 420 (reference Figure 2 )The composite image generated is input to the determiner 400.

[0120] Sounds that are usually generated in the image forming apparatus 100 are basically normal sounds.

[0121] In this embodiment, generally, a synthetic image generated based on normal sound is input to the determiner 400 .

[0122] In the determiner 400 using the autoencoder, learning is performed using a synthetic image generated from normal sound. In other words, in this embodiment, the autoencoder is basically learned using training data consisting of synthetic images generated from normal sound.

[0123] In the determiner 400 of this embodiment, learning is basically performed to make the input composite image and the output composite image consistent.

[0124] At this time, no difference is generated between the synthesized image before being input to the determiner 400 and the synthesized image output from the determiner 400 .

[0125] Hereinafter, in this specification, a synthesized image before being input to the determiner 400 may be referred to as an “input image,” and a synthesized image output from the determiner 400 may be referred to as an “output image.”

[0126] In this embodiment, when determining abnormal noise, the determiner 400 generates a difference image from the input image and the output image.

[0127] More specifically, the determiner 400 performs processing for each pixel by, for example, subtracting the pixel value of the pixel constituting the output image from the pixel value of the pixel constituting the input image to generate a differential image representing the difference between the input image and the output image.

[0128] In this embodiment, the input image and the output image each represent the intensity of sound using pixel values. In this embodiment, the determiner 400 generates a difference image representing the difference between the input and output images by, for example, subtracting the pixel value of the pixel constituting the output image from the pixel value of the pixel constituting the input image for each pixel.

[0129] At this time, if the input image input to the determiner 400 is an input image generated based on normal sound, image disturbance or the like does not occur in the difference image.

[0130] In contrast, if a synthesized image generated based on occasional abnormal noise is input to determiner 400, image disturbances caused by the abnormal noise will appear in the difference image. In other words, if a synthesized image generated based on occasional abnormal noise is input to determiner 400, images that would not appear when normal sound is present will appear in the difference image.

[0131] At this time, the determiner 400 determines that abnormal noise is generated in the image forming apparatus 100. Then, at this time, information indicating that the abnormal noise is generated is output from the determiner 400.

[0132] The determiner 400 determines abnormal noise generated in the image forming apparatus 100 based on an image appearing in the differential image.

[0133] In this embodiment, the determiner 400 performs machine learning using images generated for each acoustic sensor 120 based on information obtained by the acoustic sensor 120. The determiner 400 then uses the machine-learned abnormality detection model to calculate latent variables based on the features extracted from the images and uses these latent variables for restoration to generate an output image. The determiner 400 then determines abnormal noise generated in the image forming device 100 by comparing the input and output images. In other words, the determiner 400 determines an abnormality in the image forming unit 100A by comparing the input and output images.

[0134] When information indicating that abnormal noise has occurred is output from the determiner 400, the information indicating that abnormal noise has occurred is transmitted to the server device 200 (see FIG. 1 ), which is an example of an external device. Figure 1 ).

[0135] Although the description is omitted above, in this embodiment, Figure 2 As shown in FIG. 2 , a transmitting unit 430 is provided as an example of a transmitting means for transmitting information indicating that abnormal noise has occurred to the server device 200 .

[0136] When abnormal noise is determined by the determiner 400 , the transmitter 430 transmits information indicating that abnormal noise has occurred to the server device 200 .

[0137] The transmission unit 430 is composed of a computer (not shown) and a known transmission device (not shown) for transmitting information.

[0138] This computer includes a CPU (not shown) as an example of a processor. In the transmitter 430, a transmitter is operated according to an instruction from the CPU.

[0139] When abnormal noise is determined by the determiner 400 , the transmitter 430 transmits information indicating that abnormal noise has occurred to the server device 200 .

[0140] In this embodiment, when information indicating the occurrence of abnormal noise is output from determiner 400, information indicating that an abnormality has occurred in image forming unit 100A is also output from determiner 400. In this embodiment, information indicating that an abnormality has occurred in image forming unit 100A is also transmitted to server device 200.

[0141] In this embodiment, the difference images sequentially generated by the determiner 400 are not transmitted to the server device 200. Furthermore, information indicating that abnormal noise has not occurred is not transmitted to the server device 200 either.

[0142] In the present embodiment, only when abnormal noise is determined by the determiner 400 , information indicating that abnormal noise has occurred is transmitted to the server device 200 .

[0143] In this case, in addition to the information indicating the occurrence of abnormal noise, other information may be transmitted to the server device 200. For example, the analysis content by the determiner 400 may also be transmitted to the server device 200.

[0144] In the present embodiment, information on the content of the determination of abnormality in image forming unit 100A is also transmitted to server device 200 only when abnormal noise is determined by determination unit 400 .

[0145] Figure 5 4 is a diagram showing a sound corresponding image generated by the corresponding image generation unit 410 .

[0146] In the corresponding image generating unit 410 (refer to Figure 2 ), the information obtained from the acoustic sensor 120 is subjected to STFT (short-time Fourier transform) processing. In other words, the corresponding image generating unit 410 performs STFT processing on the information obtained from the acoustic sensor 120.

[0147] Thus, generate Figure 5 The sound shown corresponds to image 81 .

[0148] Figure 5 The sound corresponding image 81 shown has a horizontal axis 81A and a vertical axis 81B as two axes that are orthogonal to each other.

[0149] In the sound corresponding image 81, the horizontal axis 81A is the time axis, and the vertical axis 81B is the axis corresponding to the magnitude of the frequency component value. The vertical axis 81B can also be called the frequency axis.

[0150] In the present embodiment, the sound corresponding image 81 is an image having a time axis and a frequency axis, and expressing the intensity of the sound using pixel values.

[0151] In the sound corresponding image 81 , frequency component values ​​corresponding to high frequencies are displayed on the side away from the time axis, and frequency component values ​​corresponding to low frequencies are displayed on the side close to the time axis.

[0152] A white portion in the sound corresponding image 81 indicates that a sound is generated, and a black portion in the sound corresponding image 81 indicates that no sound is generated.

[0153] A vertical axis 81B of the sound corresponding image 81 corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound sensor 120 .

[0154] Each sound corresponding image 81 has a first side 81C extending along a time axis, or horizontal axis 81A, and a second side 81D extending along a time axis, or horizontal axis 81A.

[0155] In the direction in which the longitudinal axis 81B extends, the position of the first side 81C is different from the position of the second side 81D.

[0156] In this embodiment, the first side 81C, which is one of the first side 81C and the second side 81D, is a low-frequency side side 81X located on the low-frequency side. The low-frequency side side 81X is located on the side where the low-frequency component value is displayed in the sound corresponding image 81 .

[0157] The other of the first side 81C and the second side 81D is a high-frequency side side 81Y located on the high-frequency side. The high-frequency side side 81Y is located on the side of the sound corresponding image 81 where the high-frequency component value is displayed.

[0158] In this embodiment, as described above, a plurality of sound sensors 120 (see Figure 2 In this case, a method of providing the determiner 400 for each sound sensor 120 may also be considered.

[0159] In addition, in this case, the number of determiners 400 increases according to the number of sound sensors 120 , and the cost is likely to increase.

[0160] Furthermore, if the determiner 400 is provided for each acoustic sensor 120 , the processing load on the entire image forming apparatus 100 will be increased.

[0161] In contrast, in the present embodiment, the number of determination devices 400 installed is smaller than the number of acoustic sensors 120. In this case, an increase in cost due to an increase in the number of determination devices 400 can be suppressed.

[0162] In this case, the processing load on the entire image forming apparatus 100 is reduced. If the number of determiners 400 installed is reduced compared to the number of acoustic sensors 120 installed, the scale of the circuit installed in the image forming apparatus 100 is reduced.

[0163] In this embodiment, the number of determiners 400 provided is smaller than the number of sound sensors 120 provided.

[0164] In this embodiment, in order to reduce the number of determination units 400 to be provided, synthetic images, which are smaller in number than the total number of the sound sensors 120 , are generated based on the sound corresponding images 81 generated for each sound sensor 120 .

[0165] Then, the synthesized image is input to the determiner 400. Then, the determiner 400 analyzes the synthesized image, thereby determining abnormal noise.

[0166] The synthetic image generating unit 420 (see Figure 2 ) An image obtained by synthesizing a plurality of sound corresponding images 81 generated by a plurality of corresponding image generating units 410 is generated, thereby generating an image obtained by the synthesis, namely a synthesized image.

[0167] In the present embodiment, as described above, the sound corresponding image 81 is generated for each sound sensor 120 based on the information obtained by the sound sensor 120 .

[0168] The synthetic image generation unit 420 synthesizes an image based on the sound corresponding image 81 , which is the sound corresponding image 81 generated by the corresponding image generation unit 410 and is a plurality of the sound corresponding images 81 generated corresponding to the respective sound sensors 120 provided.

[0169] As a result, a synthesized image, which is an image obtained by synthesizing the sound-corresponding images 81 , is generated.

[0170] In this embodiment, it is possible to obtain a smaller number of synthetic images than the total number of the sound sensors 120. Then, in this embodiment, the determiner 400 determines abnormal noise using an autoencoder based on the generated synthetic images.

[0171] In this embodiment, the image forming apparatus 100 (refer to Figure 1 ) accesses the server device 200 via the user terminal 300 and refers to the information about abnormal noise stored in the server device 200.

[0172] In the present embodiment, when abnormal noise is determined by the determiner 400 , information indicating that abnormal noise has occurred in the image forming apparatus 100 is stored in the server apparatus 200 .

[0173] When information indicating that abnormal noise has occurred in image forming apparatus 100 is stored in server device 200 , a service person of image forming apparatus 100 performs work on image forming apparatus 100 such as replacing parts.

[0174] Figure 6 (A) Figure 6 (B) is a diagram for explaining the generation process of the composite image by the composite image generation unit 420 .

[0175] exist Figure 6 (A) shows each sound corresponding image 81 before synthesis. Figure 6 (B) shows a composite image 83 .

[0176] like Figure 6 As shown in FIG. 8(B), the composite image generation unit 420 generates a plurality of sound corresponding images 81 with respective horizontal axes 81A (in Figure 6 (B) (not shown) A composite image 83 of the sound corresponding image 81 is arranged in the form of a specific direction 6A.

[0177] In this embodiment, if Figure 5As shown, each sound corresponding image 81 generated for each sound sensor 120 has a horizontal axis 81A as a time axis and a vertical axis 81B corresponding to the magnitude of a frequency component value obtained at every specific time.

[0178] When generating the composite image 83, as Figure 6 As shown in FIG. 8(B) , the synthesized image generating unit 420 arranges the sound corresponding images 81 so that the horizontal axis 81A extends along a specific direction 6A.

[0179] exist Figure 6 In the example shown, the “one direction 6A” is a direction from the left side toward the right side in the figure.

[0180] exist Figure 6 In (B), the plurality of sound corresponding images 81 are arranged in the intersecting direction 6B which intersects the one direction 6A. More specifically, the plurality of sound corresponding images 81 are arranged in a direction orthogonal to the one direction 6A.

[0181] The synthetic image generation unit 420 generates a synthetic image 83 in which the sound corresponding images 81 are arranged such that their respective horizontal axes 81A extend along one direction 6A and the plurality of sound corresponding images 81 are arranged in a line in a crossing direction 6B that crosses the one direction 6A.

[0182] Furthermore, when generating the composite image 83, as shown in FIG. Figure 6 As shown in FIG. 8 (B), the composite image generation unit 420 connects a plurality of sound corresponding images 81 to each other.

[0183] The arrangement of the sound corresponding images 81 is not limited to a configuration where multiple sound corresponding images 81 are adjacent to each other. For example, multiple sound corresponding images 81 may partially overlap each other. Alternatively, multiple sound corresponding images 81 may be arranged with gaps between them.

[0184] Then, in this embodiment, the generated Figure 6 The composite image 83 shown in (B) is input to the determiner 400 (refer to Figure 2 ). Thus, the determination process of abnormal noise is performed by the determiner 400.

[0185] Specifically, the determiner 400 performs abnormal noise determination processing using the synthesized image 83 and the autoencoder.

[0186] More specifically, in the determiner 400 , abnormal noise determination processing is performed based on the above-described differential image obtained based on the synthesized image 83 and the output image output via the autoencoder.

[0187] Figure 7 (A) Figure 7 (B) is a diagram illustrating synchronization between the sound corresponding images 81.

[0188] In addition, Figure 7 (A) Figure 7 In (B), each sound corresponding image 81 is simplified and displayed. Figure 7 (A) Figure 7 In (B), vertical stripe images 88 caused by the pulse sound generated in the image forming apparatus 100 are included in each sound corresponding image 81 .

[0189] When generating the composite image 83, as Figure 7 As shown in FIG. 8(B), the composite image generation unit 420 generates a composite image 83 in which a plurality of sound corresponding images 81 are arranged in a temporally synchronized manner.

[0190] exist Figure 7 (A) illustrates a case where a composite image 83 is generated in a form where two sound-corresponding images 81 are not synchronized in time.

[0191] In contrast, in Figure 7 (B) illustrates a case where a composite image 83 is generated in a manner where two sound-corresponding images 81 are synchronized in time.

[0192] In this embodiment, a pulse sound may be generated as a normal sound in the image forming apparatus 100. Each time the paper P is conveyed, the paper P may abut against a member on the conveyance path, thereby generating a pulse sound as a normal sound.

[0193] exist Figure 7 (A) Figure 7 In (B), a vertical stripe image 88 caused by the pulse sound is included in each sound corresponding image 81.

[0194] A case where two sound corresponding images 81 are arranged in a temporally synchronized manner will be described.

[0195] At this time, if Figure 7 As shown in (B), when comparing the positions in the time axis direction, the position of the image 88 corresponding to the pulse sound appearing on one sound corresponding image 81 coincides with the position of the image 88 corresponding to the pulse sound appearing on the other sound corresponding image 81.

[0196] Here, if Figure 7 As shown in (A) of FIG. 8 , it is assumed that two sound corresponding images 81 are arranged in a time-asynchronous manner.

[0197] At this time, although the pulse sound is generated at the same timing, the position of the image 88 corresponding to the pulse sound appearing on one sound corresponding image 81 is offset from the position of the image 88 corresponding to the pulse sound appearing on the other sound corresponding image 81 .

[0198] In this case, the accuracy of identifying abnormal noise generated in the image forming apparatus 100 may decrease.

[0199] If the sound corresponding images 81 are arranged in a time-asynchronous manner, the accuracy of identifying abnormal noise may be reduced due to the interaction between the sound corresponding images 81 .

[0200] If the position of the image 88 corresponding to the pulse sound is shifted, the learning by the autoencoder may be inaccurate due to the interaction between the sound corresponding images 81. In this case, the accuracy of the determination of abnormal noise by the determination unit 400 may be reduced.

[0201] If the learning by the autoencoder is incorrect, a sound that is originally a normal sound will be identified as abnormal noise, and the accuracy of identifying the abnormal noise will easily decrease.

[0202] In contrast, in this embodiment, the sound corresponding image 81 is arranged in a temporally synchronized manner when generating the synthesized image 83. In this case, the problem of a decrease in the accuracy of specifying abnormal noise is less likely to occur.

[0203] Here, when the two sound corresponding images 81 are not synchronized in time, for example, one sound corresponding image 81 is shifted in the time axis direction, thereby synchronizing the two sound corresponding images 81 in time.

[0204] More specifically, for example, the following case can be considered: since the corresponding image generation unit 410 (refer to Figure 2 ) and the two sound-corresponding images 81 are not synchronized in time.

[0205] At this time, for example, the sound corresponding image 81 outputted from the faster processing corresponding image generating unit 410 is delayed. Thus, the two sound corresponding images 81 are synchronized in time.

[0206] [Another example of synthesis processing]

[0207] Figure 8 (A) Figure 8 (B) is a diagram showing another processing example for generating the composite image 83 .

[0208] exist Figure 8 (A) shows the state before the composite image 83 is generated. Figure 8(B) shows a state where the composite image 83 is generated.

[0209] exist Figure 8 In the example shown, when generating the composite image 83, as shown in FIG. Figure 8 As shown in (B), the synthesized image generation unit 420 positions the high-frequency side 81Y of one sound corresponding image 81E on the side of the other sound corresponding image 81F. Furthermore, the synthesized image generation unit 420 positions the high-frequency side 81Y of the other sound corresponding image 81F on the side of the one sound corresponding image 81E.

[0210] More specifically, in this example, the synthetic image generating unit 420 makes the Figure 8 After the sound corresponding image 81E on the upper side of (A) is reversed, the reversed sound corresponding image 81E and the sound corresponding image 81E located at Figure 8 A synthesized image 83 is generated by using another sound corresponding image 81F on the lower side of (A).

[0211] In this example, the synthesized image 83 is generated such that the high-frequency side 81Y of one sound corresponding image 81E is located on the other sound corresponding image 81F side, and the high-frequency side 81Y of the other sound corresponding image 81F is located on the one sound corresponding image 81E side.

[0212] At this time, in the portion where the one sound corresponding image 81E and the other sound corresponding image 81F are in contact with each other in the synthesized image 83 , a difference in density between the one sound corresponding image 81E and the other sound corresponding image 81F is unlikely to occur.

[0213] like Figure 6 As shown in (B), consider the following situation: low-frequency side 81X of one sound-corresponding image 81E is located on the side of another sound-corresponding image 81F, while high-frequency side 81Y of the other sound-corresponding image 81F is located on the side of one sound-corresponding image 81E. In this case, a difference in density between one sound-corresponding image 81E and the other sound-corresponding image 81F is likely to occur at the point where one sound-corresponding image 81E and the other sound-corresponding image 81F meet.

[0214] In contrast, if Figure 8 In this case, the accuracy of determining abnormal noise can be suppressed from decreasing.

[0215] It is assumed that there is a large difference between the density of the one sound corresponding image 81E and the density of the other sound corresponding image 81F at a portion where the one sound corresponding image 81E and the other sound corresponding image 81F are in contact with each other.

[0216] In this case, similar to the case of the impulse sound described above, the accuracy of identifying abnormal noise may decrease due to the interaction between the sound-corresponding images 81. More specifically, the learning of the autoencoder may be inaccurate, and the accuracy of identifying abnormal noise may decrease accordingly.

[0217] In contrast, Figure 8 As shown in (B), when the high-frequency side 81Y of one sound corresponding image 81E is located on the side of another sound corresponding image 81F and the high-frequency side 81Y of another sound corresponding image 81F is located on the side of one sound corresponding image 81E, the accuracy of identifying abnormal noise is not likely to decrease.

[0218] Figure 9 (A) Figure 9 (B) is a diagram showing another processing example.

[0219] exist Figure 9 (A) shows the sound corresponding images 81 before synthesis, etc. Figure 9 (B) shows a composite image 83 .

[0220] In this processing example, similarly to the above, when generating the composite image 83, Figure 9 As shown in (B), the high-frequency side 81Y of one sound corresponding image 81E is positioned on the side of the other sound corresponding image 81F.

[0221] Furthermore, the high-frequency side 81Y of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E.

[0222] Furthermore, in this processing example, if Figure 9 (A) Figure 9 As shown in FIG. 8(B) , the composite image generation unit 420 arranges the intermediate image 78 between one sound-corresponding image 81E and another sound-corresponding image 81F.

[0223] The intermediate image 78 is an image composed of images excluding the one sound-corresponding image 81E and the other sound-corresponding image 81F.

[0224] In this processing example, the intermediate image 78 is arranged so as to further reduce the density difference at the boundary between the one sound-corresponding image 81E and the other sound-corresponding image 81F. Here, the density of the intermediate image 78 becomes uniform.

[0225] Figure 10 is Figure 9 An enlarged view of the portion indicated by the symbol X.

[0226] exist Figure 10, intermediate image 78 , a portion of one sound-compatible image 81E, and a portion of another sound-compatible image 81F are shown.

[0227] Here, consider a plurality of pixels 208 constituting one sound corresponding image 81E. The plurality of pixels 208 are arranged in the above-mentioned one direction 6A (refer to Figure 6 (B)) and a plurality of pixels arranged on the intermediate image 78.

[0228] In this embodiment, the density of intermediate image 78 is lower than the density of pixel 208A, which has the highest density among the plurality of pixels 208. Furthermore, in this embodiment, the density of intermediate image 78 is higher than the density of pixel 208B, which has the lowest density among the plurality of pixels 208.

[0229] The plurality of pixels 208 include pixels having a high density value and pixels having a low density value.

[0230] In this embodiment, the density of intermediate image 78 is lower than the density of pixel 208A having the highest density among the plurality of pixels 208. Furthermore, the density of intermediate image 78 is higher than the density of pixel 208B having the lowest density among the plurality of pixels 208.

[0231] In this case, the shading at the boundary between the one sound corresponding image 81E and the other sound corresponding image 81F becomes smaller compared to the case where the one sound corresponding image 81E and the other sound corresponding image 81F are directly in contact with each other.

[0232] When one sound-corresponding image 81E and another sound-corresponding image 81F are directly connected, pixels with high density values ​​in the one sound-corresponding image 81E may be adjacent to pixels with low density values ​​in the other sound-corresponding image 81F.

[0233] In this case, the density at the boundary between the one sound-corresponding image 81E and the other sound-corresponding image 81F becomes larger, and this may reduce the accuracy of identifying the abnormal noise.

[0234] On the other hand, if the intermediate image 78 is provided, it is possible to suppress the increase in shading. In this case, the accuracy of identifying abnormal noise is unlikely to decrease.

[0235] In addition, in this embodiment, the other sound corresponding image 81F also has the same structure.

[0236] Here, it is assumed that a plurality of pixels 209 constitute another sound corresponding image 81F. Specifically, it is assumed that Figure 6 (B)) and a plurality of pixels 209 arranged on the intermediate image 78.

[0237] In this embodiment, the intermediate image 78 has a lower density than the pixel 209A having the highest density among the plurality of pixels 209. Furthermore, the intermediate image 78 has a higher density than the pixel 209B having the lowest density among the plurality of pixels 209.

[0238] The density of the intermediate image 78 is set in advance based on, for example, the density of one previously acquired sound-corresponding image 81E and the density of another previously acquired sound-corresponding image 81F.

[0239] Specifically, the concentration of the intermediate image 78 is set, for example, based on the first average value, which is the average value of the concentrations of the respective pixels of the above-mentioned multiple pixels 208 of one sound corresponding image 81E, and the second average value, which is the average value of the concentrations of the respective pixels of the above-mentioned multiple pixels 209 of another sound corresponding image 81F.

[0240] More specifically, for example, the average value of the first average value and the second average value is set as the density of intermediate image 78 .

[0241] Alternatively, for example, the density of the intermediate image 78 may be determined based on the number of pixels in the sound corresponding image 81E. Figure 10 The first average value is the average value of the respective concentrations of pixels included in the specific area indicated by symbol 10A, and the second average value is the average value of the respective concentrations of pixels included in the specific area indicated by symbol 10B among the pixels constituting another sound corresponding image 81F.

[0242] More specifically, at this time, the average value of the first average value and the second average value may be further calculated, and the calculated average value may be set as the density of intermediate image 78 .

[0243] Figure 11 (A) Figure 11 (B) is a diagram showing another processing example of generating the composite image 83 .

[0244] exist Figure 11 (A) shows the state before the composite image 83 is generated. Figure 11 (B) shows a state where the composite image 83 is generated.

[0245] exist Figure 11 In the example shown, when generating the composite image 83, as shown in FIG. Figure 11 As shown in (B) of FIG. 1 , the composite image generation unit 420 positions the low-frequency side 81X of one sound corresponding image 81E on the side of the other sound corresponding image 81F. Figure 11As shown in FIG. 8(B), the synthesized image generating unit 420 positions the low-frequency side 81X of the other sound corresponding image 81F on the side of the one sound corresponding image 81E.

[0246] More specifically, the composite image generating unit 420 makes the Figure 11 The other sound corresponding image 81F at the lower side of (A) is reversed vertically. Then, the composite image generation unit 420 generates a composite image 83 based on the other sound corresponding image 81F that has been reversed vertically and the one sound corresponding image 81E at the upper side in the figure.

[0247] In this example, the synthesized image 83 is generated such that the low-frequency side 81X of one sound corresponding image 81E is located on the other sound corresponding image 81F side, and the low-frequency side 81X of the other sound corresponding image 81F is located on the one sound corresponding image 81E side.

[0248] In this case, even at the portion where the one sound corresponding image 81E and the other sound corresponding image 81F are in contact with each other, a difference in density between the one sound corresponding image 81E and the other sound corresponding image 81F is unlikely to occur.

[0249] Figure 12 (A) Figure 12 (B) is a diagram showing another processing example.

[0250] exist Figure 12 (A) shows the sound corresponding images 81 before synthesis, etc. Figure 12 (B) shows a composite image 83 obtained by the synthesis.

[0251] In this processing example, when generating the composite image 83, as shown in FIG. Figure 12 As shown in (B), the low-frequency side 81X of one sound corresponding image 81E is positioned on the side of the other sound corresponding image 81F. Furthermore, the low-frequency side 81X of the other sound corresponding image 81F is positioned on the side of the one sound corresponding image 81E.

[0252] Furthermore, in this processing example, when generating the synthesized image 83 , the synthesized image generating unit 420 places the intermediate image 78 between the one sound-corresponding image 81E and the other sound-corresponding image 81F, similarly to the above.

[0253] In this processing example, the density at the boundary between the one sound-corresponding image 81E and the other sound-corresponding image 81F is also reduced by the intermediate image 78. Here, as described above, the density of the intermediate image 78 is uniform.

[0254] Figure 13 is Figure 12 An enlarged view of the portion indicated by symbol XIII.

[0255] exist Figure 13 2 also shows the intermediate image 78, a portion of one sound corresponding image 81E, and a portion of another sound corresponding image 81F.

[0256] Here, we also consider the plurality of pixels 212 constituting one sound corresponding image 81E. As described above, the plurality of pixels 212 are arranged in the one direction 6A (refer to Figure 6 (B)) and a plurality of pixels arranged on the intermediate image 78.

[0257] In this embodiment, the density of intermediate image 78 is lower than the density of pixel 212A having the highest density among the plurality of pixels 212. Furthermore, in this embodiment, the density of intermediate image 78 is higher than the density of pixel 212B having the lowest density among the plurality of pixels 212.

[0258] As described above, the plurality of pixels 212 include pixels having a high density value and pixels having a low density value.

[0259] In the present embodiment, the density value of the intermediate image 78 is smaller than the density value of the pixel 212A having the highest density value among the plurality of pixels 212 .

[0260] Furthermore, in the present embodiment, the density value of the intermediate image 78 is larger than the density value of the pixel 212B having the smallest density value among the plurality of pixels 212 .

[0261] At this time, similarly to the above, the shading at the boundary between the one sound corresponding image 81E and the other sound corresponding image 81F becomes smaller than when the one sound corresponding image 81E and the other sound corresponding image 81F are directly in contact with each other.

[0262] The same applies to the other sound corresponding image 81F.

[0263] Here, it is assumed that a plurality of pixels 213 constitute another sound corresponding image 81F. Specifically, it is assumed that Figure 6 (B)) and a plurality of pixels 213 arranged on the intermediate image 78.

[0264] In this embodiment, the density of intermediate image 78 is lower than the density of pixel 213A having the highest density among the plurality of pixels 213. Furthermore, in this embodiment, the density of intermediate image 78 is higher than the density of pixel 213B having the lowest density among the plurality of pixels 213.

[0265] As described above, the density of the intermediate image 78 is set in advance based on, for example, the density of one previously acquired sound-corresponding image 81E and the density of another previously acquired sound-corresponding image 81F.

[0266] Specifically, the concentration of the intermediate image 78 is set, for example, based on the first average value, which is the average value of the concentrations of the respective pixels of the above-mentioned multiple pixels 212 of one sound corresponding image 81E, and the second average value, which is the average value of the concentrations of the respective pixels of the above-mentioned multiple pixels 213 of another sound corresponding image 81F.

[0267] More specifically, for example, the average value of the first average value and the second average value is set as the density of intermediate image 78 .

[0268] Alternatively, for example, the concentration of the intermediate image 78 can also be set based on the average value of the respective concentrations of pixels contained in a specific area indicated by symbol 13A among the pixels constituting one sound corresponding image 81E, that is, the first average value, and the average value of the respective concentrations of pixels contained in a specific area indicated by symbol 13B among the pixels constituting another sound corresponding image 81F, that is, the second average value.

[0269] More specifically, at this time, the average value of the first average value and the second average value may be further calculated, and the calculated average value may be set as the density of intermediate image 78 .

[0270] In addition, although the illustration is omitted, Figure 6 In the composite image 83 shown, the intermediate image 78 may be arranged between one sound-corresponding image 81E and another sound-corresponding image 81F constituting the composite image 83 .

[0271] exist Figure 6 In the example shown, when the intermediate image 78 is arranged, for example, it is preferable that the density of the intermediate image 78 gradually increases from the side of one sound corresponding image 81E toward the side of the other sound corresponding image 81F.

[0272] exist Figure 6 In the example shown, the low-frequency side 81X of one sound corresponding image 81E is located on the side of the other sound corresponding image 81F, and the high-frequency side 81Y of the other sound corresponding image 81F is located on the side of the one sound corresponding image 81E.

[0273] For example, at this time, it is preferable that the density of the intermediate image 78 gradually increases from the side of the one sound corresponding image 81E toward the side of the other sound corresponding image 81F.

[0274] Furthermore, in the above description, the composite image 83 is generated from the two sound-corresponding images 81 , but the present invention is not limited to this.

[0275] For example, Figure 14 As shown in (Figure showing another example of a composite image), a composite image 83 may be generated based on three or more sound corresponding images 81.

[0276] exist Figure 14 In the illustrated example, a case is illustrated where one composite image 83 is formed based on four sound-corresponding images 81 generated in response to four sound sensors 120 provided.

[0277] In addition, Figure 14 In the example shown, similarly to the above, between two sound corresponding images 81 adjacent to each other in the intersecting direction 6B that intersects the one direction 6A, the low-frequency sides 81X are in contact with each other, or the high-frequency sides 81Y are in contact with each other.

[0278] And, in addition, Figure 14 In the example shown, it is also possible to Figure 9 、 Figure 12 In the same manner as shown, the intermediate image 78 is arranged between two adjacent sound corresponding images 81 .

[0279] (Note) (1)

[0281] An image forming apparatus comprising:

[0282] an image forming unit for forming an image on a recording medium;

[0283] a plurality of sound detection units for detecting sounds of the image forming unit; and

[0284] a determination unit that determines an abnormality in the image forming unit based on information output from the plurality of sound detection units,

[0285] The number of the determination sections is set to be smaller than the number of the plurality of sound detection sections. (2)

[0287] The image forming apparatus according to (1) further includes a transmitting unit configured to transmit information indicating that an abnormality has occurred to an external device when the abnormality is identified by the identifying unit. (3)

[0289] The image forming apparatus according to (1) or (2), wherein

[0290] The determination unit performs machine learning on an image generated based on information obtained by the sound detection unit and for each image generated by the sound detection unit, uses an abnormality detection model after machine learning to calculate latent variables based on feature quantities extracted from the image, and uses the latent variables for restoration to generate an output image, and determines the abnormality of the image forming unit by comparing the input image and the output image. (4)

[0292] The image forming apparatus according to any one of (1) to (3), further comprising a generating unit that generates an image obtained by synthesizing an image based on a plurality of images, namely, a synthesized image, wherein the image is generated based on information obtained by the sound detection unit and is generated for each sound detection unit, and the plurality of images are generated in accordance with the plurality of sound detection units provided.

[0293] The determining unit determines an abnormality based on the synthesized image generated by the generating means. (5)

[0295] The image forming apparatus according to (4), wherein

[0296] Each of the images generated by the sound detection unit has a time axis and a frequency axis and represents the intensity of the sound using pixel values.

[0297] The generating means generates the composite image in which the plurality of images are arranged such that their respective time axes are along a specific direction. (6)

[0299] The image forming apparatus according to (5), wherein

[0300] When the composite image is generated by synthesizing the plurality of images, the generating means generates the composite image in which the plurality of images are arranged such that their respective time axes are along the one direction and the plurality of images are arranged in a direction intersecting the one direction. (7)

[0302] The image forming apparatus according to (5), wherein

[0303] When the synthesized image is generated by synthesizing the plurality of images, the generating means generates the synthesized image in which the plurality of images are arranged in a temporally synchronized manner. (8)

[0305] The image forming apparatus according to (6), wherein

[0306] The frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit.

[0307] Each of the images has a first side extending along the direction in which the time axis extends, and a second side extending along the direction in which the time axis extends and in the direction in which the frequency axis extends, and having a position different from that of the first side.

[0308] One of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where a low-frequency component value is displayed in the image, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where a high-frequency component value is displayed in the image.

[0309] When the composite image is generated by synthesizing at least one image and another image included in the plurality of images, the generating component generates a composite image in which the high-frequency side of the one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image. (9)

[0311] The image forming apparatus according to (8), wherein

[0312] When generating the composite image by synthesizing images using at least the one image and the other image, the generating means arranges an intermediate image, which is an image other than the one image and the other image, between the one image and the other image. (10)

[0314] The image forming apparatus according to (9), wherein

[0315] The density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the multiple pixels arranged in the one direction among the multiple pixels constituting the one image and connected to the intermediate image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the multiple pixels arranged in the one direction,

[0316] The concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction. (11)

[0318] The image forming apparatus according to (6), wherein

[0319] The frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit.

[0320] Each of the images has a first side extending along the direction in which the time axis extends, and a second side extending along the direction in which the time axis extends and in the direction in which the frequency axis extends, and having a position different from that of the first side.

[0321] One of the first side and the second side is a low-frequency side side located on a low-frequency side and is arranged on a side where a low-frequency component value is displayed in the image, and the other side is a high-frequency side side located on a high-frequency side and is arranged on a side where a high-frequency component value is displayed in the image.

[0322] When the composite image is generated by synthesizing at least one image and another image included in the plurality of images, the generating component generates a composite image in which the low-frequency side of the one image is located on the side of the other image and the low-frequency side of the other image is located on the side of the one image. (12)

[0324] The image forming apparatus according to (11), wherein

[0325] When generating the composite image by synthesizing images using at least the one image and the other image, the generating means arranges an intermediate image, which is an image other than the one image and the other image, between the one image and the other image. (13)

[0327] The image forming apparatus according to (12), wherein

[0328] The density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the multiple pixels arranged in the one direction among the multiple pixels constituting the one image and connected to the intermediate image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the multiple pixels arranged in the one direction,

[0329] The concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction. (14)

[0331] An image forming apparatus comprising:

[0332] an image forming unit for forming an image on a recording medium;

[0333] a plurality of sound detection units for detecting sounds of the image forming unit;

[0334] a generating means for generating an image obtained by synthesizing an image based on an image, wherein the image is generated based on information obtained by the sound detection unit and is generated for each sound detection unit, and a plurality of such images are generated corresponding to the number of sound detection units provided; and

[0335] The identifying unit identifies an abnormality in the image forming unit based on the synthesized image.

[0336] According to the image forming apparatus of (1), the hardware required for abnormality identification can be reduced compared to a case where a functional unit for identifying abnormality is separately provided corresponding to each of a plurality of sound detection units.

[0337] According to the image forming apparatus according to (2), the amount of information transmitted to the external device can be reduced compared to a case where information related to an abnormality is transmitted to the external device regardless of whether an abnormality is confirmed or not.

[0338] According to the image forming apparatus according to (3), even when only normal sounds are basically generated in the image forming apparatus and abnormal noise is unlikely to be generated, an abnormality occurring in the image forming apparatus can be identified.

[0339] According to the image forming apparatus according to (4), the load of processing required for determining an abnormality can be reduced compared to a case where an abnormality is determined for each image obtained by each sound detection unit.

[0340] According to the image forming apparatus of (5), the accuracy of identifying an abnormality can be improved compared to a case where the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends.

[0341] According to the image forming apparatus of (6), the accuracy of identifying an abnormality can be improved compared to a case where the direction in which the time axis of one image included in a plurality of images extends is different from the direction in which the time axis of another image extends.

[0342] According to the image forming apparatus of (7), the accuracy of identifying abnormalities can be improved compared to a case where a plurality of images are not arranged in a temporally synchronized manner.

[0343] According to the image forming apparatus of (8), the accuracy of abnormality identification can be improved compared to a case where the low-frequency side of one image is located on the other image side and the high-frequency side of the other image is located on the one image side.

[0344] According to the image forming apparatus of (9), the accuracy of identifying an abnormality can be improved compared to a case where one image and another image are directly connected.

[0345] According to the image forming apparatus of (10), the accuracy of identifying an abnormality can be improved compared to a case where one image and another image are directly connected.

[0346] According to the image forming apparatus of (11), the accuracy of abnormality determination can be improved compared to a case where the low-frequency side of one image is located on the other image side and the high-frequency side of the other image is located on the one image side.

[0347] According to the image forming apparatus of (12), the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0348] According to the image forming apparatus of (13), the accuracy of identifying abnormalities can be improved compared to a case where one image and another image are directly connected.

[0349] According to the image forming apparatus of (14), the hardware required for determining an abnormality can be reduced compared to a case where an abnormality is determined without generating a synthetic image.

[0350] The above-described embodiments of the present invention are provided for the purpose of illustration and explanation. In addition, the embodiments of the present invention do not fully and exhaustively include the present invention, and do not limit the present invention to the disclosed embodiments. It is obvious that various modifications and variations are self-evident to those skilled in the art to which the present invention belongs. The present embodiment is selected and described in order to most easily explain the principles of the present invention and its application. Thus, other technical personnel in this field can understand the present invention through various modifications optimized for specific uses of the assumed various embodiments. The scope of the present invention is defined by the above claims and their equivalents.

Claims

1. An image forming apparatus comprising: an image forming unit for forming an image on a recording medium; a plurality of sound detection units for detecting sounds of the image forming unit; and a determination unit that determines an abnormality in the image forming unit based on information output from the plurality of sound detection units, The number of the determination sections is set to be smaller than the number of the plurality of sound detection sections. 2 . The image forming apparatus according to claim 1 , further comprising a transmitting unit configured to transmit information indicating that the abnormality has occurred to an external device when the abnormality has been identified by the identifying unit.

3. The image forming apparatus according to claim 1 or 2, wherein: The determination unit performs machine learning using an image generated based on information obtained by the sound detection unit and for each of the sound detection units, uses an abnormality detection model after machine learning to calculate latent variables based on feature quantities extracted from the image, and uses the latent variables for restoration to generate an output image, and determines the abnormality of the image forming unit by comparing the input image and the output image.

4. The image forming apparatus according to any one of claims 1 to 3, further comprising a generating unit configured to generate a composite image by synthesizing images from a plurality of images, wherein: The image is generated based on the information obtained by the sound detection unit and is generated for each sound detection unit, and is a plurality of images generated corresponding to the plurality of sound detection units provided. The determining unit determines an abnormality based on the synthesized image generated by the generating means.

5. The image forming apparatus according to claim 4, wherein Each of the images generated by the sound detection unit has a time axis and a frequency axis and represents the intensity of the sound using pixel values. The generating unit generates the composite image in which the plurality of images are arranged such that their respective time axes are along a specific direction.

6. The image forming apparatus according to claim 5, wherein When the composite image is generated by synthesizing the images based on the plurality of images, the generating component generates the composite image in which the plurality of images are arranged in such a manner that their respective time axes are along the one direction and the plurality of images are arranged in such a manner that they are aligned in a direction intersecting the one direction.

7. The image forming apparatus according to claim 5, wherein: When the composite image is generated by synthesizing the plurality of images, the generating unit generates the composite image in which the plurality of images are arranged in a temporally synchronized manner.

8. The image forming apparatus according to claim 6, wherein The frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit. Each of the images has a first side extending along the direction in which the time axis extends, and a second side extending along the direction in which the time axis extends and in the direction in which the frequency axis extends, and having a position different from that of the first side. One of the first side and the second side is a low-frequency side side located on a low-frequency side and is disposed on a side where a low-frequency component value is displayed in the image, and the other side is a high-frequency side side located on a high-frequency side and is disposed on a side where a high-frequency component value is displayed in the image. When the composite image is generated by synthesizing images using at least one image and another image included in the plurality of images, the generating component generates a composite image in which the high-frequency side of the one image is located on the side of the other image and the high-frequency side of the other image is located on the side of the one image.

9. The image forming apparatus according to claim 8, wherein When generating the composite image by synthesizing images using at least the one image and the other image, the generating means arranges an intermediate image, which is an image other than the one image and the other image, between the one image and the other image.

10. The image forming apparatus according to claim 9, wherein The density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the one image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the multiple pixels arranged in the one direction, The concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction.

11. The image forming apparatus according to claim 6, wherein The frequency axis of the image corresponds to the magnitude of the frequency component value obtained by analyzing the information obtained by the sound detection unit. Each of the images has a first side extending along the direction in which the time axis extends, and a second side extending along the direction in which the time axis extends and in the direction in which the frequency axis extends, and having a position different from that of the first side. One of the first side and the second side is a low-frequency side side located on a low-frequency side and is disposed on a side where a low-frequency component value is displayed in the image, and the other side is a high-frequency side side located on a high-frequency side and is disposed on a side where a high-frequency component value is displayed in the image. When the composite image is generated by synthesizing images using at least one image and another image included in the plurality of images, the generating component generates a composite image in which the low-frequency side of the one image is located on the side of the other image and the low-frequency side of the other image is located on the side of the one image.

12. The image forming apparatus according to claim 11, wherein When generating the composite image by synthesizing images using at least the one image and the other image, the generating means arranges an intermediate image, which is an image other than the one image and the other image, between the one image and the other image.

13. The image forming apparatus according to claim 12, wherein: The density value of the intermediate image is smaller than the density value of the pixel with the largest density value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the one image, and the density value of the intermediate image is larger than the density value of the pixel with the smallest density value among the multiple pixels arranged in the one direction, The concentration value of the intermediate image is smaller than the concentration value of the pixel with the largest concentration value among the multiple pixels arranged in the one direction and connected to the intermediate image among the multiple pixels constituting the other image, and the concentration value of the intermediate image is larger than the concentration value of the pixel with the smallest concentration value among the multiple pixels arranged in the one direction.

14. An image forming apparatus comprising: an image forming unit for forming an image on a recording medium; a plurality of sound detection units for detecting sounds of the image forming unit; A generating component synthesizes images based on the images to generate an image obtained by the synthesis, namely a synthesized image, wherein: The image is generated based on information obtained by the sound detection unit and is generated for each sound detection unit, and a plurality of images are generated corresponding to the number of sound detection units provided; and The identifying unit identifies an abnormality in the image forming unit based on the synthesized image.

Citation Information

Patent Citations

  • Image forming apparatus with self-checking function

    JP2006184722A