Generation method, detection method, generation device, detection system, control program, and recording medium

The detection system generates reference images through learning models to complement mask areas, addressing the challenge of detecting infrequent or unpredictable abnormalities by ensuring accurate comparisons and improving abnormality detection accuracy.

WO2025220629A1PCT designated stage Publication Date: 2025-10-23KYOCERA CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/014596
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately detecting abnormalities in objects, particularly those that occur infrequently or unpredictably, due to the difficulty in collecting sufficient images with varying conditions and the inherent changes in living organisms over time, leading to inaccurate comparisons between reference and first images.

Method used

A detection system that generates reference images using a learning model to complement mask areas in first images, allowing for accurate comparison and detection of abnormalities by setting mask regions and using pseudo-images to match shooting conditions, even in living organisms.

Benefits of technology

Enables precise detection of abnormalities by generating reference images that faithfully reproduce overall object features while avoiding local anomalies, enhancing the accuracy of abnormality detection in first images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025014596_23102025_PF_FP_ABST
    Figure JP2025014596_23102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention generates a reference image for comparison with an image captured of a subject, said reference image making it possible to accurately detect an abnormality occurring in the subject. A generation method according to the present disclosure generates a reference image used as a comparison target for a first image in order to detect the condition of a first subject appearing in the first image. The generation method includes: an acquisition step for acquiring the first image; a second setting step for setting one or more mask regions in the first image; and a generation step for generating one or more reference images using a second image generated by a learning model as image(s) of the mask region(s).
Need to check novelty before this filing date? Find Prior Art

Description

Generation method, detection method, generation device, detection system, control program, and recording medium

[0001] The present disclosure relates to a generation method and generation device for generating a reference image used as a comparison object for detecting the state of an object shown in an image, a detection method and detection system for detecting the state of an object shown in an image, etc.

[0002] Conventionally, when detecting an abnormality in an object based on an image of the object, a technique using a learning model that has been machine-learned in advance based on images showing the abnormality in the object is known. For example, Patent Literature 1 discloses a patrol inspection system that detects an abnormality in an object to be inspected from an image using a learning model. Furthermore, Patent Literature 2 discloses a cerebrovascular disease learning device that constructs a learning model to detect cerebrovascular disease from images obtained using MR (magnetic resonance) angiography.

[0003] Japanese Patent Publication No. 2021-090348 Japanese Special Publication No. 2022-516146

[0004] <1> An information processing method according to one aspect of the present disclosure includes an acquisition step of acquiring a first image that shows a first object, a second setting step of setting one or more mask areas in the first image, and a generation step of generating one or more reference images using a second image generated by a learning model as an image of the mask area, wherein the reference images are used as objects for comparison with the first image, and the state of the first object is detected based on the comparison result between the reference image and the first image.

[0005] <2> A detection method according to one aspect of the present disclosure includes a comparison step of comparing the first image with the reference image generated by the generation method of <1> and outputting a comparison result, and a detection step of detecting the state based on the comparison result.

[0006] <3> A generation device according to one aspect of the present disclosure includes an acquisition unit that acquires a first image that shows a first object, a second setting unit that sets one or more mask areas in the first image, and a generation unit that generates one or more reference images using second images generated by a learning model as images of the mask areas, the reference images being used as objects for comparison with the first image, and detecting the state of the first object based on the comparison result between the reference image and the first image.

[0007] <4> A detection system according to one aspect of the present disclosure includes an acquisition unit that acquires a first image containing a first object; a second setting unit that sets one or more mask areas in the first image; a generation unit that generates one or more reference images using a second image generated by a learning model as an image of the mask area; a comparison unit that compares the first image with the reference image and outputs a comparison result; and a detection unit that detects the state of the first object based on the comparison result.

[0008] The generating device according to each aspect of the present disclosure may be realized by a computer. In this case, the control program of the generating device that causes the computer to operate as each part (software element) of the generating device to realize the generating device, and the computer-readable recording medium on which it is recorded, also fall within the scope of the present disclosure.

[0009] The detection system according to each aspect of the present disclosure may be realized by a computer. In this case, the control program of the detection system that causes the computer to operate as each part (software element) of the detection system to realize the detection system, and the computer-readable non-transitory recording medium on which the control program is recorded, also fall within the scope of the present disclosure.

[0010] FIG. 1 is a functional block diagram showing an example of a schematic configuration of a detection system according to an embodiment of the present disclosure; FIG. 2 is a flowchart showing an example of a processing flow performed by a generation device; FIG. 3 is an image diagram showing an example of a processing flow performed by a generation device; FIG. 4 is a flowchart showing an example of a processing flow performed by a detection device; FIG. 5 is an image diagram showing an example of a processing flow performed by a detection device; FIG. 6 is a functional block diagram showing an example of a schematic configuration of a detection system according to another embodiment of the present disclosure; FIG. 7 is a flowchart showing an example of a processing flow performed by a generation device; and FIG. 8 is a sequence diagram showing an example of a processing flow performed by a detection system.

[0011] First Embodiment Hereinafter, an embodiment of the present disclosure will be described in detail using a detection system 100 as an example.

[0012] In the technology described in Patent Literature 1, in order to perform machine learning on a learning model, it is necessary to collect a large number of images showing abnormalities that are expected to occur in the object. When abnormalities occur in the object infrequently or when a wide variety of abnormalities occur in the object, it is difficult to collect the required number of images showing various abnormalities. In addition, it is difficult to detect abnormalities that have not been anticipated in advance.

[0013] In order to detect anomalies that occur infrequently, a wide variety of anomalies, and anomalies that cannot be predicted in advance, it is effective to compare an image showing an object for which anomalies are to be detected (hereinafter referred to as the first image) with an image showing an object in which no anomalies to be detected exist (hereinafter referred to as the reference image).

[0014] Examples of images that are expected to be used as the reference image include the following (1) and (2).

[0015] (1) An image showing an object that has been confirmed to be free of abnormalities, but is different from the object for which abnormalities are to be detected.

[0016] (2) An image of the object for which abnormalities are to be detected, taken at a time in the past when it was confirmed that there were no abnormalities.

[0017] However, whether the image (1) or (2) is used as the reference image, the shooting conditions (e.g., angle of view, posture, other shooting settings, etc.) do not match those of the first image, so it may be impossible to accurately detect an abnormality from the difference between the first image and the reference image. Furthermore, when the object is a living organism, it is often virtually impossible to acquire the image described in (1) above as the reference image. Furthermore, when the object is a living organism, the object itself may change over time. Therefore, even if the image described in (2) above is acquired as the reference image, there will be a difference between an image of the object taken in the past and an image of the object currently captured, due to the change in the object over time. Therefore, it may be difficult to accurately detect an abnormality from the difference between the first image and the reference image.

[0018] According to one aspect of the present disclosure, it is possible to generate a reference image for comparison with an image of an object, the reference image being capable of accurately detecting an abnormality occurring in the object.

[0019] (Detection System 100) First, the detection system 100 will be described with reference to Fig. 1. Fig. 1 is a functional block diagram showing an example of the schematic configuration of the detection system 100.

[0020] The detection system 100 includes a generating device 1 and a detecting device 3. Both the generating device 1 and the detecting device 3 may be cloud-based devices, or may be on-premise devices installed in a medical facility or a company that provides analysis services.

[0021] 1 illustrates one generation device 1 that can communicate with the detection device 3, but the present invention is not limited to this configuration. For example, in the detection system 100, the detection device 3 may be able to communicate with multiple generation devices 1.

[0022] 1 illustrates one information processing device 7 that can communicate with the generating device 1 and the detecting device 3, but the present invention is not limited to this configuration. For example, in the detection system 100, the generating device 1 and the detecting device 3 may be able to communicate with multiple information processing devices 7.

[0023] The generation device 1, the detection device 3, and the information processing device 7 can communicate with each other via a communication network 9. The communication network 9 may be the Internet, an intranet, a blockchain network, a wireless LAN (Local Area Network), a WAN (Wide Area Network), or the like. For example, if the generation device 1 and the detection device 3 are cloud-based devices, the communication network 9 may be the Internet, a blockchain network, or the like. If the generation device 1 and the detection device 3 are on-premise devices, the communication network 9 may be a wireless LAN provided within a medical facility or company.

[0024] The detection system 100 is a system that detects the state of a first object captured in the first image 711 based on a comparison result between the first image 711 and a reference image (described later). The first object may be a living object or an inanimate object. The first image 711 may be an image captured by any imaging method.

[0025] For example, the first object may be a part or the whole of the subject's body. For example, the first object may be the subject's head, chest, waist, arms, legs, etc. The first object may also be the subject's face, eyes, nose, mouth, lips, ears, hair, beard, skin, etc. For example, if the first object is skin and the first image 711 is a photographic image of the subject's skin, the detection system 100 can detect the condition of the subject's skin shown in the first image 711.

[0026] For example, if the first object is an organ, muscle, bone, joint, or the like of a subject, and the first image 711 is a medical image of the subject, the detection system 100 can detect the condition of the subject's organ, muscle, bone, joint, or the like shown in the first image 711. When the first object is the subject's bone and / or joint, the first image 711 may be, for example, at least one of a plain X-ray image of the subject's bone, an ultrasound image, an image obtained by a dual energy X-ray absorptiometry (DXA) method, a computed tomography (CT) image, and an image obtained by dual energy subtraction (DES). When the first object is the subject's organ and / or muscle, the first image 711 may be, for example, at least one of a plain X-ray image of the subject's organ and / or muscle, an ultrasound image, an image obtained by a DXA method, a CT image, and an image obtained by DES. The first image 711 may be at least one of a front image of the first object from the front (e.g., an image obtained by irradiating the first object with X-rays in the front-to-back direction, etc.) and a side image of the first object from the side (e.g., an image obtained by irradiating the first object with X-rays in the left-to-right direction, etc.).

[0027] For example, if the first object is a building and an item, and the first image 711 is an image of the building and the item, the detection system 100 can detect the state of the building and the product shown in the first image 711. Here, the building may include, for example, a building, a road, a wall, etc. The item may include, for example, utensils and containers used for various purposes. For example, if the first object is food and processed products, and the first image 711 is an image of the food and processed products, the detection system 100 can detect the state of the food and processed products shown in the first image.

[0028] For example, if the first image 711 is a medical image of the subject's body, the detection system 100 can detect the condition of the first object of the subject that appears in the first image 711. Examples of the condition of the first object of the subject that the detection system 100 can detect include, for example, a disease of the subject, a foreign object or device placed on the subject's body, and the like. Here, diseases that the detection system 100 can detect may include bone fractures and rare diseases, and bone fractures may include stress fractures or fragility fractures. Rare diseases are rare diseases. Rare diseases may include, for example, eosinophilic gastroenteritis, eosinophilic colitis, eosinophilic pneumonia, and the like. Examples of foreign objects that the detection system 100 can detect include implants, etc. Furthermore, examples of devices that the detection system 100 can detect include pacemakers, etc.

[0029] In the present disclosure, a case will be described where the first image 711 is a medical image such as a plain X-ray image, for example. In the following, the detection system 100 is a system for detecting the condition of bones in the first image 711 that shows the bones of a subject (first object).

[0030] Furthermore, although the present disclosure will be described with reference to an example in which the subject is a human, the subject is not limited to humans. The subject may be a non-human mammal, such as an equine, feline, canine, bovine, or porcine animal, or may be a non-mammalian animal (e.g., a bird, reptile, amphibian, or fish).

[0031] <Information Processing Device 7> The information processing device 7 may be a computer provided in a medical facility or company that uses the detection system 100. Alternatively, the information processing device 7 may be a computer used by a medical professional such as a doctor that uses the detection system 100.

[0032] The information processing device 7 includes a control unit 70, a storage unit 71, a communication unit 72, and a display unit 73. In Fig. 1, the information processing device 7 includes the display unit 73 as an example, but is not limited to this. For example, instead of including the display unit 73, an external display device (not shown) may be used.

[0033] The control unit 70 may be, for example, a CPU (Central Processing Unit) and performs overall control of the information processing device 7. The control unit 70 reads a control program, which is software stored in the storage unit 71, expands it in a memory such as a RAM (Random Access Memory), and controls each component included in the information processing device 7. The control unit 70 includes an acquisition unit 701 that acquires the detection result 313 from the detection device 3, and a display control unit 702 that displays the detection result 313 on the display unit 73.

[0034] The storage unit 71 is a storage device that stores various control programs and various data, and may store a first image 711. For the sake of simplicity, the control programs are not shown in the storage unit 71 illustrated in FIG.

[0035] The communication unit 72 transmits and receives various data to and from the detection device 3 and the generation device 1. For example, the information processing device 7 transmits a first image 711 showing the bones of the subject to the generation device 1 via the communication unit 72. The information processing device 7 also receives a detection result 313 of the detected bone condition of the subject from the detection device 3 via the communication unit 72.

[0036] The display unit 73 displays various data. The display unit 73 displays, for example, the detection result 313 received from the detection device 3 and the first image. The detection result 313 may include information indicating the location where the bone condition of the subject was detected, information indicating the detected condition, etc. If the first object is a bone, the detection result 313 may include information indicating whether the detected condition is a fracture, a foreign body, or a device. A medical professional using the information processing device 7 can examine the bone condition of the subject based on the detection result 313 received from the detection device 3 and the first image displayed on the display unit 73.

[0037] <Generation device 1> The generation device 1 is a device that generates a reference image used as a comparison target for the first image 711 in order to detect the state of the first object shown in the first image 711. The generation device 1 includes a control unit 10, a storage unit 11, and a communication unit 12.

[0038] The control unit 10 may be a CPU, for example. The control unit 10 reads a control program, which is software stored in the storage unit 11, expands it into a memory such as a RAM, and controls each component included in the generation device 1. The control unit 10 may include an acquisition unit 101, a first setting unit 102a, a second setting unit 102b, and a generation unit 103.

[0039] The acquisition unit 101 acquires a first image 711 from the information processing device 7. The acquired first image 711 may be stored in the image data 111 of the storage unit 11.

[0040] The first setting unit 102a sets a first region corresponding to a bone in the first image 711. The first setting unit 102a may set as the first region the entirety or a part of the first image 711. When setting as the first region a part of the first image 711, the first setting unit 102a may set as the first region one or more of the central region, upper region, lower region, right region, and left region of the first image 711.

[0041] The first setting unit 102a may set a region in the first image 711 in which bones are captured as the first region by segmentation. Segmentation refers to dividing an image into several regions. If the first image is an X-ray image of the subject's lumbar vertebrae, for example, a range including the region in which the lumbar vertebrae are captured may be segmented. If the first image is an X-ray image of a knee joint, for example, a range including the regions in which the femur, tibia, fibula, patella, etc. are captured may be segmented.

[0042] If the first object is a bone, the first setting unit 102a sets, as the first region, an area in which the bone appears in the first image 711. If the first object is any one of an organ, a muscle, and a joint, the first setting unit 102a may set, as the first region, an area in which any one of an organ, a muscle, and a joint appears in the first image 711.

[0043] The first region may be set using a segmentation model generated by machine learning using the same type of image as the first image 711. As the segmentation model, for example, a convolutional neural network (CNN), a fully convolutional network (FCN), a U-Net, a V-Net, a Vision Transformer (ViT), or the like can be applied.

[0044] The second setting unit 102b sets one or more mask areas in the first area set by the first setting unit 102a. The image of the mask area is complemented using a second image generated by a learning model (described later), and one or more reference images are generated. The shape of the mask area may be a circle, an ellipse, or a polygon other than a rectangle. The second setting unit 102b, which sets the first area in the first image 711, is not an essential component of the generation device 1.

[0045] The second setting unit 102b sets one or more mask areas in the first image 711. If the first area has not been set by the first setting unit 102a, the second setting unit 102b may set a mask area in the first image 711. In this case, the second setting unit 102b may set a plurality of mask areas in the first image 711, each of which corresponds to the entire first image 711.

[0046] Alternatively, when the first setting unit 102a has set a first region, the second setting unit 102b may set a mask region in the first region. In this case, the second setting unit 102b may set a plurality of mask regions in the first region, each mask region corresponding to the entire first region.

[0047] The mask region may be set using a trained segmentation model. The segmentation model may be generated by machine learning using a fourth image that is the same type of image as the first image 711. For example, if the first image 711 is a plain X-ray image, the plain X-ray image is used as the fourth image.

[0048] As the segmentation model, for example, a convolutional neural network (CNN), a fully convolutional network (FCN), a U-Net, a V-Net, a Vision Transformer (ViT), etc. can be applied. The mask region may be the same region as the first region, or may be a region that includes a part of the first region and is smaller than the first region.

[0049] The generation unit 103 generates one or more reference images using the second image generated by the learning model as an image of the mask region set by the second setting unit 102b.

[0050] The generating device 1 may be configured to repeat the setting of one or more mask areas by the second setting unit 102b and the complementation of the mask areas with the second image by the generating unit 103 until a predetermined condition is satisfied. Here, the predetermined condition may be, for example, that the entire first image 711 is complemented with the second image.

[0051] Alternatively, the generating device 1 may be configured to first set one or more mask areas in a batch using the second setting unit 102b, and then also to complement the mask areas with the second image in a batch using the generating unit 103.

[0052] When the first object is a bone, the machine learning of the learning model uses images of the bone of the second subject that are free of conditions (e.g., abnormalities, foreign objects, devices, etc.) that would be detected by the detection device 3. When the first object is an organ, a muscle, a joint, etc., the machine learning of the learning model uses images of the organ, a muscle, a joint, etc. of the second subject that are free of conditions that would be detected by the detection device 3. As a result, the generation unit 103 generates a second image that is free of conditions (e.g., abnormalities, foreign objects, devices, etc.) that would be detected by the detection device 3.

[0053] The training model may be a model that generates new pseudo-images and pseudo-data from input images and data. The training model may be a model generated by training using, for example, a generative adversarial network (GAN), a diffusion model, or a masked autoencoder (MAE). Here, as the GAN, a co-modulated generative adversarial network (CoModGAN), a large mask inpainting (LaMa), or the like may be used. As the nucleic acid model, GLIDE, stable diffusion, RealFill, or the like may be used.

[0054] After machine learning, the learning model may be fine-tuned using a fourth image and / or a fifth image of a bone previously captured. Alternatively, after machine learning, the learning model may be fine-tuned using a sixth image of a bone (second object) of the same type as the bone of the subject captured in the first image 711. The fine-tuned learning model can generate a second image that does not look unnatural as an image of a bone. For example, PEFT (Parameter-Efficient Fine-Tuning) may be used in the fine-tuning. For example, a method such as LORA (Low-Rank Adaptation) can be used as PEFT.

[0055] The second image is a pseudo-image generated by a learning model as an image to complement the mask region. If the mask region includes a bone image, the second image is an image that includes a bone image. If the mask region does not include a bone image, the second image is an image that does not include a bone image. The second image may be, for example, an image in which the contours and textures of bones in an area adjacent to or close to the mask region match.

[0056] If the first object is a bone of the subject, the learning model may be fine-tuned using a sixth image of a bone of the same type as the bone of the subject. For example, if the first object is a lung of the subject, the learning model may be fine-tuned using a sixth image of the lung. For example, if the first object is a muscle of the subject's thigh, the learning model may be fine-tuned using a sixth image of the muscle of the thigh. For example, if the first object is an elbow joint of the subject, the learning model may be fine-tuned using a sixth image of the elbow joint.

[0057] If the first object is a bone, generation unit 103 generates a second image that is a pseudo image that does not look strange as an image of a bone corresponding to the mask region set in first image 711. If the first object is any of an organ, a muscle, and a joint, generation unit 103 generates a second image that is a pseudo image that does not look strange as an image of any of an organ, a muscle, and a joint corresponding to the mask region set in first image 711.

[0058] The learning model may be further fine-tuned using the fifth or sixth image after fine-tuning using the fifth or sixth image. The additional fine-tuning may use the same method as the fine-tuning described above, or a different method. When the first image is an image of a subject, and the subject includes a bone (first object) shown in the first image and a bone (third object) paired with the bone shown in the first image, for example, the learning model may be fine-tuned using the seventh image after fine-tuning using the fifth or sixth image. The seventh image is an image of a bone (third object) paired with the bone (first object) of the same subject as in the first image. For example, if the first object is a bone in the subject's right arm, the third object may be a bone in the same subject's left arm.

[0059] Fine tuning is also effective when the first object is something other than bone. For example, the learning model may be fine-tuned after machine learning using a fifth image previously captured of the subject's organs, muscles, joints, etc. Alternatively, the learning model may be fine-tuned after machine learning using a sixth image captured of the subject's organs, muscles, joints, etc. (second object) that are the same as the subject's organs, muscles, joints, etc. (first subject) captured in the first image 711. Furthermore, the learning model may be fine-tuned using a seventh image after fine-tuning using the fifth or sixth image. The seventh image is an image captured of the organs, muscles, joints, etc. (third object) that are paired with the organs, muscles, joints, etc. (first subject) of the same subject as those captured in the first image. For example, if the first object is the subject's right lung, the third object may be the subject's left lung. For example, if the first object is the subject's right elbow joint, the third object may be the subject's left elbow joint.

[0060] That is, after fine tuning, the learning model may be further fine-tuned using a seventh image of a third object, and the additional fine-tuned learning model can generate a second image that is more natural as an image of the subject.

[0061] Because the second image is generated by a learning model as an image that complements the mask area set in the first image, if a local condition (e.g., a disease, a foreign body, a device, etc.) exists in the mask area, the second image does not faithfully reproduce this feature. Therefore, the reference image generated using the second image faithfully reproduces the overall image of the bone shown in the first image, but does not capture the local features of the bone shown in the first image. Therefore, the detection system 100 (detection device 3) can detect the condition of the bone shown in the first image by comparing the first image with the reference image generated from the first image.

[0062] The storage unit 11 is a storage device that stores various control programs and various data, and may store image data 111. For the sake of simplicity, the control programs are not shown in the storage unit 11 shown in FIG.

[0063] The communication unit 12 transmits and receives various data between the detection device 3 and the information processing device 7. For example, the generation device 1 transmits the generated reference image to the detection device 3 via the communication unit 12. In this case, the generation device 1 may transmit the first image 711 used to generate the reference image together with the reference image to the detection device 3 via the communication unit 12.

[0064] (Processing Performed by Generating Device 1) Next, the processing performed by the generating device 1 will be described using Fig. 2 with reference to Fig. 3. Fig. 2 is a flowchart showing an example of the flow of processing performed by the generating device 1. Fig. 3 is an image diagram showing an example of processing performed by the generating device 1 (particularly, the second setting unit 102b and the generating unit 103).

[0065] First, the acquisition unit 101 acquires a first image (step S1: acquisition step).

[0066] Next, the second setting unit 102b sets one or more mask regions in the first image 711 (step S2: second setting step). Before step S2, the first setting unit 102a may set a first region corresponding to a first object in the first image 711 (first setting step). In this case, the second setting unit 102b may set one or more mask regions in the first region.

[0067] Next, the generation unit 103 generates one or more reference images using the second image generated by the learning model as an image of the mask region (step S3: generation step).

[0068] In steps S2 and S3, the generating device 1 may generate a reference image by repeating the following processes: setting one or more mask regions in the first region; generating second images that complement each of the set mask regions; and complementing the corresponding mask regions using each of the generated second images. Here, the reference image may be an image in which the entire first region in the first image 711 or the entire first image 711 is complemented with the second image. That is, the reference image may be an image in which only the first region set by segmentation in the first image is complemented with the second image. The reference image may be an image in which the entire first image is complemented with the second image. The reference region may be an image in which a mask region including at least a portion of the first region set by segmentation in the first image is complemented with the second image. With this configuration, the generating device 1 can easily generate a reference image from the first image.

[0069] FIG. 3 shows how, after mask areas M1 to M3 are set at various positions in a first image 711, second images P1 to P3 corresponding to each of the mask areas M1 to M3 are generated using a learning model, and a reference image is generated using each second image (e.g., by combining each second image).

[0070] 1 , the detection device 3 detects the condition of the bones of the subject shown in the first image based on the comparison result between the first image and the reference image. The detection device 3 includes a control unit 30, a storage unit 31, and a communication unit 32. The control unit 30 includes an acquisition unit 301, a comparison unit 302, and a detection unit 303.

[0071] The acquisition unit 301 acquires a reference image from the generation device 1. The acquisition unit 301 may acquire a first image 711 and a reference image from the generation device 1. The acquisition unit 301 may store the acquired images (for example, the first image 711 and the reference image) in image data 311 in the storage unit 31.

[0072] The comparison unit 302 compares the first image with a reference image generated from the first image by the generation device 1 and outputs the comparison result. Depending on the condition of the bone to be detected, the condition may be made clearer by changing one or more of the bit depth, gradation, saturation, and brightness of the image. Therefore, the comparison unit 302 may change one or more of the bit depth, gradation, saturation, and brightness of the first image 711 and the reference image, and compare the changed first image with the reference image.

[0073] Here, bit depth refers to the number of quantization bits constituting one pixel (picture element). In other words, bit depth refers to the amount of information per discrete signal of a digital signal (i.e., the number of quantization bits). Generally, the higher the bit depth, the better the image quality. For example, medical images have a high bit depth. However, when detecting the condition of a first object, changing (e.g., decreasing) the bit depth of the first image 711 and the reference image may more clearly indicate the condition to be detected, which may not be clearly displayed on the screen of a display device. Changing one or more of the gradation, saturation, and brightness of the first image 711 and the reference image may also more clearly indicate the condition to be detected.

[0074] The detection unit 303 detects the bone condition of the subject based on the comparison result output from the comparison unit 302. The detection unit 303 may detect the bone condition of the subject based on detection criteria 312, which will be described later. The detection unit 303 may store the detection result 313 in the storage unit 31.

[0075] Alternatively, the detection unit 303 may be configured to estimate the condition of the subject's bones from the comparison results output from the comparison unit 302 using a machine-learned estimation model. In this case, the estimation model may be machine-learned using machine-learning data that includes, as explanatory variables, the comparison results between an image showing the bones of a third subject (first object) and a reference image generated from the image, and includes, as objective variables, annotations regarding the condition of the bones of the third subject (first object). The first object is not limited to bones. For example, if the first object is the subject's organs, muscles, joints, etc., the detection unit 303 can estimate the condition of the subject's organs, muscles, joints, etc. from the comparison results output from the comparison unit 302 using the machine-learned estimation model.

[0076] The storage unit 31 is a storage device that stores various control programs and various data, and may store image data 311, detection criteria 312, and detection results 313. For the sake of simplicity, the control programs are not shown in the storage unit 31 shown in FIG.

[0077] The detection criteria 312 are criteria for determining whether the bone condition of the subject is normal. The detection criteria 312 may be set in advance for each bone condition to be detected. That is, the detection criteria 312 may include criteria for detecting a fatigue fracture and criteria for detecting a foreign object or device placed in the subject's body. The detection criteria 312 may be set based on a medical image of the fracture site of a specimen subject diagnosed with a stress fracture, or an image of a specimen subject with an implant or pacemaker placed therein, in which the implant and pacemaker are visible. Here, the specimen subject is of the same type as the subject; for example, if the subject is human, the specimen subject is also human.

[0078] The communication unit 32 transmits and receives various data between the generation device 1 and the information processing device 7. For example, the detection device 3 receives the first image 711 and the reference image from the generation device 1 via the communication unit 12. The detection device 3 also transmits the detection result 313 to the information processing device 7 via the communication unit 12. In this case, the detection device 3 may transmit, together with the detection result 313, difference information indicating the difference between the first image 711 and the reference image, which is the basis of the detection result, and an image indicating the difference to the information processing device 7 via the communication unit 12. The difference information may include information indicating an area in the first image of a disease (e.g., fracture, inflammation, tumor, etc.) detected as the condition of the first object of the subject, and information indicating an area in the first image of a foreign object, device, etc. placed on the body of the subject.

[0079] When conditions at multiple locations are detected in the first image 711 as a result of comparing the first image 711 with the reference image, the detection device 3 may transmit the detection results 313 corresponding to each of the multiple locations to the information processing device 7. For example, when the first object is a bone and a fatigue fracture is detected in an area A in the first image 711 and an implant is detected in an area B different from the area A, the detection device 3 may transmit the detection results 313 including information indicating the condition of the area A (e.g., a fatigue fracture) and information indicating the condition of the area B (e.g., an implant) to the information processing device 7.

[0080] If the first object is a bone and the detected object is a fracture, the detection result 313 may include estimated information regarding the severity of the fracture. Here, the severity of the fracture may be estimated from the comparison result output from the comparison unit 302 using a machine-learned estimation model. Such an estimation model may be machine-learned, for example, regarding the correspondence between a medical image showing a fracture site of a patient's bone and the severity of the fracture site.

[0081] (Processing Performed by Detection Apparatus 3) Next, the processing performed by the detection apparatus 3 will be described using Fig. 4 with reference to Fig. 5. Fig. 4 is a flowchart showing an example of the flow of processing performed by the detection apparatus 3. Fig. 5 is an image diagram showing an example of processing performed by the detection apparatus 3 (particularly the comparison unit 302 and the detection unit 303).

[0082] First, the acquisition unit 301 acquires the first image 711 and the reference image (step S11: acquisition step).

[0083] Next, the comparison unit 302 compares the first image 711 with the reference image and outputs the comparison result (step S12: comparison step). Here, the comparison unit 302 may compare the first image 711, in which one or more of bit depth, gradation, saturation, and brightness have been changed, with the reference image and output the comparison result.

[0084] Next, the detection unit 303 detects the condition of the bones of the subject shown in the first image 711 based on the comparison result output by the comparison unit 302 (step S13: detection step). The detection result may be transmitted to, for example, the information processing device 7, which is the source that provided the first image 711 to the detection system 100. Fig. 5 shows an image indicating the difference between the first image 711 and the reference image as the comparison result of comparing the first image 711 and the reference image, and shows how an abnormality (e.g., a fracture) has been detected as the condition of the bones shown in the first image 711.

[0085] [Embodiment 2] Another embodiment of the present disclosure will be described below. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0086] 6 is a functional block diagram showing an example of a schematic configuration of a detection system 100a according to another embodiment of the present disclosure. The detection system 100a includes a generation device 1a having a function of selecting a reference image that satisfies a predetermined criterion from a plurality of reference images generated from a single first image 711, and a detection device 3.

[0087] <Generation device 1> The generation device 1a includes a control unit 10a, a storage unit 11a, and a communication unit 12. The control unit 10a may be a CPU, for example. The control unit 10a reads a control program, which is software stored in the storage unit 11a, expands it into a memory such as a RAM, and controls each component included in the generation device 1a. The control unit 10a may include an acquisition unit 101, a first setting unit 102a, a second setting unit 102b, a generation unit 103, and a selection unit 104.

[0088] The selection unit 104 selects a reference image that satisfies a predetermined criterion from among the plurality of reference images generated by the generation unit 103. The predetermined criterion may be a picture quality assessment index 112. The picture quality assessment index may be a blind / no-reference image spatial quality assessor (BRISQUE). The image assessment criterion may be stored in the storage unit 11a.

[0089] (Processing Performed by Generating Device 1a) Next, processing performed by generating device 1a will be described with reference to Fig. 7. Fig. 7 is a flowchart showing an example of the flow of processing executed by generating device 1a.

[0090] After one or more mask areas are set in the first image, the generation unit 103 generates a plurality of reference images using the second image generated by the learning model as images of the mask areas (step S3a: generation step).

[0091] The selection unit 104 selects a reference image whose image quality evaluation index satisfies a predetermined standard from among the generated reference images (step S4: selection step).

[0092] The quality of the reference image depends on the setting of the mask area and the quality of the second image. Therefore, the reference images generated by the generation unit 103 may include a reference image that does not meet certain standards as an image to be compared with the first image 711. Therefore, the generation device 1a generates multiple reference images and selects from among them a reference image that is suitable for comparison with the first image 711. In this way, the generation device 1a can generate a reference image for accurately detecting the condition of the bones shown in the first image.

[0093] The detection device 3 can accurately detect the condition of the bones of the subject shown in the first image 711 by using the reference image generated by the generation device 1a through the processing of steps S11 to S13 shown in FIG.

[0094] [Embodiment 3] Another embodiment of the present disclosure will be described below. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0095] In the detection system 100, a reference image generated from the first image 711 by the generation device 1 may be used to generate a third image in which a predetermined state of the first object (e.g., a foreign object, a device, etc.) appearing in the first image has been removed. The third image is an image in which a second mask area is set in a portion corresponding to a state detected by comparing the first image with the reference image, and the predetermined state appearing in the first image has been removed using a second image that complements the second mask area. The following description will be given using an example in which the present invention is applied to the detection system 100, but the processing according to this embodiment can also be applied to the detection system 100a.

[0096] First, the generating device 1 (for example, the acquiring unit 101) acquires the first image 711 (step S1: acquiring step).

[0097] Next, the generating device 1 (for example, the second setting unit 102b) sets one or more mask areas in the first image 711 (step S2: second setting step).

[0098] Next, the generating device 1 (e.g., the generating unit 103) generates one or more reference images using the second image generated by the learning model as an image of the set mask region (step S3: generating step). The first image 711 and the reference image generated from the first image 711 are transmitted to the detecting device 3 via the communication unit 12 (step S5).

[0099] Next, the detection device 3 (for example, the acquisition unit 301) receives the first image 711 and the reference image from the generation device 1 (step S21).

[0100] Next, the detection device 3 (for example, the comparison unit 302) compares the first image 711 with the reference image and outputs the comparison result (step S22).

[0101] Next, the detection device 3 (e.g., the detection unit 303) transmits difference information indicating an area where a predetermined difference has been detected to the generation device 1 via the communication unit 32 based on the comparison result by the comparison unit 302 (step S23). The difference information may include information indicating an area in the first image of a disease (e.g., a fracture, inflammation, tumor, etc.) detected as the condition of the first object of the subject, and information indicating an area in the first image of a foreign object, device, etc. placed on the body of the subject.

[0102] Next, the generating device 1 (for example, the acquiring unit 301) receives the difference information from the detecting device 3 (step S6).

[0103] Next, the generating device 1 (e.g., the setting unit 102) sets a mask area in the area indicated by the difference information in the first image 711. Then, the generating device 1 (e.g., the generating unit 103) generates a third image using the second image generated by the learning model as an image of the set mask area (step S7).

[0104] The third image generated in this manner is an image obtained by removing from the first image 711 a predetermined state of the first object detected by the detection device 3, the state being captured in the first image 711. When the first object is a bone, internal organ, muscle, joint, or the like, the predetermined state to be removed may be an implant or device installed in the subject's body. When detecting the state of the first object captured in the first image (e.g., a fracture, a disease, or the like), the implant or device captured in the first image may be an obstacle. Therefore, the detection device 3 may acquire a third image generated from the first image 711 and detect the state of the first object of the subject captured in the first image based on the comparison result obtained by comparing the third image with the first image. In other words, the third image can be used as an image to be compared with the first image, and therefore can also be considered a second reference image.

[0105] The reference image is an image in which a mask area is randomly set in the entire first image or in the first region, and the set mask area is complemented with the second image. In other words, the reference image is an image in which parts unrelated to the state to be detected are complemented with the second image, and may be a coarser image than the first image. If the comparison image is a coarse image, the detection result may contain noise in the comparison result between the first image and the comparison image.

[0106] Therefore, as described above, the detection device 3 first generates a reference image, and identifies the rough position of the state to be detected based on the comparison result between the reference image and the first image. Thereafter, the detection device 3 sets a mask area for the area corresponding to the identified position, and generates a third image using the second image generated as an image of the mask area.

[0107] After the rough position is identified, a mask region may be set for the region corresponding to the identified position using a trained segmentation model, such as a segment anything model (SAM).

[0108] Returning to FIG. 8, the generating device 1 transmits the third image to the detecting device 3 (step S8).

[0109] The detection device 3 receives the third image (step S24), compares the first image 711 with the third image, and outputs the comparison result (step S25).The detection device 3 then detects the state based on the comparison result (step S26).

[0110] Other Embodiments In the detection systems 100 and 100a, the generation device 1 may be configured to have some or all of the functions of the detection device 3. Furthermore, in the detection systems 100 and 100a, the information processing device 7 may be configured to have some or all of the functions of the generation devices 1 and 1a and some or all of the functions of the detection device 3.

[0111] [Example of implementation using software] The functions of the generation device 1, 1a (hereinafter referred to as the "first device") can be realized by a program for causing a computer to function as the device, and a program for causing a computer to function as each control block of the device (particularly each part included in the control unit 10, 10a).

[0112] The functions of the detection device 3, 3a (hereinafter referred to as the "second device") can be realized by a program for causing a computer to function as the second device, and a program for causing a computer to function as each control block of the second device (particularly each part included in the control unit 30, 30a).

[0113] The functions of the information processing device 7 (hereinafter referred to as the "third device") can be realized by a program for causing a computer to function as the third device, and a program for causing a computer to function as each control block of the third device (particularly each part included in the control unit 70).

[0114] In this case, the first, second, and third devices each include a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The functions described in each of the above embodiments are realized by executing the program using the control device and storage device.

[0115] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The first, second, and third devices may or may not have these recording media. In the latter case, the program may be provided to the first, second, and third devices via any wired or wireless transmission medium.

[0116] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.

[0117] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0118] The invention according to the present disclosure has been described above based on the drawings and examples. However, the invention according to the present disclosure is not limited to the above-described embodiments. In other words, the invention according to the present disclosure can be modified in various ways within the scope of the present disclosure, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the invention according to the present disclosure. In other words, it should be noted that a person skilled in the art can easily make various modifications or corrections based on the present disclosure. It should also be noted that these modifications or corrections are included in the scope of the present disclosure.

[0119] [Summary] The generation method according to aspect 1 of the present disclosure includes an acquisition step of acquiring a first image that shows a first object, a second setting step of setting one or more mask areas in the first image, and a generation step of generating one or more reference images using second images generated by a learning model as images of the mask areas, wherein the reference images are used as objects for comparison with the first image, and the state of the first object is detected based on the comparison result between the reference image and the first image.

[0120] In a generation method according to aspect 2 of the present disclosure, in the above aspect 1, the mask region may be set by segmentation of the first image.

[0121] In a generation method according to aspect 3 of the present disclosure, in the above-mentioned aspect 1 or 2, the reference image may be an image in which the entire first image is complemented by the second image.

[0122] The generation method according to aspect 4 of the present disclosure, in any of aspects 1 to 3 above, may further include a first setting step of setting a first region in the first image corresponding to the first object, and in the second setting step, setting one or more mask regions in the first region.

[0123] A generation method according to aspect 5 of the present disclosure may be such that, in aspect 4 above, the reference image is an image in which the entire first region is complemented by the second image.

[0124] A generation method according to aspect 6 of the present disclosure is the same as in aspect 4 or 5 above, wherein the first region may be set by segmenting the first image.

[0125] A generation method according to a seventh aspect of the present disclosure is any one of the first to sixth aspects, wherein the mask region is a rectangular region.

[0126] A generation method according to aspect 8 of the present disclosure is any one of aspects 1 to 7 above, wherein the mask region may be set using a trained segmentation model.

[0127] A generation method according to aspect 9 of the present disclosure may be any of aspects 1 to 8 above, wherein the learning model is generated by learning using a generative adversarial network, a diffusion model, or a masked autoencoder.

[0128] In the generation method according to aspect 10 of the present disclosure, in the above-mentioned aspect 9, the learning model may be fine-tuned after the learning by (1) using a fourth image that is an image of the same type as the first image, and / or (2) using a fifth image that was previously photographed of the first object.

[0129] In a generation method according to aspect 11 of the present disclosure, in the above-mentioned aspect 9, the learning model may be fine-tuned after the learning using a sixth image of a second object of the same type as the first object.

[0130] A generation method according to aspect 12 of the present disclosure is such that, in aspect 10 or 11 above, the first image is an image of a subject having the first object and a third object paired with the first object, and the learning model may be further fine-tuned using a seventh image of the third object after the fine-tuning.

[0131] The generation method according to aspect 13 of the present disclosure, in any of aspects 1 to 12 above, may further include a selection step of generating a plurality of the reference images in the generation step and selecting a reference image from the plurality of reference images whose image quality evaluation index satisfies a predetermined standard.

[0132] A generation method according to aspect 14 of the present disclosure may be, in the above-mentioned aspect 13, wherein the image quality assessment index is a blind / no-reference image spatial quality assessor (BRISQUE).

[0133] A generating method according to aspect 15 of the present disclosure is any one of aspects 1 to 14 above, wherein the first object is at least a part of a subject's body, and the first image may be a medical image.

[0134] A generating method according to aspect 16 of the present disclosure may be, in the above-mentioned aspect 15, wherein the condition of the first object is a disease of the subject.

[0135] In the generating method according to aspect 17 of the present disclosure, in the above-mentioned aspect 16, the disease may be a fracture.

[0136] In the production method according to aspect 18 of the present disclosure, in the above-mentioned aspect 17, the fracture may be a fatigue fracture or a fragility fracture.

[0137] In the production method according to aspect 19 of the present disclosure, in the above-mentioned aspect 16, the disease may be a rare disease.

[0138] A generating method according to aspect 20 of the present disclosure, in the above aspect 15, may be such that the state of the first object is a foreign object or device placed on the body of the subject.

[0139] The detection method according to aspect 21 of the present disclosure includes a comparison step of comparing the first image with the reference image generated by any one of the generation methods of aspects 1 to 17 above and outputting the comparison result, and a detection step of detecting the state based on the comparison result.

[0140] In the detection method according to aspect 22 of the present disclosure, in the above-mentioned aspect 21, in the comparison step, one or more of the bit depth, gradation, saturation, and brightness of the first image and the reference image may be changed, and the changed first image may be compared with the reference image.

[0141] A generation device according to aspect 23 of the present disclosure includes an acquisition unit that acquires a first image that includes a first object, a second setting unit that sets one or more mask areas in the first image, and a generation unit that generates one or more reference images using a second image generated by a learning model as an image of the mask area, the reference images being used as a comparison target with the first image, and detecting the state of the first object based on the comparison result between the reference image and the first image.

[0142] A detection system according to aspect 24 of the present disclosure includes an acquisition unit that acquires a first image containing a first object; a second setting unit that sets one or more mask areas in the first image; a generation unit that generates one or more reference images using a second image generated by a learning model as an image of the mask area; a comparison unit that compares the first image with the reference image and outputs a comparison result; and a detection unit that detects the state of the first object based on the comparison result.

[0143] A detection system according to aspect 25 of the present disclosure may be such that, in aspect 24 above, the comparison result includes difference information indicating the area in the first image where the state is detected, the generation unit generates a third image using a second image generated by the learning model as an image of the mask area set in the area indicated by the difference information, and the detection unit detects the state of the first object appearing in the third image based on the comparison result of comparing the third image with the reference image.

[0144] The control program according to aspect 26 of the present disclosure is a control program for causing a computer to function as the generation device of aspect 23 above, and is a control program for causing a computer to function as the acquisition unit, the second setting unit, and the generation unit.

[0145] A recording medium according to aspect 27 of the present disclosure is a computer-readable non-transitory recording medium on which the control program of aspect 26 above is recorded.

[0146] 1, 1a Generation device 3 Detection device 100, 100a Detection system 101 Acquisition unit 102 Setting unit 103 Generation unit 104 Selection unit 302 Comparison unit 303 Detection unit S1 Acquisition step S2 Second setting step S3, S3a Generation step S4 Selection step S12 Comparison step S13 Detection step

Claims

1. A generation method comprising: an acquisition step of acquiring a first image showing a first object; a second setting step of setting one or more mask areas in the first image; and a generation step of generating one or more reference images using a second image generated by a learning model as an image of the mask area, wherein the reference image is used as an object for comparison with the first image, and the state of the first object is detected based on the comparison result between the reference image and the first image.

2. The method of claim 1, wherein the mask region is set by segmenting the first image.

3. The generation method according to claim 1 or 2, wherein the reference image is an image in which the entire first image is complemented by the second image.

4. A generation method according to any one of claims 1 to 3, further comprising a first setting step of setting a first region in the first image corresponding to the first object, and in the second setting step, setting one or more mask regions in the first region.

5. The generation method according to claim 4, wherein the reference image is an image in which the entire first region is complemented by the second image.

6. The generation method according to claim 4 or 5, wherein the first region is set by segmenting the first image.

7. The generation method according to claim 1, wherein the mask region is a rectangular region.

8. The generation method according to any one of claims 1 to 6, wherein the mask region is set using a trained segmentation model.

9. The generation method according to any one of claims 1 to 8, wherein the learning model is generated by training using a generative adversarial network, a diffusion model, or a masked autoencoder.

10. The generation method of claim 9, wherein, after the learning, the learning model is fine-tuned using (1) a fourth image that is the same type of image as the first image, and / or (2) a fifth image that is a previous photograph of the first object.

11. The generation method according to claim 9, wherein the learning model is fine-tuned after the learning using a sixth image of a second object of the same type as the first object.

12. The generation method described in claim 10 or 11, wherein the first image is an image of a subject comprising the first object and a third object paired with the first object, and the learning model is further fine-tuned after the fine-tuning using a seventh image of the third object.

13. A generation method according to any one of claims 1 to 12, further comprising: generating a plurality of reference images in the generation step; and selecting, from the plurality of reference images, a reference image whose image quality evaluation index satisfies a predetermined standard.

14. The method of claim 13, wherein the image quality assessment index is a blind / no-reference image spatial quality assessor (BRISQUE).

15. A generating method according to any one of claims 1 to 14, wherein the first object is a part or the whole of a subject's body, and the first image is a medical image.

16. The method of generating a signal according to claim 15, wherein the condition of the first object is a disease of the subject.

17. The method of claim 16, wherein the disease is a bone fracture.

18. The method of claim 17, wherein the fracture is a fatigue fracture or a fragility fracture.

19. The method of claim 16, wherein the disease is a rare disease.

20. The generating method according to claim 15, wherein the state of the first object is a foreign object or device placed in the body of the subject.

21. A detection method comprising: a comparison step of comparing the first image with the reference image generated by the generation method described in any one of claims 1 to 17 and outputting a comparison result; and a detection step of detecting the state based on the comparison result.

22. The detection method according to claim 21, wherein in the comparison step, one or more of bit depth, tone, saturation, and brightness of the first image and the reference image are changed, and the first image after the change is compared with the reference image.

23. A generation device comprising: an acquisition unit that acquires a first image that shows a first object; a second setting unit that sets one or more mask areas in the first image; and a generation unit that generates one or more reference images using second images generated by a learning model as images of the mask areas, wherein the reference images are used as objects for comparison with the first image, and the state of the first object is detected based on the comparison result between the reference image and the first image.

24. A detection system comprising: an acquisition unit that acquires a first image that shows a first object; a second setting unit that sets one or more mask areas in the first image; a generation unit that generates one or more reference images using a second image generated by a learning model as an image of the mask area; a comparison unit that compares the first image with the reference image and outputs a comparison result; and a detection unit that detects the state of the first object based on the comparison result.

25. The detection system described in claim 24, wherein the comparison result includes difference information indicating the area in the first image where the state is detected, the generation unit generates a third image using the second image generated by the learning model as an image of the mask area set in the area indicated by the difference information, and the detection unit detects the state of the first object appearing in the third image based on the comparison result between the third image and the reference image.

26. A control program for causing a computer to function as the generating device according to claim 23, the control program causing a computer to function as the acquisition unit, the second setting unit, and the generating unit.

27. A computer-readable non-transitory recording medium on which the control program according to claim 26 is recorded.

Citation Information

Patent Citations

  • Multivariate and multi-resolution retinal image anomaly detection system

    JP2020032190A

  • Visual inspection device, visual inspection system, feature quantity estimation device, and visual inspection program

    JP2021092887A

  • Learning data generating apparatus, method and program, and object detecting apparatus, method and program

    JP2023144382A

  • Apparatus and method for processing medical image

    US20200226752A1

  • Image processing device, image processing method, and recording medium

    WO2019159853A1