Image analysis method, learning image or analysis image generation method, learning model completion method, image analysis device, and image analysis program

By generating and analyzing composite images of multiple images that are continuous in time or space, and utilizing channel allocation and learning models, the accuracy and efficiency problems of dynamic image and 3D data analysis in existing technologies are solved, and high-precision motion and attribute inference is achieved.

CN115943430BActive Publication Date: 2026-04-28KOWA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KOWA CO LTD
Filing Date
2021-06-24
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies, when using neural networks to analyze dynamic images or 3D data, cannot effectively capture subtle motion features beyond joint positions, resulting in insufficient accuracy in motion category prediction, and the large amount of data makes learning and processing difficult.

Method used

By acquiring multiple images that are continuous in time or space, assigning different channels and grayscale information, generating synthetic images, and using a complete learning model for inference, the accuracy of action or attribute prediction is improved by combining location information and multimodal learning.

Benefits of technology

It achieves high-precision and high-speed analysis of dynamic images and 3D data, effectively captures subtle motion features, improves the accuracy of motion category inference, and reduces the hardware processing burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115943430B_ABST
    Figure CN115943430B_ABST
Patent Text Reader

Abstract

In order to accurately estimate the motion or the attribute of an object from a plurality of images having continuity in time or space, an image acquisition unit acquires a plurality of images having continuity in time or space; a channel allocation unit allocates mutually different channels to at least one of the gradation information of colors and / or the gradation information of brightness that can be acquired from each of the plurality of images in accordance with a predetermined rule; a composite image generation unit generates one composite image in which at least a part of the gradation information of each image can be recognized by the channels by extracting the gradation information of the channels from each of the plurality of images and performing composition; and an inference unit analyzes the composite image and infers the plurality of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an image analysis method associated with multiple images that are continuous in time or space, a method for generating learning images or analysis images, a method for generating a completed learning model, an image analysis apparatus, and an image analysis program. Background Technology

[0002] Previously, there was a demand for computers to perform analytical processing of moving images or 3D data acquired in various situations, extracting desired features. Furthermore, there is an increasing demand for performing such analysis through so-called artificial intelligence (AI), which involves enabling neural networks to learn the extraction of desired features and then using the learned model for execution. However, compared to 2D image data, moving image or 3D data has a massive volume. Directly inputting this data into a neural network to perform learning or actual analysis is, from the perspective of learning convergence and hardware processing capabilities, not easy at present.

[0003] In contrast, for example, Patent Document 1 proposes a scheme to perform analysis processing using a synthetic image synthesized from multiple frames of images constituting a dynamic image.

[0004] Patent document 1 discloses the following technology: superimposing multiple images obtained by continuous photography of the human body to generate a synthetic image, and using a learned convolutional neural network (CNN) to analyze the synthetic image to determine the joint position of the human body.

[0005] Existing technical documents

[0006] Patent documents

[0007] Patent document 1: JP 2019-003565. Summary of the Invention

[0008] The problem the invention aims to solve

[0009] According to the image analysis method described in Patent Document 1, information from multiple frames can be included in a single composite image. However, in the technology of Patent Document 1, since the composite image is generated by simply adding the brightness values ​​of multiple images, the composite image does not contain information for understanding the temporal relationship of the actions of objects contained in the image. Therefore, the following problems exist: it cannot capture subtle motion features other than joint positions, and the accuracy of inferring the action category based on the composite image is insufficient.

[0010] The present invention was made in view of the above-mentioned problems, and its purpose is to provide an image analysis method that can infer the action or attributes of an object with good accuracy and high speed based on multiple images that are continuous in time or space, a method for generating learning images or analysis images, a method for generating a learning model, an image analysis device, and an image analysis program.

[0011] means for solving problems

[0012] The image analysis method of the present invention is characterized by comprising: an image acquisition step, acquiring multiple images that are continuous in time or space; a channel allocation step, for the multiple images, allocating at least a portion of the grayscale information of color and / or the grayscale information of brightness that can be obtained from each image to different channels according to a predetermined rule; a composite image generation step, generating a composite image by extracting the grayscale information allocated to the channels from each of the multiple images and synthesizing them, thereby generating a composite image that can identify at least a portion of the grayscale information of each image through the channels; and an inference step, analyzing the composite image and inferring from the multiple images.

[0013] Furthermore, in the image analysis method of the present invention, it is characterized in that, in the channel allocation step, colors with different hues are allocated as channels to each of the plurality of images, and grayscale information corresponding to the allocated colors is used as grayscale information of the allocated channels. In the composite image generation step, by extracting the grayscale information of the allocated channels from each of the plurality of images and performing synthesis, a color composite image is generated by synthesizing grayscale information corresponding to colors with different hues.

[0014] Furthermore, in the image analysis method of the present invention, it is characterized in that, in the inference step, the completed learning model is input to the synthetic image generated in the synthetic image generation step, and the output of the completed learning model is obtained as the inference result. The completed learning model is a model that has been pre-machined based on multiple synthetic images generated from multiple sample images used for learning.

[0015] Furthermore, in the image analysis method of the present invention, the plurality of images are further characterized in that they are composed of multiple images extracted from a dynamic image at a certain time interval, the dynamic image being composed of multiple images that are continuous in time obtained by capturing dynamic images of a specified object performing an action, and in the inference step, an inference related to the action pattern of the specified object is performed based on the synthetic image generated based on the dynamic image.

[0016] Furthermore, the image analysis method of the present invention is characterized by further including a position information acquisition step for acquiring position information of the object when acquiring multiple images in the image acquisition step. When multiple objects are used as analysis objects, in the image acquisition step, the multiple images are acquired for each of the multiple objects; in the position information acquisition step, the position information is acquired for each of the multiple objects; in the channel allocation step, a channel is allocated for each of the acquired multiple images for each object; in the composite image generation step, the composite image is generated for each object; and in the inference step, the multiple composite images generated for each object and the position information of each of the multiple objects are used as input to perform inference related to the action patterns of the multiple objects.

[0017] Furthermore, in the image analysis method of the present invention, the plurality of images are further characterized by being a plurality of tomographic images that are continuous in a specific direction when a three-dimensional region is represented by the stacking of a plurality of acquired tomographic images in a specific direction, or being a plurality of tomographic images extracted from a three-dimensional model in a manner that are continuous in a specific direction when a three-dimensional region is represented by a three-dimensional model that can arbitrarily extract tomographic images. In the inference step, inference related to the three-dimensional region is performed based on the synthesized image.

[0018] The method for generating learning or analytical images according to the present invention is characterized by comprising: an image acquisition step, acquiring multiple images that are continuous in time or space; a channel allocation step, allocating different channels to at least a portion of the grayscale information of color and / or the grayscale information of brightness that can be obtained from each of the multiple images according to a predetermined rule; and a composite image generation step, generating a composite image by extracting the grayscale information allocated to the channels from each of the multiple images and synthesizing them, thereby generating a composite image that can identify at least a portion of the grayscale information of each image through the channels.

[0019] The method for generating a learning model according to the present invention is characterized by comprising: an image acquisition step, acquiring multiple images that are continuous in time or space; a channel allocation step, for the multiple images, allocating at least a portion of the grayscale information of color and / or grayscale information of brightness that can be obtained from each image to mutually different channels according to a predetermined rule; a synthetic image generation step, generating a synthetic image by extracting the grayscale information of the allocated channels from each of the multiple images and synthesizing them, thereby generating a synthetic image that can identify at least a portion of the grayscale information of each image through the channels; a correct answer data acquisition step, acquiring correct answer data when performing inference on the synthetic image; an inference step, inputting the synthetic image into a model composed of a neural network and performing inference to output an inference result; and a parameter update step, updating the parameters of the model using the inference result and the correct answer data.

[0020] The image analysis apparatus of the present invention is characterized by comprising: an image acquisition unit that acquires a plurality of images that are continuous in time or space; a channel allocation unit that, for the plurality of images, allocates at least a portion of grayscale information of color and / or grayscale information of brightness that can be acquired from each image to different channels according to a predetermined rule; a composite image generation unit that generates a composite image by extracting the grayscale information allocated to the channels from each of the plurality of images and synthesizing them, thereby generating a composite image in which at least a portion of the grayscale information of each image can be identified through the channels; and an inference unit that analyzes the composite image and infers from the plurality of images.

[0021] The image analysis program of the present invention is characterized by enabling a computer to perform: an image acquisition function, acquiring multiple images that are continuous in time or space; a channel allocation function, allocating different channels to at least a portion of the grayscale information of color and / or grayscale information of brightness that can be obtained from each of the multiple images according to a predetermined rule; a composite image generation function, generating a composite image by extracting the grayscale information allocated to the channels from each of the multiple images and synthesizing them, thereby generating a composite image that can identify at least a portion of the grayscale information of each image through the channels; and an inference function, analyzing the composite image and making inferences about the multiple images.

[0022] The effects of the invention

[0023] The implementation methods of this application address one or more shortcomings. Attached Figure Description

[0024] Figure 1 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention.

[0025] Figure 2 This is a block diagram illustrating an example of the environment in which an image analysis apparatus corresponding to at least one embodiment of the present invention is applicable.

[0026] Figure 3 This is an explanatory diagram illustrating the concept of synthetic image generation by an image analysis apparatus corresponding to at least one embodiment of the present invention.

[0027] Figure 4 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention.

[0028] Figure 5 This is a flowchart illustrating an example of a learning process corresponding to at least one embodiment of the present invention.

[0029] Figure 6 This is an explanatory diagram illustrating the generation of a synthetic image when an image analysis device 10A is used to analyze the movements of mice in a forced swimming experiment.

[0030] Figure 7 It is a more illustrative representation Figure 6 The diagram illustrates the generation of the synthesized image.

[0031] Figure 8 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention.

[0032] Figure 9 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention.

[0033] Figure 10 This is an illustrative diagram illustrating examples of the social behavior of rats used as subjects in a Social Interaction Test.

[0034] Figure 11 This is an illustrative diagram illustrating an example of a method for extracting the location of mice that are the subjects of analysis in a social interaction experiment.

[0035] Figure 12 This is an illustration showing the generation of a synthetic image for a single mouse.

[0036] Figure 13 This is a block diagram illustrating an example of the overall structure of a model used to implement multi-modal learning.

[0037] Figure 14This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention.

[0038] Figure 15 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention.

[0039] Figure 16 This is an explanatory diagram illustrating the process of determining vascular regions using image analysis processing performed by image analysis device 10C based on voxel data acquired when the fundus is captured by OCT (Optical Coherence Tomography Apparatus).

[0040] Figure 17 This is an explanatory diagram illustrating another method for generating synthetic images using an image analysis apparatus corresponding to at least one embodiment of the present invention.

[0041] Figure 18 This is an explanatory diagram showing the relationship between each color and RGB values ​​when the channels related to grayscale information are used in a 6-color manner.

[0042] Figure 19 This is an illustrative diagram showing an example of a composite image generated through overwriting.

[0043] Figure 20 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention. Detailed Implementation

[0044] Hereinafter, examples of embodiments of the present invention will be described with reference to the accompanying drawings. Furthermore, the various constituent elements in the examples of each embodiment described below can be appropriately combined without causing contradictions. Additionally, there are cases where descriptions given as examples of a particular embodiment may be omitted in other embodiments. Furthermore, there are cases where actions or processes unrelated to the feature portions of each embodiment may be omitted. Moreover, the order of various processes constituting the various flows described below may differ without causing contradictions in the processing content.

[0045] [First Implementation Method]

[0046] Hereinafter, an example of an image analysis apparatus according to a first embodiment of the present invention will be described with reference to the accompanying drawings. Figure 1 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention. For example... Figure 1As shown, an image analysis device 10A, as an example of an image analysis device 10, includes an image acquisition unit 11, a channel allocation unit 12, a composite image generation unit 13, an inference unit 14, and a storage unit 15. Alternatively, the image analysis device 10 may be a device designed as a dedicated machine, but it can be a device that can be implemented by a conventional computer. That is, the image analysis device 10 at least has a CPU (Central Processing Unit) and memory, which are typically found in conventional computers. It may also have a GPU (Graphics Processing Unit). Furthermore, it may be configured to connect via a bus to input devices such as a mouse and keyboard, output devices such as a printer, or communication devices for connecting to a communication network. The processing in each unit of the image analysis device 10 is achieved by reading a program from memory to execute the processing in each unit and executing it in the CPU or GPU, which functions as a processing circuit. In other words, the configuration is such that, through the execution of this program, the processor (processing circuit) can execute the processing of each device.

[0047] Figure 2 This is a block diagram illustrating an example of the environment in which an image analysis apparatus corresponding to at least one embodiment of the present invention is applicable. Figure 2 In this system, a server device and multiple terminal devices are interconnected via a communication network. For example, it can also be made... Figure 2 The server device functions as the image analysis device 10, and can be accessed from any of a plurality of terminal devices via a communication network. Alternatively, the terminal device may have a program for utilizing the image analysis device installed on it, or the program on the server may be accessed via a browser. Furthermore, for example, using... Figure 2 The terminal device functions as the image analysis device 10. However, in this case, it can also be configured such that a portion of the functions of the image analysis device 10 are possessed by the server device, and that portion of the functions are utilized by accessing the server device from the terminal device via a communication network.

[0048] Furthermore, the constituent elements of the image analysis apparatus 10 described below do not need to all be possessed by a single device. It can also be configured such that some components are possessed by other devices. For example, a configuration could be provided by a server device and any one of multiple terminal devices that can be connected via a communication network, and the image analysis apparatus 10 could utilize the configurations of the other devices while communicating. Additionally, the server device is not limited to a single unit; multiple server devices could be used. Furthermore, the completed learning model, described later, can be stored within the device itself that functions as the image analysis apparatus 10, or it can be distributed among other devices such as the server device and multiple terminal devices, or it can be used each time it is connected to a device with the completed learning model to be used via a communication network. In other words, as long as the completed learning model stored in certain storage units can be used, it is acceptable whether the storage units for the completed learning model are located within the image analysis apparatus itself or in other devices.

[0049] The image acquisition unit 11 has the function of acquiring multiple images that are continuous in time or space. Here, "multiple images that are continuous in time" refers to multiple images acquired sequentially in time, such as multiple images acquired from moving images according to predetermined rules. "Multiple images that are continuous in space" refers to multiple images obtained by acquiring tomographic images on each of multiple predetermined planes in a manner where multiple predetermined planes are parallel and continuous in one direction, when the information of the three-dimensional space on the plane intersects a predetermined range of three-dimensional space and is called a tomographic image. For example, multiple images acquired from voxel data representing a predetermined three-dimensional region obtained using OCT (Optical Coherence Tomography) or the like according to predetermined rules. Furthermore, "multiple images that are continuous in space" can be multiple tomographic images that are continuous in a specific direction when a three-dimensional region is represented by stacking multiple acquired tomographic images in a specific direction, or multiple tomographic images extracted from a three-dimensional model in a manner continuous in a specific direction when a three-dimensional region is represented by a three-dimensional model from which tomographic images can be arbitrarily extracted.

[0050] Furthermore, multiple images only need to be continuous in time or space; they do not need to be acquired consecutively. For example, it is not required to select consecutive frames when shooting 60fps video. For instance, acquiring one frame every 15 frames, or four frames per second, can also be considered as having temporal continuity.

[0051] The channel allocation unit 12 has the following function: for multiple images, it allocates at least a portion of the grayscale information of colors and / or grayscale information of brightness that can be obtained from each image to different channels according to a predetermined rule. Here, a channel refers to identification information allocated in the case of synthesizing multiple images for recognizing the grayscale information of colors and / or grayscale information of brightness (brightness information) that can be obtained from each image with other images. Any channel can be set as long as the grayscale information of each image can be recognized. For example, different colors with different hues can be allocated as channels for each of the multiple images, or the grayscale information corresponding to the allocated color can be used as the grayscale information for recognizing that image. As a specific example, consider acquiring three images and allocating RGB colors as channels for each of the three images. In each image, only one color of RGB is processed as the grayscale information for which a channel has been allocated.

[0052] The composite image generation unit 13 has the following function: by extracting grayscale information allocated to channels from each of multiple images and combining them, a composite image is generated that allows at least a portion of the grayscale information of each image to be identified by channels. The image synthesis method here can also vary depending on the type of channels, etc. For example, when acquiring three images and allocating the three colors of RGB as channels for each of the three images, extracting grayscale information of only one color of RGB from each image is similar to generating a color image based on the grayscale information of RGB within the same image; a color composite image is generated based on the grayscale information of RGB obtained from the three images.

[0053] The inference unit 14 has the following functions: analyzing synthetic images and performing inferences on multiple images. The content of the inference varies depending on the object being processed, and various inference methods can be employed. In the inference unit 14, inferences related to image analysis are performed to obtain inference results. Alternatively, the inference processing of the inference unit 14 can be performed based on a completed learning model obtained through prior learning. The learning processing of the completed learning model can, for example, be performed by combining the synthetic image used for learning with the correct answer data of the inference from that synthetic image. The completed learning model can be any model learned through machine learning; various models are applicable, such as neural networks learned through deep learning. Furthermore, as an example, existing completed learning convolutional neural networks (CNNs) such as ResNet or VGG can be used, and structures with additional learning (transfer learning) can also be employed as needed. The method of learning a completed learning model from scratch by preparing multiple synthetic images related to the inference object has the following advantages: obtaining a completed learning model capable of performing high-precision inferences that closely match the tendencies of the synthetic images used for learning. On the other hand, the method of using existing complete learning models has the following advantages: even if there is no time to learn from scratch, inference processing such as classification problems can be performed immediately.

[0054] The storage unit 15 has the following functions: storing information required for processing by each unit in the image analysis apparatus 10A, and storing various types of information generated during the processing of each unit. Additionally, the storage unit 15 may also store the completed learning model. Alternatively, the completed learning model may be stored in a server device that can be connected via a communication network, thereby enabling the server device to perform the functions of the inference unit 14.

[0055] Figure 3 This is an explanatory diagram illustrating the concept of synthetic image generation in an image analysis apparatus corresponding to at least one embodiment of the present invention. For example... Figure 3 As shown, in this embodiment, multiple images, such as three images, are extracted from data that are continuous in time or space, such as dynamic images or voxel data, at a predetermined interval. A composite image is generated based on the allocation of channels for each image, and inferences for image analysis are performed using the composite image.

[0056] Next, the process of image analysis processing corresponding to at least one embodiment of the present invention will be described. Figure 4 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention. Figure 4In this process, image analysis processing begins in the image analysis device 10A by acquiring multiple images that are sequential in time or space (step S101). Next, the image analysis device 10A assigns different channels to the acquired multiple images (step S102). Then, the image analysis device 10A extracts the grayscale information of the assigned channels from each of the multiple images and combines them into a single composite image (step S103). Finally, the image analysis device 10A performs inference based on the composite image, obtains inference results related to image analysis (step S104), and the image analysis processing ends.

[0057] The inference in the image analysis process described above can be any type of processing, but when using a fully learned model for inference, there may be situations where pre-learning is required. Therefore, the learning process corresponding to at least one embodiment of the present invention will be described using the case where the learning object is a model composed of a neural network as an example. Figure 5 This is a flowchart illustrating an example of a learning process corresponding to at least one embodiment of the present invention. Figure 5 In the image analysis device 10A, the learning process begins by acquiring multiple images that are sequential in time or space, as data for learning (step S201). Next, the image analysis device 10A assigns different channels to the acquired multiple images (step S202). Then, the image analysis device 10A extracts the grayscale information of the assigned channels from each of the multiple images and combines them into a single image to generate a synthetic image (step S203). Additionally, the image analysis device 10A acquires correct answer data related to the generated synthetic image (step S204). Correct answer data represents the correct answer related to the inference and is typically annotated by a human hand. Then, the image analysis device 10A inputs the synthetic image into the neural network of the learning object to perform inference and acquire inference results related to image analysis (step S205). Finally, the parameters of the neural network are updated using the obtained inference results and correct answer data, and the learning process ends. Figure 5 The flowchart shown illustrates the learning process up to updating the neural network parameters once based on a single synthesized image. However, in practice, to improve inference accuracy, the parameters need to be updated sequentially based on multiple synthesized images. Parameter updates can be approached in ways such as calculating the loss of the inference result based on a loss function, and updating the parameters in the direction that minimizes the loss.

[0058] Here, a specific example of image analysis performed using the image analysis apparatus 10A according to the first embodiment of the present invention will be described. The image analysis apparatus 10A of the first embodiment of the present invention is applicable to multiple images that are continuous in time, that is, images captured as moving images, and can be said to be applicable to the state of the captured object moving in response to the passage of time. Specifically, the image analysis apparatus 10A can be applied to analyze moving images captured of moving animals such as mice.

[0059] As experiments concerning mouse movement, there are experiments such as the forced swimming test. This experiment is conducted, for example, to investigate the effects of drugs used for depression or schizophrenia, by administering medication to mice as an investigation into the drug's effects. The duration of active behavior (mobility) and immobility (time without movement) in the forced swimming test are used to determine whether the drug's effects cause a decrease in the mice's activity. For example, the length of immobility is used as an indicator of drug efficacy. A challenge in the movement analysis of the forced swimming test is that, for example, when a mouse is seen splashing shallowly with its hind legs (one leg) against the wall of the enclosure, although the mouse is active, it needs to be identified as immobile due to a lack of activity. Determining that a mouse is still immobile despite such activity is an error in computer analysis based solely on the presence or absence of movement. Therefore, for high-precision analysis related to the movement of mice in the forced swimming test, an image analysis device 10A is used.

[0060] Figure 6 This diagram illustrates the generation of a composite image using the image analysis device 10A for motion analysis of mice in a forced swimming experiment. Three frames are extracted at predetermined intervals from motion footage obtained from the mice in the forced swimming experiment, and RGB channels are assigned to each of the three extracted frames. Grayscale information based on the assigned color is extracted from each image. The three images are then synthesized to obtain a composite image: one with only R grayscale information, one with only G grayscale information, and one with only B grayscale information. Inference is performed using this synthesized image.

[0061] Figure 7 It is a more illustrative representation Figure 6 The diagram illustrates the generation of the synthesized image. Figure 7In the three images, the elliptical portion corresponding to the mouse's torso remains stationary, while only the rectangular portion corresponding to the mouse's feet moves and changes. When RGB channels are assigned to these three images, and the grayscale information of the assigned colors is extracted to generate a composite image, the elliptical portion, having complete RGB grayscale information, is composited and thus represented in the same way as the original image. However, the rectangular portion, representing only its individual grayscale information, is represented separately using RGB monochrome (at least in a different color than the original). In this way, even when there are moving parts among the selected three images, the moving parts can be identified by visual inspection of the composite image.

[0062] In this way, for various movements of mice in a forced swimming experiment, a completion learning model is obtained by using synthetic images and correct answer data at the time, for example, to learn a neural network. By using the completion learning model, it is possible to properly determine, for example, the action of only the hind limbs making shallow splashes (one foot) at the wall as immobility time.

[0063] As described above, as one aspect of the first embodiment, it includes: an image acquisition unit that acquires multiple images that are continuous in time or space; a channel allocation unit that, for the multiple images, allocates at least a portion of the grayscale information of color and / or grayscale information of brightness that can be acquired from each image to different channels according to a predetermined rule; a composite image generation unit that extracts the grayscale information allocated to each channel from each of the multiple images and synthesizes them to generate a composite image in which at least a portion of the grayscale information of each image can be identified by the channels; and an inference unit that analyzes the composite image and infers from the multiple images; therefore, it is possible to accurately infer the action or attributes of an object based on multiple images that are continuous in time or space.

[0064] In other words, the synthetic image used for inference differs from dynamic images or 3D data, which are multiple images that are continuous in time or space. It is a 2D image, therefore, the convergence of learning can be expected when input into a neural network for learning processing or actual analysis processing. It can be said that the hardware processing power can be fully utilized using computers currently available on the market. Since the synthetic image is 2D and contains information from multiple images, high-precision inference can be performed using motion or spatial correlation information of objects that cannot be obtained from a single image.

[0065] [Second Implementation]

[0066] Hereinafter, an example of an image analysis apparatus according to a second embodiment of the present invention will be described with reference to the accompanying drawings. In this second embodiment, multiple objects are envisioned as the objects of analysis, and the positional relationship of these multiple objects also affects the analysis results. The image analysis apparatus of the present invention is applicable to this situation, and will be described in this context. Specifically, the example of applying the image analysis apparatus of the present invention to the analysis of the social behavior of multiple mice will be described.

[0067] Figure 8 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention. For example... Figure 8 As shown, the image analysis device 10B, as an example of the image analysis device 10, includes an image acquisition unit 11, a channel allocation unit 12, a composite image generation unit 13, a position information acquisition unit 16, an inference unit 14, and a storage unit 15. Regarding the configuration given the same reference numerals as in the first embodiment, it has the same functions as the configuration described in the first embodiment; detailed descriptions of the same functions are omitted. In this second embodiment, a situation is envisioned where multiple objects are simultaneously photographed. In this case, image acquisition and composite image generation are performed for each object, which differs from the first embodiment.

[0068] The image acquisition unit 11 acquires multiple images that are continuous in time or space, but also has the following functions: determining the location of multiple objects contained in each acquired image, and acquiring images of the specified range of each object from each image.

[0069] The channel allocation unit 12 has the following functions: For each object, it allocates a channel to each of the acquired multiple images; for each object, it acquires multiple images of a defined range in which the object appears, and allocates a channel to each of the multiple images for each object.

[0070] The composite image generation unit 13 has the function of generating a composite image for each object. When two objects appear in a moving image, two composite images are generated.

[0071] The position information acquisition unit 16 has the function of acquiring position information of objects when acquiring multiple images. Although multiple images are acquired from dynamic images obtained by capturing multiple objects, the position information of each object in each image is acquired. The position information may be, for example, coordinate data. When acquiring multiple images, the position information of objects in all images may be acquired, or the position information in the first and last images may be acquired, or at least the position information of objects in one image may be acquired.

[0072] The inference unit 14 has the following functions: taking into input multiple synthetic images generated for each object and the positional information of each object, it performs inference related to the action patterns of the multiple objects. Inference can be performed using various methods; as an example, one method uses a pre-learned neural network model for inference. In employing this method, to implement a model that uses different elements such as synthetic images of multiple objects and the positional information of each object as input, as in this example, it is preferable to construct a neural network for performing multimodal learning.

[0073] Next, the process of image analysis processing corresponding to at least one embodiment of the present invention will be described. Figure 9 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention. Figure 9 In this process, image analysis processing begins in the image analysis device 10B by extracting multiple frames that are sequential in time from dynamic images obtained by capturing multiple objects (step S301). Next, the image analysis device 10B acquires multiple images for each object by extracting images corresponding to the regions of each object from each frame (step S302). Next, the position information of each object in at least one frame is acquired (step S303). Next, the image analysis device 10B assigns different channels to the acquired multiple images for each object (step S304). Next, the image analysis device 10B extracts the grayscale information of the assigned channels from each of the multiple images for each object and combines them into a composite image (step S305). Then, the image analysis device 10B performs inference based on the composite image generated for each object and the position information of each object, acquires the inference result related to image analysis (step S306), and the image analysis processing ends.

[0074] Here, a specific example of image analysis performed using the image analysis apparatus 10B according to the second embodiment of the present invention will be described. The image analysis apparatus 10B of the second embodiment of the present invention is applicable to a situation where multiple images are continuous in time, that is, a situation where multiple objects are included as objects to be captured as moving images. Specifically, the image analysis apparatus 10B can be applied to the case of analyzing the social behavior of mice when multiple mice are placed together in the same cage for behavioral observation.

[0075] As an experiment to investigate the sociality of rats, there are social interaction experiments. These experiments involve placing two rats in the same cage and observing how much social behavior (sociability) they exhibit within a specified time. For example, there are cases where rats are given medication to investigate the effects of drugs for depression or schizophrenia, and social interaction experiments are conducted as an experiment to investigate the effects of the drug. Rats exhibiting symptoms similar to depression or schizophrenia tend to show reduced social behavior; therefore, social interaction experiments are used to analyze social behavior in order to determine the efficacy of medications.

[0076] Figure 10 This is an illustrative diagram representing examples of the social behavior of rats used as subjects of analysis in a social interaction experiment. (Example:) Figure 10 As shown, social behaviors of mice include sniffing, following, and grooming, which are determined by the shift in behavior of one mouse relative to the other. Even if the positional relationship or orientation of two mice in a two-dimensional image resembles "sniffing," there may be instances where this positional relationship is only occasionally established without actual sniffing. Therefore, it is difficult to accurately determine social behaviors using only still images. Thus, image analysis device 10B is suitable for dynamic images obtained from filming social interaction experiments.

[0077] Figure 11 This is an illustrative diagram illustrating an example of a method for extracting the location of mice that are the subjects of analysis in a social interaction experiment. Figure 11 The image shown is (A), a photograph taken from above of a cage containing two mice, and (B), a background image of the same area without mice. The difference between these two images is then calculated and binarized to obtain (C), a binary image. The white areas in the binary image (C) represent the positions of the mice. Positional information can also be extracted from such binary images.

[0078] Figure 12 This is an illustrative diagram showing the generation of a synthetic image of a mouse. (Through...) Figure 11 The method shown extracts the positions of the two mice by extracting only the regions containing the mice as images. When extracting the regions containing the two mice from each of the three frames, for each mouse, the following steps are taken: Figure 12The three images are shown. For example, RGB channels are assigned to these three images, and only the grayscale information of the assigned color is extracted from each image. The three images become images with only R grayscale information, images with only G grayscale information, and images with only B grayscale information, respectively. A composite image is obtained by combining them. Such a composite image is generated for each mouse.

[0079] Figure 13 This is a block diagram illustrating an example of the overall structure of a model used for multimodal learning. As input data, synthetic images of mouse 1 and mouse 2, along with their location information, are prepared, and a neural network portion is set up to input each of these data points. For the neural network portion that inputs the synthetic images, a pre-learned neural network such as VGG16 can be used. Then, after the three neural network portions corresponding to each input, a combination layer and a classification layer are set up to obtain the inference result. The inference can be configured, for example, as an example of social behavior, based on... Figure 10 The four behaviors shown are categorized and output according to their corresponding relationships.

[0080] As described above, as one aspect of the second embodiment, there is also a position information acquisition unit that acquires position information of objects when the image acquisition unit acquires multiple images. When multiple objects are the analysis objects, the image acquisition unit acquires multiple images for each of the multiple objects, the position information acquisition unit acquires position information for each of the multiple objects, the channel allocation unit allocates a channel for each of the acquired multiple images for each object, the composite image generation unit generates a composite image for each object, and the inference unit takes the multiple composite images generated for each object and the position information of each of the multiple objects as inputs to perform inference related to the action patterns of the multiple objects. Therefore, based on multiple images that are sequential in time and related to the multiple objects, it is possible to infer the action or attributes of the objects with good accuracy.

[0081] [Third Implementation Method]

[0082] Hereinafter, an example of an image analysis apparatus according to a third embodiment of the present invention will be described with reference to the accompanying drawings. In this third embodiment, the image analysis apparatus is applicable to three-dimensional data of multiple images that are spatially continuous.

[0083] Figure 14 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention. For example... Figure 14As shown, the image analysis apparatus 10C, as an example of an image analysis apparatus 10, includes an image acquisition unit 11, a channel allocation unit 12, a composite image generation unit 13, a region segmentation unit 17, an inference unit 14, and a storage unit 15. The configurations, which are given the same reference numerals as in the first embodiment, have the same functions as those described in the first embodiment, and detailed descriptions of these same functions are omitted.

[0084] The region segmentation unit 17 has the following function: it segments the synthesized image into multiple regions of a predetermined size. The size of the region is preferably determined according to the features that are to be determined through image analysis. Each of the segments becomes the object of inference by the inference unit 14.

[0085] Next, the process of image analysis processing corresponding to at least one embodiment of the present invention will be described. Figure 15 This is a flowchart illustrating an example of an image analysis processing flow corresponding to at least one embodiment of the present invention. Figure 15 In this process, image analysis processing begins in the image analysis device 10C by acquiring multiple images that are spatially continuous (step S401). Next, the image analysis device 10C assigns different channels to the acquired multiple images (step S402). Then, the image analysis device 10C extracts the grayscale information of the assigned channels from each of the multiple images and combines them into a single composite image (step S403). Next, the image analysis device 10C segments the composite image into multiple regions (step S404). Then, the image analysis device 10C performs inference for each region of the composite image and acquires an inference result related to image analysis for each region (step S405). Inference is performed sequentially by switching regions, and the image analysis processing ends when inference results have been acquired for all regions of the composite image.

[0086] Here, a specific example of image analysis performed using the image analysis apparatus 10C according to the third embodiment of the present invention will be described. The image analysis apparatus 10C of the third embodiment of the present invention is applicable to the analysis and processing of multiple images that are spatially continuous, that is, tomographic images and the like acquired in parallel and continuously based on three-dimensional data. Specifically, the image analysis apparatus 10C can be applied to the analysis of voxel data representing a specified three-dimensional region obtained using OCT (Optical Coherence Tomography) or the like.

[0087] Figure 16 This is an explanatory diagram illustrating the process of image analysis processing, performed by the image analysis device 10C, to determine vascular regions based on voxel data, where the voxel data is obtained by imaging the fundus using OCT (Optical Coherence Tomography). First, as... Figure 16 As shown, voxel data is converted into raster data of, for example, 300 slices. Three slices are extracted from the 300-slice raster data at a time, and a composite image is generated based on the images of these three slices. The composite image is then segmented into multiple regions. The size of these regions is preferably suitable for appropriately determining the size of the vascular regions. Then, inference is performed for each region, and an inference result is obtained for all regions to determine whether they correspond to vascular regions, thereby identifying the vascular regions in the composite image. The same process is repeated for other combinations of raster data. Finally, at the stage where the inference process for all composite images is completed, volumetric data is reconstructed using the finally determined vascular region groups, thereby obtaining a 3D vascular model corresponding to the voxel data.

[0088] Alternatively, the vascular region can be determined for each grid data. Compared to performing image analysis on each slice, performing image analysis after generating a composite image using multiple slices, such as three slices, makes it easier to grasp the spatial extension direction of blood vessels. Therefore, performing image analysis after forming a composite image can improve the accuracy of vascular region determination.

[0089] As described above, as one aspect of the third embodiment, the multiple images are multiple tomographic images that are continuous in a specific direction when a three-dimensional region is represented by stacking multiple acquired tomographic images in a specific direction, or multiple tomographic images extracted from a three-dimensional model in a manner that are continuous in a specific direction when a three-dimensional region is represented by a three-dimensional model from which tomographic images can be arbitrarily extracted. In the inference unit, inference related to the three-dimensional region is performed based on the synthesized image. Therefore, the characteristics, properties, etc. of the three-dimensional region can be inferred with good accuracy based on the multiple images that are continuous in space.

[0090] [Fourth Implementation Method]

[0091] In the first to third embodiments, a configuration in which one channel is allocated to an image acquired by the image acquisition unit 11 has been described, but the implementation is not limited to this. In the fourth embodiment, an example of allocating one channel to an image resulting from the synthesis of two images will be described. Figure 17 These are explanatory diagrams illustrating other methods of generating synthetic images in an image analysis apparatus corresponding to at least one embodiment of the present invention. For example... Figure 17As shown, two images can also be composited before channel allocation, for example, by using brightness and darkness comparison, and then channels can be allocated to the composite image. Composited by using brightness and darkness means, for example, when using 256 grayscale values ​​to represent the grayscale information of each color, one image converted to a color has a grayscale value converging in the range of 0–127, while the other image converted to a color has a grayscale value converging in the range of 128–255. By compositing these two images, two images can be allocated to one channel corresponding to a specific color. Using this method, the amount of information contained in the composite image can be increased. For example, when using 3 channels of RGB, by using this method, two images can be composited using brightness and darkness and then allocated to the 3 channels of RGB, allowing a composite image to contain information from six images.

[0092] [Fifth Implementation Method]

[0093] In the first to third embodiments, as an example of channel allocation for generating a composite image, the case of using the three colors of RGB as channels was described. In the fifth embodiment, as an example of various grayscale information that can be set as channels, a method for extracting luminance information was described.

[0094] (1) Methods for generating luminance images from color images

[0095] [Prerequisites]

[0096] When an image is captured by a monochrome camera and saved as grayscale, it is treated as a luminance image because it contains only a single channel. On the other hand, since grayscale images are often saved as color images, any channel can be used in this case.

[0097] [Utilization of a single channel]

[0098] An image using any one of the RGB color channels. In this case, the optimal channel can be selected based on the color of the object. Alternatively, the G channel, which best reflects brightness information, can be used in the color image.

[0099] Processing of multiple channels

[0100] A luminance image is generated by mixing the RGB channels in any proportion (this is different from simple extraction; it involves processing to obtain the luminance image). Alternatively, a luminance image can be obtained by simply adding the individual images and dividing by 3, or by mixing at a ratio that takes into account the wavelength sensitivity characteristics of the human eye. For example, as a method for converting a color image from the NTSC standard color system to grayscale, a method is known to calculate the luminance Y of each pixel using the formula Y = 0.299R + 0.587g + 0.114B.

[0101] [Processing of the extracted brightness image]

[0102] Applying processing to the extracted image to clarify the brightness relationship between the background and the object of interest is also effective. For example, when photographing a white object against a dark background (equivalent to the white mouse in the example) or a black object against a bright background (black mice are also frequently used in experiments), brightness inversion as needed is effective. Furthermore, when the background of the extracted image is gray and the object of interest is a slightly brighter gray than the background, brightness compensation by darkening the background and whitening the object of interest is very effective in improving prediction accuracy.

[0103] (2) A method of extracting or processing information other than brightness from a color image and using it as a brightness image.

[0104] [Preparatory Information]

[0105] Images obtained through RGB 3-channel color (RGB color space) can be converted to each other with color spaces consisting of hue, saturation, and value / luminance (HSV color space or the almost equivalent HLS color space).

[0106] [Use hue as a luminance image]

[0107] Because the hue of the HSV color space does not include the brightness of an image, the color of an object can be extracted as information even in the presence of shadows. This means it is highly robust against uneven lighting or shadow reflections. Furthermore, by extracting a defined range of hues, objects of interest can be extracted, making it potentially more useful than a simple brightness image. Specifically, by extracting hues near the red zone, it is easy to extract parts of human skin (palms or face). Extending this further, when multiple objects of interest with different hues exist in an image, these hue differences can be used as a brightness image (a typical hue is represented by a hue ring, a circular structure starting with red and progressing through yellow, green, light blue, blue, and purple, returning to red; by assigning a value of 60 to each of these colors, starting with red as 0, hues can be numerically represented). Therefore, object recognition and extraction can also be performed. Hues can also be modified or compensated as needed to improve prediction accuracy.

[0108] [Utilize hue as a luminance image]

[0109] Similar to hue, saturation or value can also be used. Since saturation represents the vividness of a color, it can be used to focus on vibrant areas regardless of the color type. However, unlike hue, saturation is easily affected by lighting or shadow reflections; therefore, its use is best limited to situations where such effects are minimal and the focus is on vividness. On the other hand, value is generally the brightness of the channel that has the highest value in each pixel [V = max(R, G, B)]. Therefore, although somewhat different, it can produce results similar to the naturally appearing grayscale images described in the multi-channel processing above. Brightness compensation can also be performed for these.

[0110] [Sixth Implementation Method]

[0111] In the first to third embodiments, as an example of channel allocation for generating a composite image, the case of using the three RGB colors as channels has been described, but it is not limited to this, and grayscale information of four or more colors may also be used as channels.

[0112] As a specific example of channel settings related to grayscale information, for example, when there are 4 colors or less, RGB(A) or CMYK methods can be used. Furthermore, in the fifth embodiment, a 6-color (red, yellow, green, light blue, blue, purple) example is given in the hue description, but the channel settings can also be set to 6 colors. Moreover, multiple colors can be used separately, and the colors assigned to each channel are preferably clear and easily distinguishable colors, as described above with the 6 colors.

[0113] Here, as a method of recording images, the most popular method currently is to use the RGB color space. In the CMYK mode or the 6-color (red, yellow, green, light blue, blue, purple) mode and other channel settings that include color representations other than RGB, there is a problem that the colors of overlapping parts cannot be correctly represented during the synthesis.

[0114] Figure 18 This is an explanatory diagram showing the relationship between each color and its RGB values ​​when the channels related to grayscale information are used in a 6-color configuration. Among them, Figure 18 (a) is an explanatory diagram showing the relationship between each color and RGB values ​​(unrestricted) when using a 6-color scheme of red, yellow, green, light blue, blue, and purple. Figure 18 (a) shows the upper limit of each RGB value when representing the six most common colors (red, yellow, green, light blue, blue, and violet) using 8-bit resolution (represented in brightness levels from 0 to 255). For example, if we are interested in yellow, the values ​​of R and G are 255, and B is 0. These are the same values ​​as the overlapping parts of red and green, so it is impossible to distinguish whether the yellow area in the image is an overlapping part of red and green or a separate part of yellow as a result of the composite. In addition, when multiple images are overlapped, if the RGB value exceeds 255, the color is saturated and cannot reflect the larger value, so it is rounded down to 255, thus resulting in overlapping layers that cannot be represented correctly.

[0115] Figure 18 (b) in the diagram illustrates the relationship between the six colors and their RGB values ​​when an upper limit is set to avoid damaging the image during compositing. Without this limit... Figure 18 As shown in (a), even when all color channels are combined, the RGB channel values ​​remain at a maximum of 3 (255×3) regardless of any element. Therefore, if the RGB values ​​used in the layer are each 85 (=255÷3), then even when all layers overlap, the brightness value of each pixel channel can be preserved to a maximum of 255. Thus, it is possible to preserve the color of all overlapping combinations of layers.

[0116] [Seventh Implementation Method]

[0117] In the first to third embodiments, it was described that different channels were set for multiple images, grayscale information corresponding to each channel was extracted from each image, and the extracted grayscale information was synthesized to generate a composite image. The following explanation was given: When generating the composite image, multiple grayscale information values ​​were synthesized in each pixel to determine the grayscale and brightness information of each pixel, but this is not a limitation. For example, pixel values ​​from new or old images in chronological order could be used as the grayscale and brightness information values ​​of each pixel. Similarly, pixel values ​​from spatially continuous images on the depth side or the front side could be used as the grayscale and brightness information values ​​of each pixel. That is, a composite image could also be obtained by rewriting and overlapping processes instead of performing synthesis.

[0118] Figure 19 This is an illustrative diagram showing an example of a composite image generated through a process of rewriting and overlapping. Figure 19 In this method, grayscale information of six colors is used as channels, and different color channels are assigned to six images that are sequential in time. Then, when compositing the six images, the compositing is performed by overlaying the newest images on top of each other, such as rewriting the second oldest image on top of the oldest image in time. At this time, it is necessary to determine the effective region in each image. For example, if you want to analyze the activity of a mouse, you can extract the outline of the mouse and determine the inner part of the outline as the effective region. The information of the pixels in the effective region is rewritten on the older image. The information outside the effective region can also use the information of the oldest image, or the same compositing process as in the first to third embodiments can be performed. In this way, when overlaying is performed while rewriting the older images in time sequentially, as Figure 19 As shown, it has the following effect: when an object moves, the trajectory of the movement is easier to grasp.

[0119] In addition, Figure 19 In the example, the process of overwriting all pixels within the effective area is performed and then superimposed. However, it is also possible to extract only the contour of the effective area of ​​each image with a specified pixel thickness and then... Figure 19 The same superposition is performed. This method also works with... Figure 19 Similarly, it makes it easier to grasp the trajectory of an object's movement.

[0120] [Eighth Implementation Method]

[0121] In the first to third embodiments, the inference process of the inference unit 14 was described under the premise that there exists a case where the inference processing is performed by a pre-learned complete learning model. While the entity responsible for generating the complete learning model in this case is not explicitly described, it is undoubtedly possible that the image analysis device 10 has a learning unit.

[0122] Figure 20 This is a block diagram illustrating an example of the configuration of an image analysis apparatus corresponding to at least one embodiment of the present invention. Figure 20 The image analysis apparatus 10D described in the first embodiment also includes a learning unit 18. The learning unit 18 performs learning processing by combining a synthesized image and correct answer data of the analysis results regarding the synthesized image, and updates the parameters of the neural network. That is, by providing the learning unit 18, the learning processing of the completed learning model used in the image analysis apparatus 10D can be performed within the apparatus itself, and additional learning processing can be performed on the obtained completed learning model. Furthermore, the learning processing is performed in conjunction with… Figure 5 The same process is used to execute it.

[0123] While various embodiments of the present invention have been described through the first to eighth embodiments, they are not limited thereto and can be applied to a wide variety of uses. Embodiments such as individual identification or abnormal behavior detection of people from surveillance cameras, abnormal vehicle driving detection from road surveillance cameras, behavior classification from live sports footage, and abnormal location detection from 3D data of internal organs are also conceivable.

[0124] Explanation of reference numerals in the attached figures

[0125] 10 Image Analysis Device

[0126] 11 Image Acquisition Unit

[0127] 12 Extraction Section

[0128] 13 Forecasting Department

[0129] 14. Inference Department

[0130] 15. Storage Department

[0131] 16. Location Information Acquisition Department

[0132] 17. Region Segmentation

Claims

1. An image analysis method, comprising: The image acquisition step involves acquiring multiple images that are continuous in time or space. The channel allocation step involves, for each of the plurality of images, allocating different channels to the grayscale information corresponding to colors that differ from each other in hue, according to a prescribed rule; The image synthesis generation step involves extracting and synthesizing grayscale information corresponding to the color of each of the plurality of images, which is assigned to the channel, thereby generating a color composite image that can identify at least a portion of the grayscale information of each image by synthesizing grayscale information corresponding to colors with different hues. The inference step involves analyzing the color composite image and making inferences about the multiple images.

2. The image analysis method according to claim 1, wherein, In the inference step, the input to the completed learning model is the color composite image generated in the composite image generation step, and the output of the completed learning model is obtained as the inference result. The completed learning model is a model that has been pre-machined based on multiple color composite images generated from multiple sample images used for learning.

3. The image analysis method according to claim 1 or 2, wherein, The multiple images are extracted from moving images at certain time intervals. The moving images are composed of multiple images that are continuous in time, obtained by capturing moving images of a defined object performing a movement. In the inference step, an inference related to the specified motion pattern of the object is performed based on the color composite image generated from the dynamic image.

4. The image analysis method according to claim 3, wherein, It also includes a location information acquisition step for acquiring the location information of the object when acquiring multiple images in the image acquisition step. When multiple objects are used as the analysis objects In the image acquisition step, the plurality of images are acquired for each of the plurality of objects. In the location information acquisition step, the location information is acquired for each of the plurality of objects. In the channel allocation step, for each of the multiple acquired images, a channel is allocated for each of the object objects. In the image synthesis generation step, a color composite image is generated for each of the objects. In the inference step, the multiple color composite images generated for each of the objects and the position information of each of the multiple objects are used as input to perform inference related to the action patterns of the multiple objects.

5. The image analysis method according to claim 1 or 2, wherein, The multiple images are either multiple tomographic images that are continuous in a specific direction when representing a three-dimensional region by stacking multiple acquired tomographic images in a specific direction, or multiple tomographic images extracted from a three-dimensional model in a continuous manner in a specific direction when representing a three-dimensional region by a three-dimensional model from which tomographic images can be arbitrarily extracted. In the inference step, inferences related to the three-dimensional region are performed based on the color composite image.

6. A method for generating images for learning or analysis, comprising: The image acquisition step involves acquiring multiple images that are continuous in time or space. The channel allocation step involves, for each of the plurality of images, allocating different channels to the grayscale information corresponding to colors that differ from each other in hue, according to a prescribed rule; The image synthesis generation step involves extracting and synthesizing grayscale information corresponding to the color and assigned to each of the plurality of images, thereby generating a color composite image that can identify at least a portion of the grayscale information of each image by synthesizing grayscale information corresponding to colors with different hues.

7. A method for generating a learning model, comprising: The image acquisition step involves acquiring multiple images that are continuous in time or space. The channel allocation step involves, for each of the plurality of images, allocating different channels to the grayscale information corresponding to colors that differ from each other in hue, according to a prescribed rule; The image synthesis generation step involves extracting and synthesizing grayscale information corresponding to the color of each of the plurality of images, which is assigned to the channel, thereby generating a color composite image that can identify at least a portion of the grayscale information of each image by synthesizing grayscale information corresponding to colors with different hues. The correct answer data acquisition step involves acquiring the correct answer data when performing inference on the color composite image; The inference step involves inputting the color-synthesized image into a model composed of a neural network and performing inference to output the inference result; The parameter update step involves updating the model's parameters using the inference results and correct answer data.

8. An image analysis device, comprising: An image acquisition unit acquires multiple images that are continuous in time or space; The channel allocation unit, for each of the plurality of images, allocates different channels to grayscale information corresponding to colors with different hues, according to a prescribed rule; The image synthesis unit extracts and synthesizes grayscale information corresponding to the color of each of the plurality of images by extracting grayscale information of each of the channels assigned to each of the plurality of images and synthesizing it, thereby generating a color composite image that can identify at least a portion of the grayscale information of each image by synthesizing grayscale information corresponding to colors with different hues. The inference unit analyzes the color composite image and makes inferences about the plurality of images.

9. A storage medium storing an image analysis program that enables a computer to perform the following functions: Image acquisition function, which acquires multiple images that are continuous in time or space; The channel allocation function, for each of the multiple images, allocates different channels to the grayscale information corresponding to colors that are different from each other in hue, according to a prescribed rule; The image synthesis generation function extracts and synthesizes grayscale information corresponding to the color of each of the plurality of images by extracting grayscale information of each channel and synthesizing it, thereby generating a color composite image that can identify at least a portion of the grayscale information of each image by synthesizing grayscale information corresponding to colors with different hues. The inference function analyzes the color composite image and makes inferences about the multiple images.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method and image processing program

    JP2019003565A

  • Method and device for producing vehicle operational data based on deep learning techniques

    US20180074493A1