State determination method, state determination program, state determination device, learned model generation method and learned model generation program
A state determination method using a large language model trained on document information addresses the challenge of varying object environments by determining object states directly from image data without generating abnormal images, enhancing accuracy and efficiency.
Patent Information
- Application Number
- JP2024000405
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-17
AI Technical Summary
Existing methods for determining the state of objects, such as plants, require time and effort to generate abnormal images, leading to decreased accuracy when diagnosing on-site image data due to varying cultivation environments.
A state determination method using a large language model trained on document information to generate image feature information and determine object states without requiring abnormal object generation, by inputting image and prompt data into a learned model.
Enables accurate state determination on captured image data without the need for generating abnormal objects, improving efficiency and reducing the time and effort typically required in conventional methods.
Smart Images

Figure 2025106831000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a state determination method, a state determination program, a state determination device, a learned model generation method, and a learned model generation program.
Background Art
[0002] Conventionally, for example, in Patent Document 1, as a machine learning method for detecting an object with high accuracy, the distance values of at least some of the pixels included in the detection target portion of the original distance image composed of a plurality of pixels indicating the distance values to the object are changed to abnormal values. A method of constructing a learned model has been proposed by performing supervised learning with the processed distance image as an input parameter and the output parameter as the label given to the detection target portion and the original distance image.
[0003] Further, in Patent Document 2, first image information, which is normal image information of a portion to be detected of a building, is generated from BIM information, which is three-dimensional information of the building, and second image information obtained by actually imaging the detection target location of the building is obtained. By performing image processing on at least one of the first image information and the second image information, third image information with newly added abnormal locations is generated. The third image information is learned by a generation model using a GAN, and fourth image information as image information for augmentation is generated from the first image information by the learned generation model. A technique for learning an anomaly detection model using the second image information and the fourth image information as teacher data has been proposed.
[0004] According to the above-described conventional techniques, for example, by imaging an object such as a plant with an imaging device in a test field or the like, after collecting a normal image in which the object is in a normal state and an abnormal image in which the object is in an abnormal state, machine learning is performed to generate a learning model. Then, a program for performing an abnormality diagnosis as to whether an arbitrary object image is normal or abnormal can be generated by the generated learning model.
Prior Art Documents
Patent Documents
[0005] [Patent Document 1] International Publication No. 2020 / 105225 [Patent Document 2] Japanese Unexamined Patent Application Publication No. 2021 - 140379 [Summary of the Invention] [Problems to be Solved by the Invention]
[0006] However, for example, in the case of plants, if the cultivation environment is different, not only the appearance of the images obtained by imaging but also the form and color of the plants will differ. Therefore, when diagnosing the image data (on - site image data) taken at the location (on - site) where an object such as a plant exists, using the learning model obtained by performing machine learning with normal images and abnormal images in a test field, the accuracy of abnormal diagnosis may decrease. At this time, there are problems such as the need for time and effort to collect abnormal images on - site, create abnormal states in plant growth, or create abnormal images using a generation model such as a GAN. These problems are the same even if the object is something other than a plant, and there has been a demand for the development of a technology that can determine the state of the imaged object without the time and effort required for generating abnormal objects.
[0007] The present invention has been made in view of such circumstances, and its object is to provide a state determination method, a state determination program, a state determination device, a learned model generation method, and a learned model generation program that can perform state determination on the captured image data obtained by imaging an object without the time and effort required for generating abnormal objects. [Means for Solving the Problems]
[0008] In order to solve the above-described problems and achieve the object, a state determination method according to one aspect of the present invention includes an image feature information generation step of generating image feature information including a plurality of language information representing different modes among the modes of an object depicted in an image to be determined, an input reception step of receiving an input of a prompt instructing the state determination of the object, and a determination result acquisition step of inputting the image feature information read from a storage unit and the prompt into a learned model to obtain a determination result of determining the state of the object to be determined. The learned model is generated by training a large language model with a plurality of feature data sets which are document information composed of sentences indicating the modes of the object and the determination of the state in the modes.
[0009] The state determination method according to one aspect of the present invention further includes, in the above invention, a numerical information generation step of generating numerical information regarding an object depicted in an image to be determined. The determination result acquisition step inputs the image feature information, the numerical information, and the prompt into the learned model to obtain a determination result of determining the state of the object to be determined.
[0010] In the state determination method according to one aspect of the present invention, in the above invention, the image feature information generation step generates the image feature information based on language information obtained by inputting an image in which an object to be determined is depicted and a prompt designating different modes into a multimodal large language model.
[0011] In the state determination method according to one aspect of the present invention, in the above invention, the determination result is language information indicating an answer to a prompt instructing the state determination of the object.
[0012] The state determination program according to one aspect of the present invention causes a computer to execute an image feature information generation step of generating image feature information including a plurality of language information representing different modes of an object shown in an image to be determined, an input reception step of receiving an input of a prompt instructing the state determination of the object, and a determination result acquisition step of inputting the image feature information read from a storage unit and the prompt into a learned model to acquire a determination result of determining the state of the object to be determined. The learned model is generated by causing a large language model to learn a plurality of feature data sets which are document information constituted by sentences indicating the mode of the object and the determination of the state in the mode.
[0013] The state determination device according to one aspect of the present invention includes an image feature information generation unit that generates image feature information including a plurality of language information representing different modes of an object shown in an image to be determined, an input reception unit that receives an input of a prompt instructing the state determination of the object, and a state determination unit that inputs the image feature information read from a storage unit and the prompt into a learned model to acquire a determination result of determining the state of the object to be determined. The learned model is generated by causing a large language model to learn a plurality of feature data sets which are document information constituted by sentences indicating the mode of the object and the determination of the state in the mode.
[0014] A method for generating a learned model according to one aspect of the present invention is a method for generating a learned model that outputs a determination result of determining the state of an object shown in an image, and includes a generation step of causing a large language model to learn a plurality of feature data sets which are document information constituted by sentences indicating the mode of the object and the determination of the state in the mode, read from a storage unit, to generate a learned model.
[0015] A learned model generation program according to an aspect of the present invention is a learned model generation program that generates a learned model that outputs a determination result for determining the state of an object depicted in an image. The program causes a computer to execute a generation step of causing a large language model to learn a plurality of feature data sets that are document information composed of the aspect of the object and a sentence indicating the determination of the state in the aspect, thereby generating a learned model.
Advantages of the Invention
[0016] According to the state determination method, state determination program, state determination device, learned model generation method, and learned model generation program of the present invention, it is possible to perform state determination on imaging image data obtained by imaging an object without requiring time and effort due to the generation of abnormal objects.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In all the drawings of the following embodiment, the same or corresponding parts are denoted by the same reference numerals. Also, the present invention is not limited to the embodiment described below.
[0019] First, in explaining the state determination according to an embodiment of the present invention, the intensive study conducted by the present inventor will be described. That is, according to the findings of the present inventor, as a state determination process which is a program for determining the state of a predetermined object, it is conceivable to adopt a learned model generated by machine learning. The learned model can be generated by normal / abnormal feature data created by specialized books or the like. That is, a method of generating a state determination model for determining normal / abnormal by performing machine learning using at least the normal / abnormal feature data created by specialized books or the like as teacher data is conceivable.
[0020] In the prior art, it is possible to acquire abnormal images in a test field, but in the field, while it is possible to generate so-called normal state objects such as normal crops and organisms, it has been extremely difficult to generate so-called abnormal state objects such as abnormal crops and organisms. At this time, as described above, when diagnosing an object generated in the field using a model trained with data (normal or abnormal data) including abnormal images acquired in the test field, the accuracy of abnormal diagnosis decreases.
[0021] Based on the above points, the present inventor also examined a method of additionally performing machine learning using normal data and abnormal data obtained by imaging an object in the field, but creating an abnormal object in the field has problems that require time and effort. Furthermore, it is itself difficult to generate an abnormal object, and a problem of causing significant losses by generating an abnormal object also occurs.
[0022] Therefore, the present inventor has further intensively studied the generation of a model that can accurately determine the state even when there is no local abnormal data. That is, the present inventor came up with a method of training a large language model (LLM) to create a state determination model and using the LLM for determination. Specifically, the present inventor obtained information related to state determination from documents such as literature and textbooks, and used these as teacher data to train the LLM to obtain a trained model (state determination model). In the trained model, the language information obtained by inputting an image of the object into the multimodal LLM is input, and the features corresponding to the set prompt are output as language information. In this way, since the trained model is generated by machine learning of specialized feature data, the determination accuracy can be improved.
[0023] Note that the object of state determination is applicable not only to plant cultivation and biological growth, but also to objects where the environment where abnormal data is easily obtained is different from the environment where it is actually desired to be applied. Specifically, in the case of agricultural crops, it is applicable to programs for discriminating diseases, physiological disorders, and pest damage from local images. In addition, it is also applicable to flames in incinerators, buildings such as plant facilities in water treatment plants, bridges, and fish farming. Examples of the state determination of buildings and bridges include metal corrosion. The following embodiment is devised based on the above intensive studies by the present inventor.
[0024] (State Determination System) FIG. 1 shows an information processing apparatus according to an embodiment of the present invention. As shown in FIG. 1, the state determination system 1 includes a learning device 2 and a state determination device 3. The state determination system 1 is configured such that the learning device 2 and the state determination device 3 can input and output data to each other.
[0025] Here, the input and output of data of the learning device 2 and the state determination device 3 can be executed by network communication via a network, cloud, etc., or can be executed by contactless communication such as Bluetooth (registered trademark). In addition, it is also possible to execute the transfer of data via a disk recording medium such as a USB (Universal Serial Bus) memory, a CD (Compact Disc), a DVD (Digital Versatile Disc), or a BD (Blu-ray (registered trademark) Disc). The network is configured by appropriately combining wired communication and wireless communication, and is composed of a communication network such as an Internet line network or a mobile phone line network. The network consists of, for example, a dedicated line, a public communication network such as the Internet, a telephone communication network such as a LAN (Local Area Network), a WAN (Wide Area Network), a mobile phone, a public line, one or a combination of a plurality of VPNs (Virtual Private Network), etc.
[0026] FIG. 2 is a block diagram showing the configuration of a learning device included in a state determination system according to an embodiment of the present invention. The learning device 2 includes a communication unit 21, a learning unit 22, an input / output unit 23, a control unit 24, and a storage unit 25.
[0027] The communication unit 21 is, for example, a LAN interface board, a wired communication circuit for wired communication, or a wireless communication circuit for wireless communication. The LAN interface board, the wired communication circuit, and the wireless communication circuit are connected to the network. The communication unit 21 as a transmission unit and a reception unit is connected to the network and communicates with the state determination device 3.
[0028] The learning unit 22 performs learning using learning data to generate a learned model. The learned model in the present embodiment is a large language model (LLM) generated by performing machine learning on document information (feature data set described later) using a multi-layer neural network including an input layer, an intermediate layer, and an output layer. The machine learning performed by the learning unit 22 can use known methods.
[0029] The input / output unit 23 can be composed of, for example, a touch panel display, a speaker microphone, etc. The input / output unit 23 as input means may include an interface that inputs various information transmitted from an external server through the communication unit and outputs it to the control unit 24. Further, the input / output unit 23 includes a user interface such as a keyboard, input buttons, a lever, a touch panel for manual input provided superimposed on a display such as a liquid crystal, or a microphone for voice recognition. By an operator or the like operating the input / output unit 23, it is configured to be able to input predetermined information to the control unit 24. The input / output unit 23 as output means displays a predetermined image or the like on a display monitor, displays characters, figures, etc. on the screen of the touch panel display, or outputs voice from a speaker according to the control by the control unit 24. That is, the input / output unit 23 is configured to be able to notify external parties of predetermined information. Note that the input unit and the output unit in the input / output unit 23 may be configured separately.
[0030] Specifically, the control unit 24 includes a processor having hardware such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field-Programmable Gate Array), and a main storage unit such as a RAM (Random Access Memory) and a ROM (Read Only Memory) (none of which are shown).
[0031] The storage unit 25 is composed of a storage medium selected from a volatile memory such as a RAM, a non-volatile memory such as a ROM, an EPROM (Erasable Programmable ROM), a hard disk drive (HDD), and a removable medium. Note that as the removable medium, for example, a USB memory or a disk recording medium such as a CD, a DVD, or a BD can be adopted. Also, the storage unit 12 may be configured using a computer-readable recording medium such as a memory card that can be externally attached.
[0032] The storage unit 25 can store an operating system (OS), various programs such as a learning application, various tables, various databases, etc., for executing the operations of the learning device 2. These various programs can also be recorded on a computer-readable recording medium such as a hard disk, a flash memory, a CD-ROM, a DVD-ROM, or a flexible disk and widely distributed.
[0033] FIG. 3 is a block diagram showing the configuration of a state determination device included in the state determination system according to an embodiment of the present invention. The state determination device 3 includes a communication unit 31, an image feature information generation unit 32, a state determination unit 33, an input / output unit 34, a control unit 35, and a storage unit 36.
[0034] The communication unit 31 is, for example, a LAN interface board, a wired communication circuit for wired communication, or a wireless communication circuit for wireless communication. The LAN interface board, the wired communication circuit, and the wireless communication circuit are connected to a network. The communication unit 31 as a transmission unit and a reception unit connects to the network and communicates with the learning device 2.
[0035] The image feature information generation unit 32 generates image feature information regarding the object depicted in the image. For example, the image feature information generation unit 32 inputs the image and a prompt into a multimodal large language model (MLLM) to obtain language information regarding the aspect of the object as the image feature information. Note that as the multimodal MLLM, known ones such as GPT-4V, LLaVA-1.5, Ferret, and MM-VID can be used. Also, as long as language information based on the image can be obtained, not limited to the multimodal MLLM, other known methods may be used.
[0036] The state determination unit 33 determines the state of the object depicted in the image using the learned model (LLM) generated by the learning device 2. The state determination unit 33 inputs the image feature information generated by the image feature information generation unit 32 and a prompt (linguistic prompt), which is language information indicating an instruction or question from the user, into the learned model, and outputs a determination result based on the result output by the learned model.
[0037] The input / output unit 34 can be composed of, for example, a display, a touch panel display, a speaker microphone, etc. The input / output unit 34 as input means may include an interface that inputs various information transmitted from an external server through the communication unit and outputs it to the control unit 35. Also, the input / output unit 23 includes a user interface such as a keyboard, input buttons, a lever, a touch panel for manual input provided superimposed on a display such as a liquid crystal, or a microphone for voice recognition. By an operator or the like operating the input / output unit 34, it is configured to be able to input predetermined information to the control unit 35. The input / output unit 34 as output means displays a predetermined image or the like on a display monitor, displays characters, figures, etc. on the screen of the touch panel display, or outputs voice from a speaker according to the control by the control unit 35. That is, the input / output unit 34 is configured to be able to notify external parties of predetermined information. Note that the input unit and the output unit in the input / output unit 34 may be configured separately.
[0038] Specifically, the control unit 35 includes a processor having hardware such as a CPU, a DSP, and an FPGA, and a main storage unit (both not shown) such as a RAM and a ROM.
[0039] The storage unit 36 is composed of a storage medium selected from volatile memories such as a RAM, non-volatile memories such as a ROM, an EPROM, an HDD, and a removable medium. Note that as the removable medium, for example, a USB memory or a disk recording medium such as a CD, a DVD, or a BD can be adopted. Also, the storage unit 36 may be configured using a computer-readable recording medium such as a memory card that can be externally attached.
[0040] The storage unit 36 can store various programs such as an OS and a state determination application, various tables, and various databases for executing the operation of the state determination device 3. These various programs can also be recorded on a computer-readable recording medium such as a hard disk, a flash memory, a CD-ROM, a DVD-ROM, or a flexible disk and widely distributed.
[0041] (Learning Process) FIG. 4 is a flowchart for explaining the learning process according to an embodiment of the present invention. Hereinafter, an example of generating a learned model that outputs the result of determining the state of the "leaf" shown in the image will be described.
[0042] First, the learning unit 22 acquires a plurality of feature datasets (step S11). This feature dataset is created using document information composed of, for example, academic information that is learned or scientifically proven from specialized books, etc., or a state of an object and a sentence indicating a determination result of that state based on hearsay information. Specifically, it is data that, for example, takes the state of a leaf (aspect) and the determination result of normal or abnormal (disease name) in this aspect as a pair. Specifically, the feature dataset is a dataset that takes, for "leaves", the aspect information of the shape, color, and pattern of the leaf and the normal / abnormal (disease name) of the leaf in that aspect as a pair. For example, for multiple types of aspect information such as "round shape" and "many brown spots on the leaf", a determination result of "is a disease called ○○" is associated. At this time, as the feature dataset, language information obtained by inputting an image and a prompt (for example, "Extract the features of the shape, color, and pattern from the leaf image respectively") into the multimodal LLM may be adopted.
[0043] After that, the learning unit 22 generates a learned model by learning using a plurality of feature datasets and the LLM (step S12). Specifically, the learning unit 22 performs machine learning on the LLM with a plurality of datasets having the aspect information of the leaf as an explanatory variable and the determination result indicating the normal / abnormal (disease name) of the leaf as an objective variable to generate a learned model.
[0044] (State determination process) Next, an image diagnosis method by a state determination program executed by the state determination device 3 configured as described above will be described. FIG. 5 is a flowchart for explaining the state determination process according to an embodiment of the present invention. When an execution instruction for the state determination process is input, the control unit 35 reads the program from the storage unit 36 and executes the state determination process.
[0045] First, an image to be determined is acquired (step S101). In step S101, the control unit 35 acquires the image to be determined via the communication unit 31 or the input / output unit 34. Hereinafter, an example of determining the state of "leaves" shown in the image will be described.
[0046] Subsequently, the image feature information generation unit 32 generates image feature information from the image to be determined (step S102). The image feature information generation unit 32 inputs, for example, an image and a plurality of types of prompts related to the states of leaves, such as "leaf shape", "leaf color", and "leaf pattern", into the multimodal LLM, and obtains language information related to the shape, color, and pattern of the leaves as image feature information. Specifically, for example, when the image feature information generation unit 32 inputs the image to be determined and a prompt such as "This plant is ◇◇. Regarding this image, please tell me the following. Leaf shape. Leaf color. Leaf pattern." into the multimodal LLM, language information such as "It has a rounded shape, is generally green, and has many brown spots" is obtained as image feature information. Note that when the multimodal LLM can determine the type of plant from the leaf image, it is not necessary to input the type of plant into the prompt.
[0047] After obtaining the image feature information, the input / output unit 34 accepts the input of a prompt related to the state determination (step S103). The input / output unit 34 accepts the input of a prompt such as "Please tell me the state of this plant." via the input unit, for example. Note that step S103 may be executed before step S102 or may be executed simultaneously.
[0048] After acquiring the image feature information and accepting the input of the prompt, the state determination unit 33 determines the state of the object depicted in the image using the learned model (LLM) generated by the learning device 2 (step S104). The state determination unit 33 inputs the image feature information generated by the image feature information generation unit 32 into the learned model, and outputs a determination result based on the result output by the learned model. For example, when language information regarding the leaves is input into the learned model, a state determination result for those leaves is output. For example, regarding the leaves, language information such as "This plant is ◇◇, has a rounded shape, is entirely green, and has many brown spots", and the prompt "Please tell me the state of this plant." is input into the learned model, and determination results such as "The state of the leaves is abnormal, and the disease name is estimated to be ○○", "Growth is good (normal)", "Insufficient (or excessive, etc.) moisture (or fertilizer, sunlight, etc.) is estimated" are output from the learned model. Note that in this step S104, the state determination unit 33 may acquire the learned model from the learning device 2, or if the learned model is stored in the storage unit 36 of the state determination device 3, the learned model may be read out from the storage unit 36.
[0049] Then, the input / output unit 34 outputs the determination result (step S105). The input / output unit 34 causes the determination result to be displayed on a display, for example.
[0050] According to the above-described embodiment, regarding an object, based on a plurality of aspect information of different types generated by a multimodal LLM and a trained model obtained by training the LLM with a feature dataset from a specialized document or the like, the state of the object is determined. Here, in conventional image diagnosis, it was necessary to consider image features and select a machine learning model. However, in this embodiment, in order to cause the LLM to interpret, it is specialized in extracting features such as those that an expert would mention when looking at the image (here, a leaf) from a specialized book or the like. As a result, it becomes possible to perform state determination without collecting a large amount of on-site images. According to this embodiment, in order to interpret the features of the image interpreted by the multimodal LLM and the quantitative information, fine-tuning specialized for determination is performed on the existing LLM. Therefore, it is possible to perform state determination on the captured image data obtained by imaging the object without taking time and effort due to the generation of abnormal objects.
[0051] (Modification example) Next, a modification example of the embodiment of the present invention will be described with reference to FIGS. 6 to 8. Since the configuration of the learning device according to this modification example is the same as the configuration of the learning device 2 according to the embodiment, the description thereof will be omitted. Hereinafter, the same components as those of the state determination system 1 according to the embodiment will be described with the same reference numerals. In this modification example, as information generated based on an image, numerical information is further included to perform a trained model and state determination processing.
[0052] FIG. 6 is a block diagram showing the configuration of a state determination device included in a state determination system according to a modification example of the embodiment of the present invention. The state determination device 3A includes a communication unit 31, an image feature information generation unit 32, a state determination unit 33, an input / output unit 34, a control unit 35, a storage unit 36, and a numerical information generation unit 37.
[0053] The numerical information generation unit 37 generates numerical information regarding the object shown in the image. The numerical information generation unit 37 generates numerical information, for example, by image processing, a learned model, or the like. Specifically, for the "leaf", the numerical information generation unit 37 generates numerical information by counting the number of spots for each color, digitizing the size and distribution of the spots, etc. The numerical information generation unit 37 may further add numerical information such as color tone based on image processing and environmental information such as temperature and humidity.
[0054] (Learning process) In the learning process according to this modification example, a learned model is generated according to the learning process shown in FIG. 4. First, the learning unit 22 acquires a plurality of feature data sets (step S11). The feature data set in this modification example is a data set that combines, for the "leaf", the shape of the leaf, the color of the leaf and the pattern information of the leaf, and numerical information such as the number of spots, and the normal / abnormal (disease name) of the leaf in that state. For example, for a plant named ◇◇, for the information "rounded shape" and "□□ brown spots on the leaf", the determination result of "if there are □□ spots, it is the disease △△" is associated.
[0055] After that, the learning unit 22 generates a learned model by learning using a plurality of feature data sets and an LLM (step S12). Specifically, the learning unit 22 causes the LLM to perform machine learning on a plurality of data sets having the mode information and numerical information of the leaf as explanatory variables and the determination result indicating the normal / abnormal (disease name) of the leaf as the target variable to generate a learned model.
[0056] (State determination process) Next, an image diagnosis method by a state determination program executed by the state determination device 3 configured as described above will be described. FIG. 7 is a flowchart for explaining the state determination process according to the embodiment of the present invention. When an execution instruction for the state determination process is input, the control unit 35 reads the program from the storage unit 36 and executes the state determination process.
[0057] First, obtain the image to be determined (step S201). In step S201, the control unit 35 obtains the image to be determined via the communication unit 31 or the input / output unit 34. Hereinafter, an example of determining the state of the "leaf" shown in the image will be described.
[0058] After that, the image feature information generation unit 32 generates image feature information from the image to be determined (step S202). The image feature information generation unit 32 inputs, for example, the image and prompts such as "leaf shape", "leaf color", and "leaf pattern" into the multimodal LLM in the same manner as in step S102 described above, and obtains language information regarding the shape, color, and pattern of the leaf as image feature information.
[0059] Also, the numerical information generation unit 37 generates numerical information from the image to be determined (step S203). The numerical information generation unit 37 obtains, for example, the "number of brown spots" from the "leaf" shown in the image as numerical information. In addition, the numerical information generation unit 37 can calculate the number of spots for each color, the occupied area of the spots on the surface of the leaf, etc., according to the information input to the learned model.
[0060] After obtaining the image feature information and the numerical information, the input / output unit 34 accepts the input of the prompt related to the state determination (step S204). The input / output unit 34 accepts the input of a prompt such as "Please tell me the state of this plant." via the input unit, for example. Note that in steps S202 to S204, step S203 or step S204 may be executed first, or they may be executed simultaneously.
[0061] After acquiring the image feature information and numerical information and receiving the input of the prompt, the state determination unit 33 determines the state of the object depicted in the image using the learned model (LLM) generated by the learning device 2 (step S205). The state determination unit 33 inputs the image feature information generated by the image feature information generation unit 32, the numerical information generated by the numerical information generation unit 37, and the prompt into the learned model, and outputs a determination result based on the result output by the learned model. For example, when language information about a leaf, numerical information, and a prompt indicating to output a determination result are input into the learned model, a state determination result for that leaf is output. For example, for a leaf, information such as "This plant is ◇◇, has a rounded shape, is entirely green, and has □□ brown spots", and a prompt "Please tell me the state of this plant." are input into the learned model, and a determination result such as "If there are □□ spots, it is presumed to be the disease △△" is output from the learned model. In this step S205, the state determination unit 33 may acquire the learned model from the learning device 2, or if the learned model is stored in the storage unit 36 of the state determination device 3, the learned model may be read from the storage unit 36.
[0062] Then, the input / output unit 34 outputs the determination result (step S206). The input / output unit 34 causes, for example, the determination result to be displayed on a display.
[0063] According to the modified example described above, similar to the embodiment, regarding the object, based on a plurality of image feature information of different types generated by the multimodal LLM and the learned model learned by the LLM using the feature dataset from specialized documents, etc., the state of the object is determined. Therefore, it is possible to perform state determination on the captured image data obtained by imaging the object without spending time and effort due to the generation of abnormal objects.
[0064] Also, according to the modification example, since the feature data set and the numerical information regarding "leaves" are further included as the information used for the determination process, the state can be determined in a more detailed and accurate manner.
[0065] In the above-described embodiments and modification examples, an example of determining the state regarding the "leaves" shown in the image has been described. However, in addition to the leaves, it is possible to determine the state from other parts such as flowers and stems, or objects other than plants.
[0066] (Recording medium) In the above-described one embodiment, a program for executing the processing method executed by the learning device 2 and the state determination device 3 can be recorded on a recording medium readable by a device such as a computer, other machines, or a wearable device (hereinafter referred to as a computer, etc.). By causing a computer or the like to read and execute the program of this recording medium, the computer or the like functions as a movement control device. Here, a recording medium readable by a computer or the like refers to a non-temporary recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer or the like. Examples of removable recording media of such among these include flexible disks, magneto-optical disks, CD-ROMs, CD-R / Ws, DVDs, BDs, DATs, magnetic tapes, memory cards such as flash memories, etc. Also, as recording media fixed to a computer or the like, there are hard disks, ROMs, etc. Furthermore, SSDs can be used as both removable recording media from a computer or the like and recording media fixed to a computer or the like.
[0067] As described above, one embodiment of the present invention has been specifically explained. However, the present invention is not limited to the above-described embodiment, and various modifications based on the technical idea of the present invention are possible. For example, the models cited in the above-described embodiment are merely examples, and different models may be used as necessary. The present invention is not limited by the descriptions and drawings that form part of the disclosure of the present invention according to this embodiment. For example, in the above-described embodiment, an example of performing state determination using a trained model in which a large language model is trained with a feature dataset has been described. However, a configuration using a pre-created (untrained) large language model or a model based on a large language model such as a model (framework) using RAG (Retrieval-augmented Generation) may also be used. For example, in RAG, a prompt with reference information is input to an untrained large language model to obtain a determination result. The reference information at this time is, for example, the latest information related to the state determination of the object.
[0068] Also, in the above-described embodiment, deep learning using a neural network as an example of machine learning has been used. However, machine learning based on other methods may also be performed. For example, other supervised learning methods such as support vector machines, decision trees, naive Bayes, and k-nearest neighbor methods may be used. Also, semi-supervised learning may be used instead of supervised learning.
[0069] Also, in one embodiment, the above-mentioned "part" can be read as "circuit" or the like. For example, the control unit can be read as a control circuit.
[0070] Further effects and modification examples can be easily derived by those skilled in the art. The broader aspects of the present disclosure are not limited to the specific details and representative embodiments represented and described as above. Therefore, various changes are possible without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.
Description of Reference Numerals
[0071] 1 State determination system 2 Learning device 3 State determination device 21, 31 Communication unit 22 Learning unit 23, 34 Input / output unit 24, 35 Control unit 25, 36 Memory unit 32 Image feature information generation unit 33 State determination unit 37 Numerical information generation unit
Claims
1. An image feature information generation step of generating image feature information including a plurality of language information representing a plurality of different modes, which are modes of an object depicted in an image to be determined; An input reception step of receiving an input of a prompt instructing a state determination of the object; A determination result acquisition step of inputting the image feature information read from a storage unit and the prompt into a learned model, and acquiring a determination result of determining the state of the object to be determined; Including, The learned model is generated by training a large language model with a plurality of feature datasets which are document information composed of sentences indicating the mode of the object and the determination of the state in the mode; A state determination method.
2. A numerical information generation step of generating numerical information regarding an object depicted in an image to be determined; Further including, In the determination result acquisition step, the image feature information, the numerical information, and the prompt are input into the learned model to acquire a determination result of determining the state of the object to be determined. The state determination method according to Claim 1.
3. The image feature information generation step generates the image feature information based on language information obtained by inputting an image in which an object to be determined is depicted and a prompt specifying different modes into a multimodal large language model. The state determination method according to Claim 1.
4. The determination result is language information indicating an answer to a prompt instructing a state determination of the object. The state determination method according to Claim 1.
5. An image feature information generation step of generating image feature information including a plurality of language information representing a plurality of different modes, which are modes of an object depicted in an image to be determined; An input reception step of receiving an input of a prompt instructing a state determination of the object; A determination result acquisition step of inputting the image feature information read from a storage unit and the prompt into a learned model, and acquiring a determination result of determining the state of the object to be determined; To be executed by a computer, The learned model is generated by training a large language model with a plurality of feature datasets which are document information composed of sentences indicating the mode of the object and the determination of the state in the mode; A state determination program.
6. An image feature information generation unit that generates image feature information including a plurality of language information representing different modes of an object shown in an image to be determined; An input reception unit that receives an input of a prompt instructing the determination of the state of the object; A state determination unit that inputs the image feature information read from the storage unit and the prompt into a learned model, and obtains a determination result for determining the state of the object to be determined; Comprising; The learned model is generated by causing a large language model to learn a plurality of feature data sets which are document information composed of sentences indicating the mode of the object and the determination of the state in the mode; A state determination device.
7. A learned model generation method for generating a learned model that outputs a determination result for determining the state of an object shown in an image, A generation step of generating a learned model by causing a large language model to learn a plurality of feature data sets which are document information composed of sentences indicating the mode of the object and the determination of the state in the mode, read from a storage unit; A learned model generation method including this.
8. A learned model generation program for generating a learned model that outputs a determination result for determining the state of an object shown in an image, A generation step of generating a learned model by causing a large language model to learn a plurality of feature data sets which are document information composed of sentences indicating the mode of the object and the determination of the state in the mode; A learned model generation program that causes a computer to execute this.
Citation Information
Patent Citations
Learning method for abnormality detection model of building, learning device, generation method for landscape information of building, generation device, and computer program
JP2021140379A
Machine learning method, trained model, control program, and object detection system
WO2020105225A1