Information processing method, program, and information processing apparatus

By using a language model to generate classification information and training an object classification model, the method addresses the limitation of existing technologies in generating classification information, achieving improved accuracy and reliability in object classification.

JP2026005025APending Publication Date: 2026-01-15TOKYO ELECTRON DEVICE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024103203
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing information processing methods, such as Patent Document 1, are unable to generate classification information related to the classification of objects using a language model.

Method used

An information processing method that utilizes a computer to acquire images of objects and apply them to a language model to generate classification information, followed by training an object classification model using the generated classification information.

Benefits of technology

Enables the generation of detailed and accurate classification information for objects, improving the reliability of classification results by combining language model-generated information with a trained object classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026005025000001_ABST
    Figure 2026005025000001_ABST
Patent Text Reader

Abstract

To provide an information processing method or the like capable of generating classification information related to the classification of an object by a language model.SOLUTION: In an information processing method according to one aspect, a computer 1 executes a process of acquiring an image of a target object captured in a plant and providing the acquired image to a language model to generate classification information regarding classification of the target object. This makes it possible to generate classification information regarding the classification of the object by the language model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing method, a program, and an information processing device. [Background technology]

[0002] In recent years, there has been active development of technology for classifying objects based on captured images. For example, Patent Document 1 discloses an information processing method that acquires an image of an object captured using visible light and an image captured using infrared light, classifies the object from the visible light image, and detects the state of the object from the infrared image based on the classification result. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7223194 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the invention of Patent Document 1 has a problem in that it is not possible to generate classification information relating to the classification of objects using a language model.

[0005] One aspect of the present invention is to provide an information processing method and the like that can generate classification information related to classification of objects using a language model. [Means for solving the problem]

[0006] An information processing method according to one aspect is characterized in that a computer executes a process of acquiring an image of an object photographed in a plant, and applying the acquired image to a language model to generate classification information regarding a classification of the object. [Effects of the Invention]

[0007] In one aspect, the language model allows for the generation of classification information regarding object classes. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is an explanatory diagram showing an outline of an automatic operation system. [Figure 2] FIG. 1 is a block diagram illustrating an example of the configuration of a computer. [Figure 3] FIG. 10 is an explanatory diagram showing an example of a record layout of a training data DB. [Figure 4] FIG. 1 is a block diagram showing an example of the configuration of an apparatus. [Figure 5] FIG. 2 is an explanatory diagram illustrating an example of the data structure of a class definition file. [Figure 6] FIG. 10 is an explanatory diagram illustrating a process for generating training data. [Figure 7] 10 is a flowchart showing a processing procedure for generating training data. [Figure 8] 10 is a flowchart showing a processing procedure when both the first classification information and the second classification information are output. [Figure 9] FIG. 10 is an explanatory diagram illustrating a process for generating training data in the second embodiment. [Figure 10] 10 is a flowchart showing a processing procedure for generating training data in the second embodiment. [Figure 11] FIG. 11 is a block diagram showing an example of the configuration of a computer according to a third embodiment. [Figure 12] 11 is a flowchart showing a processing procedure for generating training data in the third embodiment. [Figure 13] 10 is a flowchart showing a processing procedure for generating classification information of an object based on an image and time-series data of the object. [Figure 14] 10 is a flowchart showing a processing procedure when outputting the state of an object. DETAILED DESCRIPTION OF THE INVENTION

[0009] The present invention will be described in detail below with reference to the drawings showing embodiments thereof.

[0010] (Embodiment 1) The first embodiment relates to a form in which classification information relating to classification of objects is generated using a language model. Fig. 1 is an explanatory diagram showing an overview of an automatic operation system. The system of this embodiment includes an information processing device 1, an imaging device 2, and a device 3, and each device transmits and receives information via a network N such as the Internet.

[0011] The information processing device 1 is an information processing device that processes, stores, and transmits / receives various types of information. The information processing device 1 is, for example, a personal computer or a server device. In the following description of the present embodiment, the information processing device 1 will be referred to as a computer 1 for the sake of simplicity.

[0012] The imaging device 2 is installed to capture images of objects in a plant, and is, for example, an imaging device such as a CCD (Charge Coupled Device) camera or a CMOS (Complementary Metal Oxide Semiconductor) camera.

[0013] The imaging device 2 includes a wireless communication unit. The wireless communication unit is a wireless communication module for performing communication-related processing, and transmits the captured image of the object to the computer 1 via the network N. The imaging device 2 may be connected to the computer 1 via a wired connection. The imaging device 2 may be replaced by a personal computer capable of capturing images, a smartphone, or a mobile surveillance robot capable of capturing images of the object.

[0014] The objects include food products or industrial products, or substances or materials used to manufacture food products or industrial products. Food products include, for example, confectionery (chocolate, ice cream, biscuits, rice crackers, candy, etc.), processed meat products (processed meat, minced meat, hamburger steak, sausage, etc.), kamaboko products (fish paste products made by shaping and heating minced fish paste), tamagoyaki (rolled eggs), rice balls, etc. Industrial products are molded products made of metal, resin, etc. (cast products, forged products, injection molded products, extrusion molded products, etc.).

[0015] The device 3 is a device within a plant. The plant may be, for example, an energy plant, an industrial plant (e.g., a food plant, a petrochemical plant, or a pharmaceutical plant), or an environmental plant. An energy plant is a device that produces energy from thermal or nuclear power. An industrial plant is a device that produces chemical products, food, medicines, metals, or the like used in daily life. An environmental plant is a device that treats sewage or waste and further recycles resources.

[0016] The equipment in the plant is a plurality of pieces of equipment such as manufacturing equipment, processing equipment, electrical equipment, or piping that are combined to produce a specific product or material. As an example, the equipment 3 in this embodiment is equipment that manufactures food products or industrial products. The equipment 3 is manufacturing equipment that includes a kettle, a shaft, or various sensors (e.g., a weight measurement sensor, a temperature sensor, a pressure data sensor, or a photoelectric sensor). The kettle is a container used to heat, bake, or dry a substance or material.

[0017] In automated operations, it is important to classify objects photographed in plants. Therefore, to automate classification, it is necessary to create (construct) a learning model that generates classification information for objects using artificial intelligence (AI). When building or updating such a learning model, a large amount of training data is required. Therefore, an efficient method for generating training data is required.

[0018] A computer 1 according to this embodiment acquires an image of an object photographed in a plant. The computer 1 provides the acquired image to a language model to generate classification information relating to the classification of the object.

[0019] The computer 1 acquires training data including a plurality of sets of images of objects and classification information of the objects generated by a language model. Based on the acquired training data, the computer 1 trains a learning model that generates classification information regarding the classification of an object when an image of the object photographed in a plant is input. The language model (language model 171) and the learning model (object classification model 172) will be described later.

[0020] 2 is a block diagram showing an example of the configuration of the computer 1. The computer 1 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, a display unit 15, a reading unit 16, and a large-capacity storage unit 17. Each component is connected by a bus B.

[0021] The control unit 11 includes an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), or a quantum processor. The control unit 11 reads and executes a control program 1P (program product) stored in the storage unit 12, thereby performing various information processing or control processing related to the computer 1. The control program 1P described in this embodiment may be provided on a recording medium or may be distributed from an external computer.

[0022] It should be noted that the control program 1P can be deployed to run on a single computer, or on multiple computers located at one site, or distributed across multiple sites and interconnected by a communications network.

[0023] 2, the control unit 11 is described as a single processor, but it may be a multi-processor. The control unit 11 may execute various information processing or control processes by the same processor within the computer 1, or may execute various processes by different processors within the computer 1.

[0024] The storage unit 12 includes memory elements such as RAM (Random Access Memory) and ROM (Read Only Memory), and stores the control program 1P or data required for the control unit 11 to execute processing. The storage unit 12 also temporarily stores data required for the control unit 11 to execute arithmetic processing. The communication unit 13 is a communication module for performing communication-related processing, and transmits and receives information to and from the imaging device 2 or device 3 via the network N.

[0025] The input unit 14 may be a keyboard, a mouse, or a touch panel integrated with the display unit 15. The display unit 15 is a liquid crystal display, an organic EL (electroluminescence) display, or the like, and displays various information according to instructions from the control unit 11.

[0026] The reading unit 16 reads a portable storage medium 1a including a CD (Compact Disc)-ROM or a DVD (Digital Versatile Disc)-ROM. The control unit 11 may read the control program 1P from the portable storage medium 1a via the reading unit 16 and store it in the mass storage unit 17. Alternatively, the control unit 11 may download the control program 1P from another computer via a network N or the like and store it in the mass storage unit 17. Furthermore, the control unit 11 may read the control program 1P from the semiconductor memory 1b.

[0027] The mass storage unit 17 includes a recording medium such as a hard disk drive (HDD) or a solid state drive (SSD), etc. The mass storage unit 17 includes a language model 171, an object classification model 172, a training data database (DB) 173, and a class definition file 174.

[0028] The language model 171 is a language generation model constructed by pre-training using large-scale text data (dataset). As the language model 171, for example, large language models (LLMs) such as Transformer, ALBERT (A Lite BERT), GPT (Generative Pre-trained Transformer)-2, GPT-3, GPT-4, LLaVA (Large Language and Vision Assistant), MiniGPT-4, or BERT (Bidirectional Encoder Representations from Transformers) can be used.

[0029] The object classification model 172 is a generator (estimator) that generates classification information regarding the classification of an object based on an image of the object photographed in a plant, and is a trained model generated by machine learning. The training data DB 173 stores training data for constructing (generating) the object classification model 172. The class definition file 174 is a file that defines classification classes for identifying the classification of an object.

[0030] In this embodiment, the storage unit 12 and the large-capacity storage unit 17 may be configured as an integrated storage device. Furthermore, the large-capacity storage unit 17 may be configured with a plurality of storage devices. Furthermore, the large-capacity storage unit 17 may be an external storage device connected to the computer 1.

[0031] The computer 1 may perform various information processing and control processing by itself or may perform the processing in a distributed manner across multiple computers. The computer 1 may also be implemented by multiple virtual machines installed in a single computer, or may be implemented using a cloud server.

[0032] 3 is an explanatory diagram showing an example of a record layout of the training data DB 173. The training data DB 173 includes an input data string and an output data string. The input data string stores images of objects photographed in a plant. The output data string stores classification information relating to the classification of the objects. The classification information will be described later.

[0033] 4 is a block diagram showing an example of the configuration of the device 3. The device 3 includes a control unit 31, a storage unit 32, and a communication unit 33. Note that the control unit 31, the storage unit 32, and the communication unit 33 are similar to the control unit 11, the storage unit 12, and the communication unit 13 of the computer 1, and therefore a description thereof will be omitted.

[0034] 5 is an explanatory diagram showing an example of the data configuration of the class definition file 174. A class definition file 174 that defines classification classes for identifying the classification of objects is stored in advance in the mass storage unit 17 of the computer 1. Note that in this embodiment, the contents of the classification classes are stored in the form of a file, but this is not limiting and they may also be stored in a database.

[0035] As shown in the figure, as an example, the classification classes are classified according to the attributes of the object, and include a white class (1), a yellow class (2), a yellow-black class (3), a red class (4), a red-black class (5), and a black class (6). The white class is defined as "background color: invisible, color inside the pot: white and yellow, overall brightness: bright." The yellow class is defined as "background color: visible, color inside the pot: white and yellow, overall brightness: bright."

[0036] The yellow-black class defines the content as "Background color: visible, color inside the pot: yellow and black, overall brightness: bright." The red class defines the content as "Background color: visible, color inside the pot: red or orange, overall brightness: normal." The red-black class defines the content as "Background color: visible, color inside the pot: red and black, overall brightness: dark." The black class defines the content as "Background color: visible or invisible, color inside the pot: black, overall brightness: dark."

[0037] It should be noted that the classification classes are not limited to those described above, and other classification classes may be set according to, for example, attribute information (color, state, etc.) of the object.

[0038] The above-mentioned storage formats of the DB and files are merely examples, and other storage formats may be used as long as the relationships between the data are maintained.

[0039] 6 is an explanatory diagram illustrating the process of generating training data. Computer 1 acquires images of objects photographed in a plant from imaging device 2. Based on the acquired images of the objects, computer 1 generates prompts to be given to language model 171.

[0040] The language model 171 is a language model that uses images of objects taken in a plant, and is used as a program module that is part of artificial intelligence software. The language model 171 in this embodiment is a constructed language model (language generation model) that receives as input an image of an object and a prompt including instructions (commands) for generating classification information related to the classification of the object, and outputs the classification information of the object.

[0041] Instead of storing the language model 171 in the mass storage unit 17, the computer 1 may access an external language processing server or language processing platform and read it out.

[0042] A prompt is an instruction or input sentence created in a format that can be understood by the language model 171 and given as input to the language model 171. The language model 171 interprets the input image of the object and the prompt, and outputs an appropriate response (e.g., classification information of the object).

[0043] Specifically, when an image of an object and a prompt are input, the language model 171 analyzes the input image of the object and identifies the object contained in the image. The language model 171 generates classification text (labels) for each of multiple attributes (such as color, shape, size, or state) of the identified object. For example, the classification text for the attribute "brightness" may be "bright," the classification text for the attribute "upper half color" may be "yellow," the classification text for the attribute "lower half color" may be "white," or the classification text for the attribute "background" may be "invisible / dark." The language model 171 obtains additional information from the context of the prompt to generate a more accurate response.

[0044] As an example, language model 171 divides the prompt into tokens to convert it into a format that can be processed by language model 171. Language model 171 performs context understanding by calculating the association of each token with other tokens in the prompt.

[0045] The language model 171 classifies objects identified from images and generates responses to prompts based on linguistic knowledge obtained through pre-learning and fine-tuning. For example, the language model 171 selects optimal tokens using generation techniques such as greedy decoding, beam search, or sampling. The language model 171 decodes the selected tokens to return them to text format and generates classification information regarding the classification of the objects.

[0046] In this embodiment, the prompt includes instructions for generating classification information for the object, such as instructions for describing the image, instructions for brightness, instructions for colors of the top and bottom halves, or inquiries about background visibility.

[0047] As an example, the generated prompt is: "Please explain this image. Specifically, please tell me the brightness, the color of the upper half, and the color of the lower half. Or, "Can you see the background?"

[0048] The computer 1 inputs the image of the object and the generated prompt into the language model 171, and generates text relating to the classification of each of the multiple attributes of the object as classification information relating to the classification of the object.

[0049] As shown in the figure, as an example, the generated classification information is Brightness: Bright Upper half color: Yellow Lower half color: White Background: invisible / dark" is also acceptable.

[0050] The computer 1 identifies a classification class that indicates the classification of the object based on the classification-related text for each of the multiple attributes of the generated object. Specifically, the computer 1 acquires the class definition file 174 from the mass storage unit 17. The computer 1 reads the classification class stored in the acquired class definition file 174. The computer 1 identifies the classification class that indicates the classification of the object based on the classification information generated by the language model 171, with reference to the read classification class. For example, the identified classification class is the "white class."

[0051] In the present embodiment, an example of an image of an object has been described, but the present invention is not limited to this and may be a video of the object. For example, the computer 1 acquires a video of the object captured by the imaging device 2. The computer 1 extracts a plurality of frame images from the acquired video at predetermined intervals (for example, 10 seconds). The computer 1 provides the extracted plurality of frame images to the language model 171, thereby generating classification information regarding the classification of the object.

[0052] The computer 1 generates training data in which images of objects are associated with classification information including the identified classification class. The computer 1 stores the generated training data in the training data DB 173. The input data string of the training data DB 173 stores images of objects. The output data string stores classification information including classification classes that identify the classification of the objects and text related to the classification for each of multiple attributes of the objects. Note that the output data string may include either the classification class or the text related to the classification.

[0053] Next, a process for training the object classification model 172 using the generated training data will be described. The object classification model 172 is used as a program module that is part of artificial intelligence software. The object classification model 172 is a generator (estimator) that has a pre-constructed neural network that receives an image of an object as input and outputs classification information related to the classification of the object.

[0054] The computer 1 performs deep learning to learn feature quantities of objects in an image to generate the object classification model 172. For example, the object classification model 172 is a convolution neural network (CNN), which has an input layer that receives input of an image of the object, an output layer that outputs classification information of the object, and an intermediate layer that has been trained by backpropagation.

[0055] The input layer has a plurality of neurons that receive input of pixel values ​​of each pixel included in the image of the object, and passes the input pixel values ​​to the intermediate layer. The intermediate layer has a plurality of neurons that extract features of the image of the object, and passes the extracted image features to the output layer.

[0056] The intermediate layer is configured by alternating convolution layers that convolve the pixel values ​​of each pixel input from the input layer and pooling layers that map the pixel values ​​convolved in the convolution layers, thereby compressing the pixel information of the image of the target object and ultimately extracting image features.

[0057] The intermediate layer then generates (estimates) classification information about the classification of the object corresponding to the image of the object using a fully connected layer whose parameters are learned by backpropagation. The generated classification information is output to an output layer having multiple neurons.

[0058] Note that the image of the object may be passed through an alternating convolution layer and a pooling layer to extract features, and then input to the input layer.

[0059] Instead of CNN, any object detection algorithm may be used, such as RCNN (Regions with Convolutional Neural Network), Fast RCNN, Faster RCNN, SSD (Single Shot Multibook Detector), YOLO (You Only Look Once), SVM (Support Vector Machine), Bayesian network, Transformer network, or regression tree.

[0060] For example, the computer 1 uses training data stored in the training data DB 173 to train (learn) the object classification model 172. The training data is data including a plurality of sets of images of objects and classification information generated by the language model 171.

[0061] Specifically, the computer 1 inputs images of objects as training data to the input layer, performs arithmetic processing in the intermediate layer, and outputs classification information of the objects from the output layer. The output layer includes, for example, a sigmoid function or a softmax function, and outputs estimated classification information of the objects based on the feature values ​​output from the intermediate layer.

[0062] The classification information includes, for example, the probability of a classification class (such as a white class, yellow class, yellow-black class, red class, red-black class, or black class), or information about the classification class to which the object belongs (such as the name of the classification class).

[0063] The computer 1 compares the classification information output from the output layer with the classification information of the object in the training data, i.e., the correct value, and optimizes the parameters used in the calculation process in the intermediate layer so that the output value from the output layer approaches the correct value. The parameters are, for example, weights (coupling coefficients) between neurons or coefficients of activation functions used in each neuron. There are no particular limitations on the method for optimizing the parameters, but for example, the computer 1 optimizes various parameters using the backpropagation algorithm.

[0064] The computer 1 performs the above processing on the image of each object included in the training data, and generates (updates) the object classification model 172. As a result, for example, the computer 1 trains the object classification model 172 using the training data, thereby generating a model that can generate (estimate) classification information of objects.

[0065] When the computer 1 acquires an image of an object, it inputs the acquired image of the object to the object classification model 172. The computer 1 performs a calculation process to extract features of the image of the object in the intermediate layer of the object classification model 172. The computer 1 inputs the extracted features to the output layer of the object classification model 172, and acquires as an output an estimation result that estimates classification information of the object.

[0066] As an example, for an image of an object, the estimation results output are "0.87", "0.03", "0.05", "0.02", "0.02", and "0.01" for the "white class", "yellow class", "yellow-black class", "red class", "red-black class", and "black class", respectively.

[0067] Furthermore, the estimation result can be output using a predetermined threshold. For example, if the computer 1 determines that the probability value of the "white class" (0.87) is equal to or greater than a predetermined threshold (e.g., 0.80), the computer 1 outputs the white class as the estimation result. Note that, without using the above-described threshold, the computer 1 may output the classification class corresponding to the highest probability value from the probability values ​​of each classification class estimated by the object classification model 172 as the estimation result.

[0068] 7 is a flowchart showing the processing steps for generating training data. The control unit 11 of the computer 1 acquires an image of an object photographed in a plant from the imaging device 2 via the communication unit 13 (step S101). The control unit 11 generates a prompt including a generation instruction for generating classification information of the object (step S102).

[0069] The control unit 11 inputs the acquired image of the object and the generated prompt to the language model 171 (step S103), and outputs classification information regarding the classification of the object (step S104). The classification information is text regarding the classification of each of a plurality of attributes of the object (for example, "Background color: visible, Color inside the pot: white and yellow, Overall brightness: bright").

[0070] Based on the output classification information, control unit 11 identifies a classification class that indicates the classification of the object (step S105). Specifically, control unit 11 acquires class definition file 174 from mass storage unit 17. Control unit 11 reads the classification class stored in acquired class definition file 174. Based on the classification information generated by language model 171, control unit 11 refers to the read classification class to identify the classification class that indicates the classification of the object.

[0071] The control unit 11 generates training data in which the image of the object is associated with classification information including the identified classification class (classification information obtained by the processes of steps S104 to S105) (step S106). The control unit 11 stores the generated training data in the training data DB 173 of the mass storage unit 17 (step S107).

[0072] The control unit 11 acquires training data including a plurality of sets of object images and classification information generated by the language model 171 from the training data DB 173 (step S108). The control unit 11 uses the acquired training data to train the object classification model 172 (step S109). The control unit 11 ends the process.

[0073] 8 is a flowchart showing the processing procedure when both the first classification information and the second classification information are output. The first classification information is classification information generated by language model 171 based on an image of the object. The second classification information is classification information generated by object classification model 172 based on an image of the object. Note that the same reference numerals are used to designate the same contents as in FIG. 7, and the description thereof will be omitted.

[0074] After executing the process of step S103, the control unit 11 of the computer 1 outputs the first classification information from the language model 171 (step S111). The first classification information includes text (labels) relating to classifications of multiple attributes of the object, and may be, for example, "Brightness: Bright, Upper half color: Yellow, Lower half color: White, Background: Invisible / Dark."

[0075] The control unit 11 inputs the image of the object into the object classification model 172 (step S112) and outputs second classification information (step S113). The second classification information is, for example, the probability of a classification class (white class, yellow class, yellow-black class, red class, red-black class, black class, etc.). For example, the second classification information may be "white class: 0.87, yellow class: 0.03, yellow-black class: 0.05, red class: 0.02, red-black class: 0.02, black class: 0.01". The control unit 11 ends the process.

[0076] According to this embodiment, by providing an image of an object taken in a plant to the language model 171, it becomes possible to generate classification information relating to the classification of the object.

[0077] According to this embodiment, it is possible to train the object classification model 172 based on training data including a plurality of sets of images of objects and classification information generated by the language model 171.

[0078] According to this embodiment, by combining photographed images of objects in a plant with classification information from the language model 171, training of the object classification model 172 is performed efficiently, making it possible to obtain more detailed and accurate classification results.

[0079] According to this embodiment, by outputting both the first classification information generated by the language model 171 and the second classification information generated by the object classification model 172, it is possible to improve the reliability of the classification of the object.

[0080] (Embodiment 2) The second embodiment relates to a form in which a plurality of images of an object is generated by a language model, and training data including a plurality of sets of the generated images and classification information relating to the classification of the object is generated. Note that a description of the contents overlapping with the first embodiment will be omitted.

[0081] 9 is an explanatory diagram illustrating the process of generating training data in embodiment 2. Based on an image of an object, a language model 171 can generate multiple images different from the image. Specifically, the computer 1 acquires an image of the object captured by the imaging device 2. The computer 1 generates a prompt including information about the acquired image of the object (for example, the file name of the image) and an instruction to generate multiple images different from the image of the object.

[0082] As an example, the generated prompt is: "Image of object: "01.jpg" "Simulate the stirring process at different temperatures and stirring speeds to generate different images." The prompt may also include instructions for the kiln background, the color inside the kiln, or the overall brightness.

[0083] The computer 1 inputs the acquired image of the object and the generated prompt into the language model 171, and generates multiple images of the object. In this case, the language model 171 may be realized by, for example, diffusion models.

[0084] The computer 1 acquires classification information related to the classification of an object. Specifically, the computer 1 generates a prompt to be given to the language model 171 based on the acquired image of the object. The prompt includes a generation instruction for generating classification information for the object. The computer 1 inputs the image of the object and the generated prompt into the language model 171, and generates text related to the classification of each of multiple attributes of the object as classification information for the object. Based on the generated classification information, the computer 1 refers to the class definition file 174 stored in the mass storage unit 17 and identifies a classification class that indicates the classification of the object.

[0085] The computer 1 generates training data in which the generated images are associated with classification information including the identified classification class. The computer 1 stores the generated training data in the training data DB 173. The input data string of the training data DB 173 stores the images of the object generated by the language model 171. The output data string stores the classification information of the object.

[0086] Next, a process of training object classification model 172 using the generated training data will be described. Specifically, computer 1 acquires training data stored in training data DB 173. The training data is data including a plurality of sets of a plurality of images of objects and classification information generated by language model 171.

[0087] The computer 1 inputs a plurality of images of an object as training data into the input layer, performs arithmetic processing in the intermediate layer, and outputs classification information of the object from the output layer. As in the first embodiment, the classification information includes, for example, the probability of the classification class, or information about the classification class to which the object belongs (for example, the name of the classification class), etc.

[0088] The computer 1 compares the classification information output from the output layer with the classification information of the object in the training data, i.e., the correct value, and optimizes the parameters used for the calculation process in the intermediate layer so that the output value from the output layer approaches the correct value. The computer 1 performs the above process on multiple images of the object included in the training data, and generates (updates) the object classification model 172.

[0089] Fig. 10 is a flowchart showing the processing procedure for generating training data in the second embodiment. Note that the same reference numerals are used to denote the same parts as in Fig. 7, and the description thereof will be omitted. After executing the process of step S105, the control unit 11 of the computer 1 generates a second prompt for generating multiple images different from the image of the object (step S121). The second prompt includes information about the image of the object (for example, the file name of the image), a command to generate multiple images different from the image of the object, and the like.

[0090] Control unit 11 inputs the acquired image of the object and the generated second prompt to language model 171 (step S122), and generates a plurality of images different from the image of the object (step S123). Control unit 11 generates training data in which the generated plurality of images are associated with the classification information obtained by the processes of steps S104 to S105 (step S124). Control unit 11 executes the process of step S107.

[0091] According to this embodiment, it is possible to use the language model 171 to generate a plurality of images that are different from the image of the object.

[0092] According to this embodiment, it is possible to train the object classification model 172 based on training data including a plurality of sets of a plurality of images and object classification information.

[0093] (Embodiment 3) The third embodiment relates to a form in which a second object classification model, which will be described later, is trained based on time-series data of objects and classification information of the objects generated by a language model. Note that a description of the contents overlapping with the first and second embodiments will be omitted.

[0094] Fig. 11 is a block diagram showing an example of the configuration of a computer 1 in embodiment 3. Note that the same reference numerals are used to denote the same components as those in Fig. 2, and descriptions thereof will be omitted. The mass storage unit 17 includes a second object classification model 175. The second object classification model 175 is a generator (estimator) that generates classification information regarding the classification of an object based on time-series data of the object obtained in the plant, and is a trained model generated by machine learning.

[0095] For example, the second object classification model 175 may be an RCNN, but instead of an RCNN, any object detection algorithm such as Seq2Seq (Sequence to Sequence), CNN, Fast RCNN, Faster RCNN, SSD, YOLO, SVM, Bayesian network, Transformer network, or regression tree may be used.

[0096] The computer 1 acquires time-series data of an object obtained in a plant from the device 3. The time-series data is time-series sensor data obtained by various sensors mounted on the device 3. The sensors include, for example, a weight measurement sensor, a temperature sensor, an acceleration sensor, a pressure sensor, a gyro sensor, an X-ray line sensor, or a humidity sensor.

[0097] For example, the computer 1 may acquire time-series weight sensor data, temperature sensor data, or acceleration sensor data of an object within a predetermined period (e.g., 20 seconds) from the device 3. The computer 1 generates a prompt including the acquired time-series data of the object and instructions for generating classification information of the object.

[0098] In this embodiment, the classification information output from the language model 171 includes, for example, a manufacturing process state classification, a characteristic classification, or a quality classification. The manufacturing process state classification is inferred from acceleration data or temperature changes, etc., and includes a manufacturing stage classification (e.g., mixing, dissolving, or kneading), a normal or abnormal state (e.g., a sudden rise in temperature, a sudden rise in pressure, or abnormal vibration), etc.

[0099] The property classification is inferred from acceleration sensor data or temperature sensor data, etc., and includes the shape, size, or physical properties of the object (e.g., density, hardness, or surface smoothness), etc. The quality classification is inferred from weight change or temperature change, etc., and includes quality class (e.g., high quality, medium quality, or low quality), or classification based on the composition or particle size of the object (e.g., classification by cocoa content), etc.

[0100] In the following, an example will be described in which the classification information output from the language model 171 is a "state classification of the manufacturing process," but the same can be applied to other types of classification information.

[0101] As an example, the generated prompt is: "Time series data includes weight, temperature, and acceleration data of the object. Weight sensor data: [10.2, 10.5, 10.3, 10.6, 10.8, ...] Temperature sensor data: [25.3, 25.5, 25.7, 30.9, 32.1, ...] Accelerometer data: [[0.1, 0.2, 9.8], [0.2, 0.3, 9.9], [0.3, 0.4, 9.7], ...] "Please use this time series data to provide classification information for the object."

[0102] The prompt may include a graph generated from time-series data such as weight sensor data, temperature sensor data, or acceleration sensor data. The computer 1 inputs the time-series data of the object and the generated prompt into the language model 171, and generates text related to the state of the manufacturing process of the object as classification information related to the classification of the object.

[0103] As an example, the generated classification information is: Condition: Temperature spike No sudden rise or fall in pressure "No abnormal vibrations" is acceptable.

[0104] Based on the generated text relating to the state of the manufacturing process of the object, the computer 1 refers to the class definition file 174 stored in the mass storage unit 17 and identifies a classification class that indicates the classification of the object. An example of the data configuration of the class definition file 174 in this embodiment is classified according to the state of the manufacturing process of the object, and includes a normal class, a temperature abnormality class, a pressure abnormality class, and a vibration abnormality class.

[0105] The normal class is defined as "the temperature does not rise or fall suddenly, the pressure does not rise or fall suddenly, and the vibration is normal." The temperature abnormality class is defined as "the temperature rises or falls suddenly, the pressure does not rise or fall suddenly, and the vibration is normal." The pressure abnormality class is defined as "the temperature does not rise or fall suddenly, the pressure does not rise or fall suddenly, and the vibration is normal." The vibration abnormality class is defined as "the temperature does not rise or fall suddenly, the pressure does not rise or fall suddenly, and the vibration is abnormal." For example, the identified classification class is the "temperature abnormality class."

[0106] Computer 1 generates training data in which time-series data of the object is associated with classification information including the identified classification class. Computer 1 stores the generated training data in training data DB 173. The time-series data of the object is stored in the input data string of training data DB 173, and the classification information of the object is stored in the output data string.

[0107] Next, a process of training the second object classification model 175 using the generated training data will be described. Specifically, the computer 1 acquires the training data accumulated in the training data DB 173. The training data is data including a plurality of sets of time-series data of objects and classification information generated by the language model 171.

[0108] The computer 1 inputs time-series data of an object, which is training data, into an input layer, undergoes arithmetic processing in an intermediate layer, and outputs classification information of the object from an output layer. As in the first embodiment, the classification information includes, for example, the probability of a classification class (e.g., normal class, temperature abnormality class, pressure abnormality class, or vibration abnormality class), or information about the classification class to which the object belongs (e.g., the name of the classification class), etc.

[0109] The computer 1 compares the classification information output from the output layer with the classification information of the objects in the training data, i.e., the correct value, and optimizes the parameters used for the calculation process in the intermediate layer so that the output value from the output layer approaches the correct value. The computer 1 performs the above process on the time-series data of the objects included in the training data, and generates (updates) the second object classification model 175.

[0110] 12 is a flowchart showing a processing procedure for generating training data in embodiment 3. The control unit 11 of the computer 1 acquires time-series data (e.g., weight sensor data, temperature sensor data, or acceleration sensor data) of an object within a predetermined period (e.g., 20 seconds) from the device 3 via the communication unit 13 (step S131). The control unit 11 generates a prompt to be given to the language model 171 (step S132). The prompt includes the time-series data of the object and a generation instruction for generating classification information of the object.

[0111] Control unit 11 inputs the generated prompt into language model 171 (step S133) and outputs classification information regarding the classification of the object (step S134). The classification information is, for example, text regarding the state of the manufacturing process of the object. Based on the output classification information, control unit 11 refers to class definition file 174 stored in mass storage unit 17 and identifies a classification class that indicates the classification of the object (step S135).

[0112] The control unit 11 generates training data in which the time-series data of the object is associated with classification information including the identified classification class (step S136). The control unit 11 stores the generated training data in the training data DB 173 of the mass storage unit 17 (step S137).

[0113] The control unit 11 acquires training data including a plurality of sets of time-series data of objects and classification information generated by the language model 171 from the training data DB 173 (step S138). The control unit 11 uses the acquired training data to train the second object classification model 175 (step S139). The control unit 11 ends the process.

[0114] According to this embodiment, by providing the time-series data of an object to the language model 171, it becomes possible to generate classification information of the object.

[0115] According to this embodiment, it is possible to train the second object classification model 175 based on training data including a plurality of sets of time-series data of objects and classification information of the objects.

[0116] <Variation 1> The process of generating classification information for an object by providing the image and time-series data of the object to the language model 171 will be described.

[0117] 13 is a flowchart showing a processing procedure for generating classification information of an object based on an image and time-series data of the object. The control unit 11 of the computer 1 acquires an image of the object taken in the plant from the imaging device 2 via the communication unit 13 (step S141). The control unit 11 acquires time-series data of the object within a predetermined period (e.g., 20 seconds) from the device 3 via the communication unit 13 (step S142).

[0118] The control unit 11 generates a prompt to be given to the language model 171 (step S143). The prompt includes time-series data of the object, a generation instruction for generating classification information of the object, and the like.

[0119] As an example, the generated prompt is: "Time series data includes weight, temperature, and acceleration data of the object. Weight sensor data: [10.2, 10.5, 10.3, 10.6, 10.8, ...] Temperature sensor data: [25.3, 25.5, 25.7, 30.9, 32.1, ...] Accelerometer data: [[0.1, 0.2, 9.8], [0.2, 0.3, 9.9], [0.3, 0.4, 9.7], ...] Based on this image and time series data, please tell us the classification information of the object. Also, please tell me the brightness, color of the upper half, and color of the lower half. Or, "Can you see the background?"

[0120] The control unit 11 inputs an image of the object and a prompt including time-series data of the object to the language model 171 (step S144), and outputs classification information regarding the classification of the object (step S145).

[0121] As an example, the generated classification information is: Condition: Temperature spike No sudden rise or fall in pressure No abnormal vibrations Brightness: Normal Upper half color: Red Bottom half color: Orange Background: Visible" is also acceptable.

[0122] Based on the output classification information, control unit 11 refers to class definition file 174 stored in mass storage unit 17, and identifies the classification class that indicates the classification of the object (step S146). Control unit 11 then ends the process.

[0123] An example of the data configuration of the class definition file 174 in this modification includes classification classes classified according to the state of the manufacturing process of the object (for example, normal class, temperature abnormality class, pressure abnormality class, and vibration abnormality class), and classification classes classified according to the attributes of the object (for example, white class, yellow class, yellow-black class, red class, red-black class, and black class). Note that the content defined for each classification class is the same as in the first and second embodiments, and therefore description thereof will be omitted. For example, the identified classification class is "temperature abnormality class + red class."

[0124] According to this modification, by providing the image and time-series data of an object to the language model 171, it becomes possible to generate classification information of the object.

[0125] (Embodiment 4) The fourth embodiment relates to a form in which the state of an object is output by providing an image of the object to the language model 171. Note that a description of the contents that overlap with the first to third embodiments will be omitted.

[0126] 14 is a flowchart showing a processing procedure for outputting the state of an object. The control unit 11 of the computer 1 acquires an image of the object photographed in the plant from the imaging device 2 via the communication unit 13 (step S151). The control unit 11 generates a prompt to be given to the language model 171 (step S152). The prompt includes an output instruction for outputting the state of the object, etc.

[0127] The state of an object is classified into, for example, a color state, a shape state, or a size state. The color state indicates whether the object has a normal color, an abnormal discoloration, or an uneven hue. The shape state indicates whether the object maintains a normal shape, is distorted or deformed, or has an irregular shape. The size state indicates whether the object is within a predetermined size range or is outside that range.

[0128] As an example, the generated prompt is: It would also be fine to say, "Based on an image of the object, please tell me the state of the object according to its attributes such as color, shape, or size."

[0129] The control unit 11 inputs the acquired image of the object and the generated prompt to the language model 171 (step S153), and outputs the state of the object (step S154).

[0130] For example, the output state of the object is: "Object status: The color shows abnormal discoloration. The shape is distorted. It can also be "Size out of range."

[0131] The control unit 11 determines whether the state of the object output from the language model 171 is normal (step S155). If the state is normal (YES in step S155), the control unit 11 outputs a command to start control of the object to the device 3 via the communication unit 13 (step S156). The control includes, for example, temperature control (rising or lowering, etc.), operation control (lowering or rotating a shaft, etc.), or squeezing control.

[0132] If the state is not normal (NO in step S155), control unit 11 generates a second prompt including an instruction to output the reason why control cannot be started (step S157).

[0133] As an example, the generated prompt is: "Please explain why you are unable to initiate control over the object." is also acceptable.

[0134] The control unit 11 inputs the generated second prompt to the language model 171 (step S158), and outputs the reason why control over the object cannot be started (step S159). The control unit 11 then ends the process.

[0135] For example, the reason for the output is "Control cannot be initiated. The object's condition is abnormal. The color is abnormally discolored, the shape is distorted, and the size is outside the specified range.

[0136] According to this embodiment, by providing an image of an object to the language model 171, it is possible to output the state of the object.

[0137] According to this embodiment, control is started only when the state of the object is normal, thereby preventing unnecessary errors and improving the reliability of the entire system.

[0138] According to this embodiment, by outputting the reason why control cannot be started when the condition is not normal, problems can be detected and resolved early, thereby improving the quality of the object.

[0139] The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.

[0140] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used. [Explanation of symbols]

[0141] 1. Information processing equipment (computer) 11 Control section 12 Storage section 13 Communications Department 14 Input section 15 Display section 16 Reading unit 17 Mass storage 171 language models 172 Object Classification Model 173 Training Data DB 174 Class Definition File 175 Second Object Classification Model 1a Portable storage media 1b semiconductor memory 1P control program 2. Imaging device 3 equipment 31 Control Unit 32 Storage section 33 Communications Department

Claims

1. Acquire images of objects taken at the plant, By providing the acquired image to a language model, classification information regarding the classification of the object is generated. An information processing method in which processing is performed by a computer.

2. The language model generates text regarding classification for each of a plurality of attributes of the object; Identifying a classification of the object based on the generated text of multiple attributes The information processing method according to claim 1 .

3. acquiring training data including a plurality of sets of images of the object and classification information generated by the language model; Based on the acquired training data, a learning model is trained that generates classification information regarding the classification of an object when an image of the object photographed in the plant is input.

3. The information processing method according to claim 1 or 2.

4. generating a plurality of images by providing the language model with an image of the object and a prompt including instructions to generate a plurality of images different from the image; acquiring training data including a plurality of sets of the generated images and classification information regarding the classification of the object; Based on the acquired training data, a learning model is trained that generates classification information regarding the classification of an object when an image of the object photographed in the plant is input.

3. The information processing method according to claim 1 or 2.

5. acquiring an image of the object; The acquired image of the object is input to the language model, or to the language model and the learning model, and classification information obtained from the language model is output, or first classification information obtained from the language model and second classification information obtained from the learning model are output. The information processing method according to claim 3 .

6. Acquire time-series data of the object obtained at the plant; The acquired time-series data is provided to the language model to generate classification information regarding the classification of the object.

3. The information processing method according to claim 1 or 2.

7. acquiring training data including a plurality of sets of time-series data of the object and classification information generated by the language model; Based on the acquired training data, a second learning model is trained to generate classification information regarding classification of the object when time-series data of the object obtained in the plant is input. The information processing method according to claim 6.

8. acquiring images and time series data of the object; The acquired image and time-series data of the object are provided to the language model, thereby generating classification information regarding the classification of the object. The information processing method according to claim 6.

9. outputting a state of the object by providing an image of the object to the language model; When the output state is normal, control of the object is started; If the output state is not normal, the reason why the process cannot be started is output using the language model.

3. The information processing method according to claim 1 or 2.

10. Acquire images of objects taken at the plant, By providing the acquired image to a language model, classification information regarding the classification of the object is generated. A program that causes a computer to perform a process.

11. An information processing device including a control unit, The control unit Acquire images of objects taken at the plant, By providing the acquired image to a language model, classification information regarding the classification of the object is generated. Information processing device.

Citation Information

Patent Citations

  • Information processing method, computer program, information processing device, and information processing system

    JP7223194B1